Candidate Derivation for Affine Merge Modes in Video Coding

By deriving control point motion vectors from neighboring blocks using affine models, the methods enhance the accuracy of motion vector prediction, improving the efficiency of video encoding and decoding processes.

JP7804765B2Active Publication Date: 2026-01-22BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024526732
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-08
Filing Date
2022-11-08
Publication Date
2026-01-22
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in accurately deriving motion vector candidates for affine motion prediction modes, leading to inefficiencies in video encoding and decoding processes.

Method used

The proposed methods involve obtaining parameters from neighboring blocks to construct affine models and derive control point motion vectors (CPMVs) using inheritance-based and construction-based derivation methods, enhancing the accuracy of motion vector prediction.

Benefits of technology

Improved motion vector candidate derivation leads to more efficient video encoding and decoding processes, reducing bit rate requirements while maintaining video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804765000016
    Figure 0007804765000016
  • Figure 0007804765000017
    Figure 0007804765000017
  • Figure 0007804765000018
    Figure 0007804765000018
Patent Text Reader

Abstract

A method of video decoding, a method of video encoding, an apparatus thereof, and a non-transitory computer-readable storage medium are provided. The method of video decoding includes obtaining one or more first parameters based on a first neighboring block of a current block, and obtaining one or more second parameters based on the first neighboring block and / or a second neighboring block of the current block. The method may further include constructing one or more affine models by using the one or more first parameters and the one or more second parameters, and obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is filed on and claims priority to U.S. Provisional Patent Application No. 63 / 277,148, entitled "Candidate Derivation for Affine Merge Mode in Video Coding," filed November 8, 2021, which is incorporated by reference in its entirety for all purposes.

[0002] This disclosure relates to video coding and compression, and more particularly, but not exclusively, to methods and apparatus for improving affine merge candidate derivation for affine motion prediction modes in video encoding or decoding processes. [Background technology]

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some currently known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMediaVideo1 (AV1) was developed by the Alliance for Open Media (AOM) as the successor to its predecessor, VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Chinese Audio and Video Coding Standard Workgroup. Most of the existing video coding standards build on the well-known hybrid video coding framework, i.e., using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce the redundancy present in a video image or sequence, and transform coding to reduce the energy of the prediction error. An important goal of video coding techniques is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation of video quality.

[0004] The first-generation AVS standard includes the Chinese national standards "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+). It can provide approximately 50% bitrate savings at the same perceptual quality compared to the MPEG-2 standard. The video part of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second-generation AVS standard includes the Chinese national standard series "Information Technology, Efficient Multimedia Coding" (known as AVS2), which is primarily targeted at the transmission of extra HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. AVS2 was published as a Chinese national standard in May 2016. Meanwhile, the video part of the AVS2 standard has been submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for applications. The AVS3 standard is a new generation video coding standard for UHD video applications that aims to exceed the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was finalized, which provides approximately 30% bitrate savings over the HEVC standard. Currently, there is one reference software, called the High Performance Model (HPM), maintained by the AVS group to certify reference implementations of the AVS3 standard. Summary of the Invention [Problem to be solved by the invention]

[0005] This disclosure provides examples of techniques related to improving motion vector candidate derivation for motion prediction modes in a video encoding or decoding process. [Means for solving the problem]

[0006] According to a first aspect of the present disclosure, there is provided a method for video decoding. The method includes: obtaining one or more first parameters based on first neighbor blocks of a current block; and obtaining one or more second parameters based on the first neighbor blocks and / or second neighbor blocks of the current block. The method may further include constructing one or more affine models by using the one or more first parameters and the one or more second parameters; and obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models.

[0007] According to a second aspect of the present disclosure, there is provided a method for video decoding. The method may include obtaining a plurality of motion vector candidates from a history-based motion vector prediction (HMVP) table, where the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate. The method may further include obtaining a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, and obtaining a plurality of control point motion vectors (CPMVs) for a current block based on the plurality of CPMVs of the virtual block.

[0008] According to a third aspect of the present disclosure, there is provided a method of video decoding, which may include obtaining one or more candidate motion vectors from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from one side of the current block, and obtaining one or more CPMVs for the current block based on the one or more candidate motion vectors.

[0009] According to a fourth aspect of the present disclosure, there is provided a method of video encoding. The method may include determining one or more first parameters based on first neighboring blocks of a current block and determining one or more second parameters based on the first neighboring blocks and / or second neighboring blocks of the current block. Further, the method may include constructing one or more affine models by using the one or more first parameters and the one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.

[0010] According to a fifth aspect of the present disclosure, there is provided a method of video encoding. The method may include determining a plurality of motion vector candidates from a history-based motion vector prediction (HMVP) table, where the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate. Further, the method may include determining a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, and obtaining a plurality of CPMVs for the current block based on the plurality of CPMVs of the virtual block.

[0011] According to a sixth aspect of the present disclosure, there is provided a method of video encoding. The method may include determining one or more candidate motion vectors from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning distance, one of the at least one scanning distance indicating a number of blocks away from one side of the current block. Further, the method may include obtaining one or more CPMVs for the current block based on the one or more candidate motion vectors.

[0012] According to a seventh aspect of the present disclosure, there is provided a method of video decoding. The method may include obtaining one or more first parameters using an inheritance-based derivation method and obtaining one or more second parameters using a construction-based derivation method. Further, the method may include constructing one or more affine models using the one or more first parameters and the one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.

[0013] According to an eighth aspect of the present disclosure, there is provided a method of video encoding. The method may include determining one or more first parameters using an inheritance-based derivation method and determining one or more second parameters using a construction-based derivation method. Further, the method may include constructing one or more affine models using the one or more first parameters and the one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.

[0014] According to a ninth aspect of the present disclosure, there is provided an apparatus for video decoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors being further configured to perform a method according to the first aspect, the second aspect, the third aspect, or the seventh aspect when the instructions are executed.

[0015] According to a tenth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors being further configured to perform a method according to the fourth, fifth, sixth, or eighth aspect when the instructions are executed.

[0016] According to an eleventh aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform a method according to any one of the preceding aspects.

[0017] A more particular description of embodiments of the present disclosure will be made by reference to specific embodiments illustrated in the accompanying drawings, in which the embodiments will be described and explained with additional specificity and detail, given that the drawings represent only some embodiments and therefore should not be considered limiting in scope. [Brief explanation of the drawings]

[0018] [Figure 1A] 1 is a block diagram illustrating a system for encoding and decoding video blocks according to some embodiments of the present disclosure. [Figure 1B] FIG. 2 is a block diagram of an encoder according to some embodiments of the present disclosure. [Figure 1C] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1D] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1E]1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1F] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to some embodiments of the present disclosure. [Figure 3A] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3B] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3C] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3D] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3E] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 4A] FIG. 1 illustrates an example four-parameter affine model, according to some embodiments of the present disclosure. [Figure 4B] FIG. 1 illustrates an example four-parameter affine model, according to some embodiments of the present disclosure. [Figure 5] FIG. 1 illustrates a six-parameter affine model, according to some embodiments of the present disclosure. [Figure 6] FIG. 10 illustrates an example of close neighboring blocks for inherited affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 7] FIG. 10 illustrates an example of close neighboring blocks for constructed affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 8]10A-10C are diagrams illustrating non-adjacent neighboring blocks for inherited affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 9] FIG. 10 illustrates the derivation of constructed affine merge candidates using neighboring blocks, according to some embodiments of the present disclosure. [Figure 10] FIG. 10 illustrates vertical scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 11] FIG. 10 illustrates horizontal scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 12] FIG. 10 illustrates combined vertical and horizontal scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 13A] FIG. 10 illustrates a neighboring block having the same size as the current block, according to some embodiments of the present disclosure. [Figure 13B] 10A and 10B are diagrams illustrating neighboring blocks having a different size than the current block, according to some embodiments of the present disclosure. [Figure 14A] A figure illustrating an example in which the bottom-left or top-right block of the bottom-most or right-most block in the previous distance is used as the bottom-most or right-most block of the current distance, according to some embodiments of the present disclosure. [Figure 14B] 10A-10C are diagrams illustrating an example in which the left or top block of the bottom or rightmost block in the previous distance is used as the bottom or rightmost block of the current distance, according to some embodiments of the present disclosure. [Figure 15A] FIG. 10 illustrates scanning positions at bottom-left and top-right positions used for non-adjacent neighboring blocks above and to the left, according to some embodiments of the present disclosure. [Figure 15B] FIG. 10 illustrates a scanning position at the bottom right position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 15C]FIG. 10 illustrates a scanning position at the bottom left position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 15D] FIG. 10 illustrates a scanning position at the right-top position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 16] FIG. 10 illustrates a simplified scanning process for deriving pre-constructed merge candidates, according to some embodiments of the present disclosure. [Figure 17A] FIG. 10 illustrates an example of spatial neighborhoods for deriving inherited affine merge candidates, according to some embodiments of the present disclosure. [Figure 17B] FIG. 10 illustrates an example of spatial neighborhoods from which constructed affine merge candidates are derived, in accordance with some embodiments of the present disclosure. [Figure 18] 1 illustrates an example of an inheritance-based derivation method for deriving affine constructed candidates, according to some embodiments of the present disclosure. [Figure 19] FIG. 1 illustrates an exemplary computing environment coupled with a user interface, according to some embodiments of the present disclosure. [Figure 20] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 21] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 22] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 23] 1 is a flowchart illustrating a method for video encoding, according to some embodiments of the present disclosure. [Figure 24] 1 is a flowchart illustrating a method for video encoding, according to some embodiments of the present disclosure. [Figure 25] 1 is a flowchart illustrating a method for video encoding, according to some embodiments of the present disclosure. [Figure 26] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 27] 1 is a flowchart illustrating a method for video encoding, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0020] The terms used in the disclosure are employed only for the purpose of describing particular embodiments and are not intended to limit the disclosure. The singular forms "a / an," "said," and "the" in the disclosure and the appended claims are intended to include the plural forms as well, unless otherwise clearly indicated throughout the disclosure. Also, the term "and / or" used in the disclosure will be understood to refer to and include one or any or all possible combinations of the associated listed items.

[0021] References throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. A feature, structure, element, or characteristic described in connection with one or some embodiments may also be applicable to other embodiments, unless expressly specified otherwise.

[0022] Throughout the disclosure, the terms "first," "second," "third," etc. all do not imply any spatial or chronological order and, unless clearly specified otherwise, are used solely as nomenclature to refer to related elements, e.g., devices, components, compositions, steps, etc. For example, a "first device" and a "second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be arbitrarily named.

[0023] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" may include memory (shared, dedicated, or a group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. The components may or may not be physically attached or located near each other.

[0024] As used herein, the terms "if" or "when" may be understood to mean "upon" or "in response to," depending on the context. When these terms appear in the claims, they may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include steps where: i) when or if condition X exists, a function or action X' is performed; and ii) when or if condition Y exists, a function or action Y' is performed. A method may be implemented with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at different times for multiple executions of the method.

[0025] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a purely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a particular function.

[0026] 1A is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel, according to some implementations of the present disclosure. As shown in FIG. 1A, system 10 includes a source device 12 that generates and encodes video data to be subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, or video streaming devices. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0027] In some implementations, destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to destination device 14. In one embodiment, link 16 may include a communication medium to enable source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from source device 12 to destination device 14.

[0028] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, Blu-ray disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. Destination device 14 may access the coded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing coded video data stored on a file server. Transmission of the coded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0029] As shown in FIG. 1A , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feeding interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As one example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, implementations described herein may be generally applicable to video coding and may be applied to wireless and / or wired applications.

[0030] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access, decoding, and / or playback by destination device 14 or another device. Output interface 22 may further include a modem and / or a transmitter.

[0031] Destination device 14 includes an input interface 28, a video decoder 30, and a display device. Input interface 28 may include a receiver and / or a modem and may receive encoded video data over link 16. The encoded video data communicated over link 16 and provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included in encoded video data transmitted over a communications medium, stored on a storage medium, or stored on a file server.

[0032] In some implementations, destination device 14 may include a display device 34, which can be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0033] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to any particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or later standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or later standards.

[0034] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, an electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in the respective device.

[0035] Like HEVC, VVC is built on a block-based hybrid video coding framework. FIG. 1B is a block diagram illustrating a block-based video encoder according to some implementations of the present disclosure. In encoder 100, an input video signal is processed block by block, called a coding unit (CU). Encoder 100 may be the video encoder 20 shown in FIG. 1A. In VTM-1.0, a CU can be up to 128 x 128 pixels. However, unlike HEVC, which partitions blocks only based on a quadtree, in VVC, a single coding tree unit (CTU) is divided into CUs based on a quadtree / binary tree / ternary tree to adapt to variable local characteristics. In addition, the concept of multiple partitioning unit types in HEVC has been removed; i.e., the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In a multi-type tree structure, a CTU is first partitioned by a quadtree structure, and then each quadtree leaf node can be further partitioned by a binary tree and a ternary tree structure.

[0036] 3A-3E are schematic diagrams illustrating multi-type tree partitioning modes according to some implementations of the present disclosure, each showing five partitioning types, including quadrant partitioning (FIG. 3A), vertical bisection partitioning (FIG. 3B), horizontal bisection partitioning (FIG. 3C), vertically extended trisection partitioning (FIG. 3D), and horizontally extended trisection partitioning (FIG. 3E).

[0037] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra-prediction") uses pixels from samples of previously coded neighboring blocks (called reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in the video signal. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") uses reconstructed pixels from previously coded video pictures to predict the current video block. Temporal prediction reduces temporal redundancy inherent in the video signal. Temporal prediction for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. If multiple reference pictures are supported, a reference picture index is also transmitted, which is used to identify which reference picture in the reference picture store the temporal prediction comes from.

[0038] After spatial and / or temporal prediction, an intra / inter mode decision circuit 121 in encoder 100 chooses the best prediction mode based on, for example, a rate-distortion optimization method. Block predictor 120 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by inverse quantization circuit 116 and inverse transformed by inverse transform circuit 118 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Additionally, in-loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed in a reference picture store in picture buffer 117 and used to code subsequent video blocks. To form the output video bitstream 114, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 106 to be further compressed and packed to form the bitstream.

[0039] For example, a deblocking filter is available in the current version of VVC, along with AVC and HEVC. In HEVC, an additional in-loop filter called SAO is defined to further improve coding efficiency. In the current version of the VVC standard, an additional in-loop filter called ALF is being actively investigated and has a good chance of being included in the final standard.

[0040] These in-loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. They can also be turned off as a decision made by the encoder 100 to save computational complexity.

[0041] It should be noted that intra prediction is typically based on unfiltered reconstructed pixels, and inter prediction is based on filtered reconstructed pixels if those filter options are turned on by the encoder 100.

[0042] FIG. 2 is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section present in encoder 100 of FIG. 1B. The block-based video decoder 200 may be the video decoder 30 shown in FIG. 1A. In the decoder 200, an incoming video bitstream 201 is first decoded through entropy decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 204 and inverse transform 206 to obtain reconstructed prediction residuals. A block predictor mechanism, implemented in an intra / inter mode selector 212, is configured to perform either intra prediction 208 or motion compensation 210 based on the decoded prediction information. The set of unfiltered reconstructed pixels is obtained by summing, using a summer 214, the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block predictor mechanism.

[0043] The reconstructed blocks may further pass through an in-loop filter 209 before being stored in a picture buffer 213, which serves as a reference picture store. The reconstructed video in the picture buffer 213 may be transmitted to drive a display device and may also be used to predict later video blocks. In situations where the in-loop filter 209 is turned on, a filtering operation is performed on the reconstructed pixels to derive the final reconstructed video output 222.

[0044] In the current VVC and AVS3 standards, motion information for the current coding block is either replicated from spatial or temporal neighboring blocks specified by merge candidate indexes or obtained through explicit signaling of motion estimation. The focus of this disclosure is to improve the accuracy of motion vectors for affine merge modes by improving the derivation method of affine merge candidates. To facilitate the explanation of this disclosure, the existing affine merge mode design in the VVC standard is used as an example to promote the proposed ideas. While the existing affine mode design in the VVC standard is used as an example throughout this disclosure, those skilled in the art of modern video coding technology will note that the proposed techniques may also be applied to different designs of affine motion prediction modes or other coding tools with the same or similar design spirit.

[0045] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame may be monochrome and therefore include only one two-dimensional array of luma samples.

[0046] As shown in FIG. 1C, video encoder 20 (or, more specifically, a partitioning unit in a prediction processing unit of video encoder 20) generates a coded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively in a raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set. As a result, all CTUs in a video sequence have the same size, which may be one of 128x128, 64x64, 32x32, and 16x16. However, it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 1D, each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements used to code the coding tree block samples. The syntax elements describe how a video sequence may be reconstructed in video decoder 30, including the nature of different types of units of coded blocks of pixels and inter or intra prediction, intra prediction mode, motion vectors, and other parameters. For monochrome pictures or pictures with three distinct color planes, a CTU may contain syntax elements used to code a single coding tree block and samples of the coding tree block. A coding tree block may be an N by N block of samples.

[0047] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of a CTU to divide the CTU into smaller CUs. As shown in FIG. 1E, 64×64 CTU 400 is first partitioned into four smaller CUs, each with a 32×32 block size. Among the four smaller CUs, CU 410 and CU 420 are each partitioned into four 16×16 CUs by block size. Two 16×16 CUs, 430 and 440, are each further partitioned into four 8×8 CUs by block size. FIG. 1F shows a quad tree data structure illustrating the final result of the partitioning process for CTU 400 as shown in FIG. 1E, where each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32×32 to 8×8. Similar to the CTU shown in Figure 1D, each CU may contain a CB of luma samples and two corresponding coding blocks of chroma samples for identically sized frames, as well as syntax elements used to code the coding block samples. In monochrome pictures or pictures with three distinct color planes, a CU may contain a single coding block and syntax elements used to code the coding block samples. It should be noted that the quadtree partitioning shown in Figures 1E-1F is for illustrative purposes only; a CTU may be divided into CUs based on quadtree, ternary tree, or binary tree partitioning to suit variable local characteristics. In a multi-type tree structure, a CTU is partitioned using a quadtree structure, and each quadtree leaf CU may be further partitioned into binary and ternary tree structures. As shown in Figures 3A-3E, there are five possible partitioning types of coding blocks with width W and height H: quadtree partitioning, horizontal binary tree partitioning, vertical binary tree partitioning, horizontal ternary tree partitioning, and vertical ternary tree partitioning.

[0048] In some implementations, video encoder 20 may further partition a coding block of a CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction, inter or intra, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In monochrome pictures or pictures with three separate color planes, a PU may include a single PB and syntax elements used to predict the PB. Video encoder 20 may generate predictive luma, Cb and Cr blocks for the luma, and Cb and Cr PBs for each PU of the CU.

[0049] Video encoder 20 may use intra prediction or inter prediction to generate predictive blocks for a PU. If video encoder 20 uses intra prediction to generate predictive blocks for a PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate predictive blocks for a PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0050] After video encoder 20 generates predictive luma, Cb, and Cr blocks for one or more PUs of a CU, video encoder 20 may generate a luma residual block for the CU by subtracting the predictive luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0051] Further, as illustrated in FIG. 1E , video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some embodiments, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In monochrome pictures or pictures with three separate color planes, a TU may include a single transform block and syntax structures used to transform the transform block samples.

[0052] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0053] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients and provide further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy encode syntax elements that indicate the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements that indicate the quantized transform coefficients. Finally, video encoder 20 may output a bitstream that includes a sequence of bits that form a representation of the coded frame and associated data, which is either stored in storage device 32 or transmitted to destination device 14.

[0054] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks for PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, video decoder 30 may reconstruct the frame.

[0055] As mentioned above, video coding achieves video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding efficiency more than intra-frame prediction due to the use of motion vectors to predict a current video block from a reference video block.

[0056] However, ever-improving video data capture techniques and more refined video block sizes for preserving content within the video data also substantially increase the amount of data required to represent motion vectors for the current frame. One way to overcome this challenge is to take advantage of the fact that a group of neighboring CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between those neighboring CUs. Therefore, by utilizing their spatial and temporal correlations, also referred to as the current CU's "motion vector predictor (MVP)," it is possible to use the motion information of spatially neighboring CUs and / or temporally co-located CUs as an approximation of the current CU's motion information (e.g., motion vector).

[0057] Instead of encoding into the video bitstream the actual motion vector of the current CU determined by the motion estimation unit as described above in connection with FIG. 1B, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to yield a motion vector difference (MVD) for the current CU. By doing so, the motion vector determined by the motion estimation unit for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0058] Similar to the process of choosing a predictive block in a reference frame during inter-frame prediction of a code block, a set of rules is required that are adopted by both video encoder 20 and video decoder 30 to construct a motion vector candidate list (also known as a "merge list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from video encoder 20 to video decoder 30; the index of the selected motion vector predictor in the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU.

[0059] Affine Model In HEVC, only the translational motion model is applied for motion-compensated prediction. In the real world, many types of motion exist, such as zoom in / out, rotation, depth motion, and other irregular motion. In VVC and AVS3, affine motion-compensated prediction is applied by signaling one flag for each inter-coding block to indicate whether the translational motion model or the affine motion model is applied for inter prediction. In the current VVC and AVS3 designs, two affine modes are supported for one affine coding block, including a 4-parameter affine mode and a 6-parameter affine mode.

[0060] The four-parameter affine model has the following parameters: two parameters for translation in the horizontal and vertical directions, one parameter for zoom motion, and one parameter for rotation motion in both directions. In this model, the horizontal zoom parameter is equal to the vertical zoom parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. To achieve a better fit of the motion vectors and affine parameters, the affine parameters are derived from two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. As shown in Figures 4A-4B, the affine motion field of a block is described by two CPMVs (V0, V1). Based on the control point motion, the motion field of one affine coded block (v x ,v y )teeth,

number

[0061] The 6-parameter affine model has the following parameters: two parameters for translation in each of the horizontal and vertical directions, two parameters for zoom and rotation in the horizontal direction, and two parameters for zoom and rotation in the vertical direction. The 6-parameter affine motion model is coded by three CPMVs. As shown in Figure 5, the three control points of a 6-parameter affine block are located at the top-left, top-right, and bottom-left corners of the block. The motion at the top-left control point is associated with translation, the motion at the top-right control point is associated with rotation and zoom in the horizontal direction, and the motion at the bottom-left control point is associated with rotation and zoom in the vertical direction. Compared to the 4-parameter affine motion model, the rotation and zoom motions in the horizontal direction of the 6-parameter affine model cannot be identical to those in the vertical direction. Assuming that (V0, V1, V2) are the MVs at the top left, top right, and bottom left corners of the current block in Figure 5, the motion vectors (v x ,v y )teeth,

number

[0062] Affine Merge Mode In affine merge mode, the CPMV for the current block is not explicitly signaled but is derived from neighboring blocks. In particular, in this mode, motion information of spatially neighboring blocks is used to generate the CPMV for the current block. The affine merge mode candidate list has a limited size. For example, in the current VVC design, there can be a maximum of five candidates. The encoder can evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. Affine merge candidates can be determined in three ways. In the first method, affine merge candidates can be inherited from neighboring affine-coded blocks. In the second method, affine merge candidates can be constructed from translational MVs from neighboring blocks. In the third method, zero MVs are used as affine merge candidates.

[0063] For the inherited method, there can be at most two candidates, which are taken from the neighboring block located to the bottom left of the current block (e.g., the scanning order is from A0 to A1 as shown in Figure 6) and the neighboring block located to the top right of the current block (e.g., the scanning order is from B0 to B2 as shown in Figure 6), if available.

[0064] For the pre-constructed method, the candidates are combinations of nearby translational MVs, which can be generated by two steps.

[0065] Step 1: Obtain four translational MVs, including MV1, MV2, MV3, and MV4, from the available neighborhood. MV1: MV from one of the three neighboring blocks closest to the top left corner of the current block. As shown in Figure 7, the scanning order is B2, B3, and A2. MV2: The two neighboring blocks closest to the top right corner of the current block home MV from one. As shown in Figure 7, the scanning order is B1 and B0. MV3: The two neighboring blocks near the bottom left corner of the current block home MV from one. As shown in Figure 7, the scanning order is A1 and A0. MV4: MV from the block that is co-located in time with the neighboring block near the bottom right corner of the current block. As shown in the figure, the neighboring block is T.

[0066] Step 2: Derive combinations based on the four translational MVs from Step 1. Combination 1: MV1, MV2, MV3, Combination 2: MV1, MV2, MV4, Combination 3: MV1, MV3, MV4, Combination 4: MV2, MV3, MV4, Combination 5: MV1, MV2, Combination 6: MV1, MV3.

[0067] When the merge candidate list is not full after filling with inherited and constructed candidates, a zero MV is inserted at the end of the list.

[0068] Affine AMVP mode Affine Advanced Motion Vector Prediction (AMVP) mode can be applied to CUs with both width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predictor CPMVP is signaled in the bitstream. The affine AMVP candidate list size is 2, and the affine AMVP candidate list is signaled by using the following four types of CPMV candidates in the following order: - Inherited affine AMVP candidates extrapolated from the CPMVs of neighboring CUs, - A pre-constructed affine AMVP candidate CPMVP derived using the translational MVs of nearby CUs, - Translational MVs from nearby CUs, - Zero MV.

[0069] The checking order for inherited affine AMVP candidates is the same as that for inherited affine merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference pictures as in the current block are considered. No pruning process is applied when inserting inherited affine motion predictors into the candidate list.

[0070] The pre-constructed AMVP candidate is derived from the same spatial neighborhood as in affine merge mode. The same checking order is used as in affine merge candidate construction. In addition, the reference picture indexes of neighboring blocks are also checked. The first block in the checking order that is inter-coded and has the same reference picture as the current CU is used. If the current CU is coded in 4-parameter affine mode and both mv0 and mv1 are available, mv0 and mv1 are added as one candidate to the affine AMVP candidate list. If the current CU is coded in 6-parameter affine mode and all three CPMVs are available, they are added as one candidate to the affine AMVP candidate list. Otherwise, the pre-constructed AMVP candidate is set as unavailable.

[0071] After valid inherited and constructed affine AMVP candidates are inserted, the affine AMVP candidates list If mv is still less than 2, mv0, mv1, and mv2 are added in order to predict all control point MVs of the current CU as translational MVs, when available. Finally, if it is still not full, zero MVs are used to fill the affine AMVP list.

[0072] History-based merge candidate derivation History-based MVP (HMVP) merge candidates are added to the merge list after spatial MVP and temporal motion vector prediction (TMVP). In this method, the motion information of a previously coded block is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0073] The HMVP table size S can be set to 6, indicating that up to five history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, and a redundancy check is first applied to find whether an identical HMVP exists in the table. If found, the identical HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the identical HMVP is inserted into the last entry of the table.

[0074] HMVP candidates are used in the merge candidate list construction process. The most recent HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.

[0075] To reduce the number of operations for redundancy check, the following simplifications are introduced: First, the last two entries in the table are checked for redundancy with respect to each of the A1 and B1 spatial candidates. Second, the process of building a merge candidate list from HMVP is completed when the total number of available merge candidates reaches a maximum of the allowed merge candidates minus one.

[0076] For the current video standards VVC and AVS, only adjacent neighboring blocks are used to derive affine merge candidates for the current block, as shown in Figures 6 and 7 for inherited and constructed candidates, respectively. To increase the diversity of merge candidates and further exploit spatial correlations, it is straightforward to extend the coverage of neighboring blocks from adjacent to non-adjacent areas.

[0077] In the current video standards VVC and AVS, each affine inherited candidate is derived from one neighboring block with affine motion information, while each affine constructed candidate is derived from two or three neighboring blocks with translational motion information. To further exploit spatial correlation, new candidate derivation methods that combine affine and translational motion can be investigated.

[0078] The proposed candidate derivation method for the affine merge mode can be extended to other coding modes, such as the affine AMVP mode and the regular merge mode.

[0079] In this disclosure, the candidate derivation process for the affine merge mode can be extended by using not only adjacent neighboring blocks but also non-adjacent neighboring blocks. The detailed methods can be summarized in the following aspects, including affine merge candidate pruning, a non-ajacent neighbor based derivation process for affine inherited merge candidates, a non-ajacent neighbor based derivation process for affine constructed merge candidates, an inheritance based derivation method for affine constructed merge candidates, an HMVP based derivation method for affine constructed merge candidates, and a candidate derivation method for affine AMVP mode and regular merge mode.

[0080] Affine Merge Candidate Pruning As the affine merge candidate list in typical video coding standards usually has a limited size, candidate pruning is a necessary process to remove redundant candidates. This pruning process is necessary for both inherited and constructed affine merge candidates. As explained in the introduction, the CPMV of the current block is not directly used for affine motion compensation. Instead, the CPMV needs to be transformed into translational MVs at the position of each sub-block within the current block. The transformation process is performed by adhering to a general affine model, as shown below.

number

[0081] For the 6-parameter affine model, three CPMVs are available, designated V0, V1, and V2. The six model parameters a, b, c, d, e, and f are then

number

[0082] For a four-parameter affine model, when the top-left corner CPMV and top-right corner CPMV, referred to as V0 and V1, are available, the six parameters a, b, c, d, e, and f are

number

[0083] For a four-parameter affine model, when the top-left corner CPMV and bottom-left corner CPMV, referred to as V0 and V2, are available, the six parameters a, b, c, d, e, and f are

number

[0084] In the above equations (4), (5), and (6), w and h represent the width and height of the current block, respectively.

[0085] When two merge candidate sets of CPMV are compared for redundancy check, it is proposed to check the similarity of six affine model parameters. Therefore, the candidate pruning process can be performed in two steps.

[0086] In step 1, given two candidate sets of CPMV, the corresponding affine model parameters for each candidate set are derived. More specifically, the two candidate sets of CPMV can be represented by two sets of affine model parameters, e.g., (a1, b1, c1, d1, e1, f1) and (a2, b2, c2, d2, e2, f2).

[0087] In step 2, a similarity check is performed between the two sets of affine model parameters based on one or more predefined thresholds. In one embodiment, when the absolute values ​​of (a1-a2), (b1-b2), (c1-c2), (d1-d2), (e1-e2), and (f1-f2) are all below a positive threshold, such as a value of 1, the two candidates are considered similar and one of them may be pruned / removed and not placed in the merge candidate list.

[0088] In some embodiments, the divide or right shift operation in step 1 may be removed to simplify the calculations in the CPMV pruning process.

[0089] In particular, the model parameters of c, d, e, and f can be calculated without dividing by the width w and height h of the current block. For example, taking the above equation (4) as an example, the approximate model parameters of c', d', e', and f' can be calculated as the following equation (7):

number

[0090] In the case where only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters that depend on the width or height of the current block. In this case, the model parameters can be transformed to take into account the influence of the width and height. For example, in the case of Equation (5), the approximate model parameters of c', d', e', and f' can be calculated based on the following Equation (8): In the case of Equation (6), the approximate model parameters of c', d', e', and f' can be calculated based on the following Equation (9):

number

[0091] In step 2 above, a threshold is required to evaluate the similarity between two candidate sets of CPMVs. There may be multiple ways to define the threshold. In one embodiment, the threshold may be defined for each comparable parameter. Table 1 is an example of this embodiment, showing thresholds defined for each comparable model parameter. In another embodiment, the threshold may be defined by considering the size of the current coding block. Table 2 is an example of this embodiment, showing thresholds defined by the size of the current coding block.

[0092] [Table 1]

[0093] [Table 2]

[0094] In another embodiment, the threshold is the current block's width The thresholds can be defined by considering the width or height of the current coding block. Tables 3 and 4 are examples of this embodiment. Table 3 shows the thresholds defined by the width of the current coding block, and Table 4 shows the thresholds defined by the height of the current coding block.

[0095] [Table 3]

[0096] [Table 4]

[0097] In another embodiment, the thresholds may be defined as a group of fixed values. In another embodiment, the thresholds may be defined by a combination of any of the above embodiments. In one example, the thresholds may be defined based on the different parameters as well as the current block's width and height. Table 5 shows the current coding block's Width and 1 is an example of this embodiment showing a threshold defined by height. In any of the above proposed embodiments, the comparable parameter may represent any parameter defined in any of equations (4) through (9), if necessary.

[0098] [Table 5]

[0099] The advantage of using the transformed affine model parameters for candidate redundancy check is that it results in a unified similarity check process for candidates with different affine model types. For example, one merge candidate may have a 6-parameter affine model with three CPMVs. use Another candidate may use a four-parameter affine model with two CPMVs, which considers the different influence of each CPMV in the merge candidate when deriving the target MV for each sub-block, and provides the similarity significance of the two affine merge candidates relative to the width and height of the current block.

[0100] Non-proximal Neighborhood-Based Resolution Process for Affine Inherited Merge Candidates For inherited merge candidates, the non-proximal neighborhood-based derivation process can be performed in three steps: Step 1 is for candidate scanning, Step 2 is for CPMV projection, and Step 3 is for candidate pruning.

[0101] In step 1, non-adjacent neighboring blocks are scanned and selected by the following method.

[0102] Scanning Area and Distance In some embodiments, non-adjacent neighboring blocks may be scanned from the area to the left and the area above the current coding block, and the scanning distance may be defined as the number of coding blocks from the scanning position to the left or top side of the current coding block.

[0103] As shown in Figure 8, multiple lines of non-adjacent neighboring blocks can be scanned either to the left or above the current coding block. The distances shown in Figure 8 represent the number of coding blocks from each candidate position to the left or top side of the current block. For example, an area with a "distance of 2" on the left side of the current block indicates that candidate neighboring blocks located within this area are two blocks away from the current block. Similar indications can be applied to other scanning areas with different distances.

[0104] In one or more embodiments, the non-adjacent neighboring blocks at each distance may have the same block size as the current coding block, as shown in FIG. 13A. As shown in FIG. 13A, non-adjacent neighboring block 1301 on the left and non-adjacent neighboring block 1302 on the top have the same size as the current block 1303. In some embodiments, the non-adjacent neighboring blocks at each distance may have a different block size than the current coding block, as shown in FIG. 13B. Neighboring block 1304 is a neighboring block of the current block 1303. As shown in FIG. 13B, non-adjacent neighboring block 1305 on the left and non-adjacent neighboring block 1306 on the top are adjacent to the current block 1307. different The neighboring blocks 1308 are neighboring blocks that are close to the current block 1307.

[0105] It should be noted that when the non-neighboring blocks at each distance have the same block size as the current coding block, the value of the block size is adaptively changed according to the partitioning granularity in each different area within the image. When the non-neighboring blocks at each distance have a block size different from that of the current coding block, the value of the block size can be predefined as a constant value, such as 4x4, 8x8, or 16x16. The 4x4 non-neighboring motion field shown in Figures 10 and 12 is an example of this case, and the motion field can be considered as a special case of, but not limited to, a sub-block.

[0106] Similarly, the non-adjacent coding blocks shown in Figure 11 may have different sizes as well. In one embodiment, the non-adjacent coding blocks may have a size as the current coding block that is adaptively changed. In another embodiment, the non-adjacent coding blocks may have a predefined size with a fixed value, such as 4x4, 8x8, or 16x16.

[0107] Based on the defined scanning distance, the current coding block The total size of the scanning area on either the left or top side may be determined by a configurable distance value. In one or more embodiments, the maximum scanning distance on the left and top sides may use the same value or different values. FIG. 13 shows an example in which the maximum distance on both the left and top sides share the same value of 2. The maximum scanning distance value(s) may be determined by the encoder side and signaled in the bitstream. Alternatively, the maximum scanning distance value(s) may be predefined as fixed value(s), such as a value of 2 or 4. When the maximum scanning distance is predefined as a value of 4, it indicates that the scanning process is completed when the candidate list is full, or that all non-adjacent neighboring blocks with a maximum distance of 4 have been scanned, whichever comes first.

[0108] In one or more embodiments, within each scanning area at a particular distance, the start and end neighborhood blocks may be position dependent.

[0109] In some embodiments, for a left scanning area, the starting neighboring block may be the adjacent lower-left block of the starting neighboring block of the neighboring scanning area with a shorter distance. For example, as shown in FIG. 8, the starting neighboring block of the "Distance 2" scanning area on the left side of the current block is the adjacent lower-left neighboring block of the starting neighboring block of the "Distance 1" scanning area. The ending neighboring block may be the adjacent left block of the ending neighboring block of the upper scanning area with a shorter distance. For example, as shown in FIG. 8, the ending neighboring block of the "Distance 2" scanning area on the left side of the current block is the adjacent left neighboring block of the ending neighboring block of the "Distance 1" scanning area above the current block.

[0110] Similarly, for the upper scanning area, the starting neighboring block may be the right-top block adjacent to the starting neighboring block of the neighboring scanning area with a shorter distance, and the ending neighboring block may be the left-top block adjacent to the ending neighboring block of the neighboring scanning area with a shorter distance.

[0111] Scanning Order When neighboring blocks are scanned in non-contiguous areas, certain order or / and rules may be followed to determine the selection of scanned neighboring blocks.

[0112] In some embodiments, the left area may be scanned first, followed by scanning the area above. As shown in Figure 8, three lines of non-adjacent area on the left side (e.g., from distance 1 to distance 3) may be scanned first, followed by scanning three lines of non-adjacent area above the current block.

[0113] In some embodiments, the left area and the top area may be scanned alternately. For example, as shown in Figure 8, the left scanning area with "Distance 1" is scanned first, followed by scanning the top area with "Distance 1".

[0114] For scanning areas located on the same side (e.g., left or top areas), the scanning order is from the area with the shortest distance to the area with the longest distance. This order can be flexibly combined with other embodiments of the scanning order. For example, the left and top areas can be scanned alternately, and the order for the areas on the same side is scheduled to be from the shortest distance to the longest distance.

[0115] Within each scanning area at a particular distance, a scanning order can be defined. In one embodiment, for the left scanning area, scanning can start from the bottom neighboring block to the top neighboring block. For the top scanning area, scanning can start from the right block to the left block.

[0116] Scanning completed For the inherited merge candidates, neighboring blocks coded in affine mode are defined as qualified candidates. In some embodiments, the scanning process may be performed iteratively. For example, scanning performed within a particular area at a particular distance may be stopped in an instance when the first X qualified candidates are identified, where X is a predefined positive value. For example, as shown in FIG. 8, scanning within the left scanning area with a distance of 1 may be stopped when the first one or more qualified candidates are identified. Then, the next iteration of the scanning process begins by targeting another scanning area, governed by a predefined scanning order / rule.

[0117] In some embodiments, the scanning process may be performed continuously, for example, scanning performed within a particular area at a particular distance may be stopped in instances when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or the maximum allowed number of candidates has been reached.

[0118] During the candidate scanning process, the non-adjacent neighboring blocks of each candidate are determined and scanned by adhering to the scanning method proposed above. For easier implementation, the non-adjacent neighboring blocks of each candidate can be indicated and identified by a specific scanning position. Once the specific scanning area and distance are determined by adhering to the method proposed above, the scanning position can be determined accordingly based on the following method.

[0119] In one method, the bottom left and top right positions are used for the top and left non-adjacent neighboring blocks, as shown in Figure 15A.

[0120] Alternatively, the bottom right position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15B.

[0121] Alternatively, the bottom left position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15C.

[0122] Alternatively, the right-top position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15D.

[0123] For easier illustration, in Figures 15A-15D, each non-adjacent neighboring block is assumed to have the same block size as the current block. Without loss of generality, this illustration can be easily extended to non-adjacent neighboring blocks having different block sizes.

[0124] Furthermore, step 2 can utilize the same process for CPMV projection as used in the current AVS and VVC standards, where it is assumed that the current block shares the same affine model with the selected neighboring blocks, and then the coordinates of two or three corner pixels (e.g., if the current block uses a 4-parameter model, two coordinates (the top-left pixel / sample location and the top-right pixel / sample location) are used, and if the current block uses a 6-parameter model, three coordinates (the top-left pixel / sample location, the top-right pixel / sample location, and the bottom-left pixel / sample location) are used) are plugged into equation (1) or (2), depending on whether the neighboring blocks are coded with a 4-parameter or 6-parameter affine model, to generate two or three CPMVs.

[0125] In step 3, any qualified candidate identified in step 1 and converted in step 2 may undergo a similarity check against all existing candidates already in the merge candidate list. The details of similarity check have already been explained in the "Affine Merge Candidate Pruning" section above. If the newly qualified candidate is found to be similar to any existing candidate in the candidate list, then this newly qualified candidate will be removed / pruned.

[0126] Non-proximal Neighborhood-Based Derivation Process for Affine Constructed Merge Candidates In the case of deriving inherited merge candidates, one neighboring block is identified at a time, and this single neighboring block needs to be coded in affine mode and may contain two or three CPMVs. In the case of deriving constructed merge candidates, two or three neighboring blocks are identified at a time, and each identified neighboring block does not need to be coded in affine mode, and only one translational MV is extracted from this block.

[0127] Figure 9 presents an example in which preconstructed affine merge candidates can be derived using non-neighboring neighboring blocks. In Figure 9, A, B, and C are the geographic positions of three non-neighboring neighboring blocks. A hypothetical coding block is formed using the position of A as the top-left corner, the position of B as the top-right corner, and the position of C as the bottom-left corner. Considering a hypothetical CU as an affine-coded block, the MVs at positions A', B', and C' can be derived by adhering to Equation (3), and the model parameters (a, b, c, d, e, f) can be calculated by the translational MVs at positions A, B, and C. Once derived, the MVs at positions A', B', and C' can be used as three CPMVs for the current block, and an existing process (such as the one used in the AVS and VVC standards) for generating preconstructed affine merge candidates can be used.

[0128] For the constructed merge candidates, the non-neighborhood-based derivation process can be performed in five steps. The non-neighborhood-based derivation process can be performed in a device such as an encoder or decoder in five steps: Step 1 is for candidate scanning; Step 2 is for affine model determination; Step 3 is for CPMV projection; Step 4 is for candidate generation; and Step 5 is for candidate pruning. In Step 1, non-neighborhood blocks can be scanned and selected by the following method:

[0129] Scanning Area and Distance In some embodiments, to maintain rectangular coding blocks, the scanning process is performed only for two non-neighboring neighboring blocks, and the third non-neighboring neighboring block may depend on the horizontal and vertical positions of the first and second non-neighboring neighboring blocks.

[0130] In some embodiments, as shown in Figure 9, the scanning process is performed only for positions B and C. The position of A can be uniquely determined by the horizontal position of C and the vertical position of B. In this case, the scanning area and distance can be defined according to a specific scanning direction.

[0131] In some embodiments, the scanning direction may be perpendicular to the side of the current block. One example is shown in FIG. 10, where the scanning area is defined as one line of contiguous motion fields to the left or above the current block. The scanning distance is defined as the number of motion fields from the scanning position to the side of the current block. It should be noted that the size of the motion fields may depend on the maximum granularity of the applicable video coding standard. In the example shown in FIG. 10, the size of the motion fields is assumed to be set to 4x4, consistent with the current VVC standard.

[0132] In some embodiments, the scanning direction may be parallel to the sides of the current block. One example is shown in Figure 11, where the scanning area is defined as one line of contiguous coding blocks to the left or above the current block.

[0133] In some embodiments, the scanning direction can be a combination of vertical and horizontal scanning to the side of the current block. One example is shown in FIG. 12. As shown in FIG. 12, the scanning direction can also be a combination of parallel and diagonal. Scanning at position B starts from left to right, then diagonally to the block to the left and above. Scanning at position B repeats as shown in FIG. 12. Similarly, scanning at position C starts from top to bottom, then diagonally to the block to the left and above. Scanning at position C repeats as shown in FIG. 12.

[0134] Scanning Order In some embodiments, the scanning order may be defined as from a position with a smaller distance to the current coding block to a position with a larger distance, which may be applied in the case of vertical scanning.

[0135] In some embodiments, the scanning order can be defined as a fixed pattern. This fixed pattern scanning order can be used for candidate positions with similar distances. One example is the case of horizontal scanning. In one example, the scanning order can be defined as a top-to-bottom direction for the left scanning area and a left-to-right direction for the top scanning area, similar to the example shown in FIG. 11.

[0136] For the case of a combined scanning method, the scanning order can be a combination of fixed pattern and distance dependent, similar to the example shown in FIG.

[0137] Scanning completed For pre-constructed merge candidates, the qualified candidates do not need to be affine coded, since only translational MVs are required.

[0138] Depending on the number of candidates required, the scanning process may be concluded when the first X qualified candidates have been identified, where X is a positive value.

[0139] As shown in Figure 9, three corners, named A, B, and C, are needed to form a virtual coding block. For easier implementation, the scanning process in step 1 can be performed only to identify non-adjacent neighboring blocks located at corners B and C, and the coordinate of A can be accurately determined by taking the horizontal coordinate of C and the vertical coordinate of B. In this way, the virtual coding block formed is constrained to be rectangular. In the case where either point B or C is unavailable, e.g., outside the boundary, or the motion information of the non-adjacent neighboring blocks corresponding to B or C is unavailable, the horizontal or vertical coordinate of C can be defined as the horizontal or vertical coordinate, respectively, of the top-left point of the current block.

[0140] In another embodiment, when corner B and / or corner C are initially determined from the scanning process in step 1, non-adjacent neighboring blocks located at corners B and / or C may be identified accordingly. Second, the position(s) of corners B and / or C may be reset to a pivot point within the corresponding non-adjacent neighboring block, such as the center of mass of each non-adjacent neighboring block. For example, the center of mass may be defined as the geometric center of each neighboring block.

[0141] For purposes of uniformity, the methods of defining scanning areas and distances, scanning order, and scanning completion proposed for deriving inherited merge candidates may be fully or partially reused for deriving constructed merge candidates. In one or more embodiments, the same methods defined for inherited merge candidate scanning, including but not limited to scanning areas and distances, scanning order, and scanning completion, may be fully reused for constructed merge candidate scanning.

[0142] In some embodiments, the same method defined for inherited merge candidate scanning can be partially reused for constructed merge candidate scanning. Figure 16 shows an example of this case. In Figure 16, the block size of each non-adjacent neighboring block is the same as the current block, which is defined similarly as inherited candidate scanning, but the overall process is a simplified version because scanning at each distance is limited to only one block.

[0143] 17A-17B show another example of this case, where both the non-adjacent inherited merging candidates and the non-adjacent constructed merging candidates are defined with the same block size as the current coding block, and the scanning order, scanning area, and scanning completion condition may be defined differently.

[0144] In FIG. 17A, the maximum distance for the non-close neighbor on the left is four coding blocks, and the maximum distance for the non-close neighbor on the top is five coding blocks. Also, at each distance, the scanning direction is bottom-top for the left side and right-left for the top side. In FIG. 17B, the maximum distance for non-close neighbors is four for both the left and top sides. Additionally, scanning at certain distances is unavailable because there is only one block at each distance. In FIG. 17A, if M qualified candidates are identified, the scanning operation within each distance can be completed. The value of M can be a predefined fixed value, such as a value of 1 or any other positive integer, or a signaled value determined by the encoder, or a configurable value at the encoder or decoder. In one embodiment, the value of M can be equal to the merge candidate list size.

[0145] 17A-17B, scanning operations at different distances may be completed when N qualified candidates are identified. The value of N may be a predefined fixed value, such as a value of 1 or any other positive integer, or a signaled value determined by the encoder, or a configurable value in the encoder or decoder. In one embodiment, the value of N may be the same as the merge candidate list size. In another embodiment, the value of N may be the same as the value of M.

[0146] In both Figures 17A and 17B, non-close spatial neighbors with closer distances to the current block may be prioritized, which indicates that a non-close spatial neighbor with distance i is scanned or checked before a neighbor with distance i+1, where i may be a non-negative integer representing a particular distance.

[0147] At a certain distance, at most two non-adjacent spatial neighbors are used, which means that if available, at most one neighbor from one side of the current block, e.g., from the left and from the top, is selected for inherited or constructed candidate derivation. As shown in Figure 17A, the checking order for the left and top neighbors is bottom-top and right-left, respectively. For Figure 17B, this rule can also be applied, the difference being that at any particular distance, there can be only one option per side of the current block.

[0148] For the constructed candidates, as shown in FIG. 17B, left The positions of the left and top non-close spatial neighbors are first determined independently. After that, the position of the left-top neighbor, which can surround the rectangular virtual block together with the left and top non-close neighbors, can be determined accordingly. Then, as shown in Figure 9, the motion information of the three non-close neighbors is used to form CPMVs at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which are finally projected to the current CU to generate corresponding constructed candidates.

[0149] In step 2, the translational MVs at the selected candidate positions after step 1 are evaluated and a suitable affine model can be determined. For easier illustration, and without loss of generality, Figure 9 is again used as an example.

[0150] Due to factors such as hardware constraints, implementation complexity, and different reference indexes, the scanning process may be completed before a sufficient number of candidates are identified, for example, motion information for the motion fields in one or more of the selected candidates after step 1 may not be available.

[0151] If motion information for all three candidates is available, the corresponding hypothetical coding block represents a 6-parameter affine model. If motion information for one of the three candidates is not available, the corresponding hypothetical coding block represents a 4-parameter affine model. If motion information for more than one of the three candidates is not available, the corresponding hypothetical coding block may be unusable to represent a valid affine model.

[0152] In some embodiments, if motion information is not available for the top left corner of the virtual coding block, e.g., corner A in FIG. 9, or if motion information is not available for both the top right corner, e.g., corner B in FIG. 9, and the bottom left corner, e.g., corner C in FIG. 9, the virtual block may be set as invalid and may be unusable to represent a valid model, and then steps 3 and 4 may be skipped during the current iteration.

[0153] In some embodiments, if either the top right corner, e.g., corner B in FIG. 9, or the bottom left corner, e.g., corner C in FIG. 9, is unavailable, but not both, the virtual block may represent a valid four-parameter affine model.

[0154] In step 3, if the hypothetical coding block is capable of representing a valid affine model, the same projection process used for the inherited merge candidates can be used.

[0155] In one or more embodiments, the same projection process used for the inherited merge candidate may be used, in this case the 4-parameter model represented by the hypothetical coding block from step 2 is projected onto a 4-parameter model for the current block, and the 6-parameter model represented by the hypothetical coding block from step 2 is projected onto a 6-parameter model for the current block.

[0156] In some embodiments, the affine model represented by the virtual coding block from step 2 is always projected onto a 4-parameter model or a 6-parameter model for the current block.

[0157] According to equations (5) and (6), there can be two types of four-parameter affine models: Type A is when the top-left corner CPMV and the top-right corner CPMV, designated as V0 and V1, are available; Type B is when the top-left corner CPMV and the bottom-left corner CPMV, designated as V0 and V2, are available.

[0158] In one or more embodiments, the type of the projected 4-parameter affine model is the same type of 4-parameter affine model represented by the virtual coding block, e.g., if the affine model represented by the virtual coding block from step 2 is a 4-parameter affine model of type A or B, then the projected affine model for the current block is also of type A or B, respectively.

[0159] In some embodiments, the 4-parameter affine model represented by the hypothetical coding block from step 2 is always projected onto a 4-parameter model of the same type for the current block. For example, a 4-parameter affine model of type A or B represented by a hypothetical coding block is always projected onto a 4-parameter affine model of type A.

[0160] In step 4, based on the projected CPMV after step 3, in one embodiment, the same candidate generation process used in the current VVC or AVS standard may be used. In another embodiment, the temporal motion vectors used in the candidate generation process used in the current VVC or AVS standard may not be used for the non-adjacent neighborhood block-based derivation method. When a temporal motion vector is not used, it indicates that the generated combination does not include any temporal motion vector.

[0161] In step 5, any newly generated candidate after step 4 may undergo a similarity check against all existing candidates already in the merge candidate list. The details of the similarity check have been previously explained in the "Affine Merge Candidate Pruning" section. If the newly generated candidate is found to be similar to any existing candidate in the candidate list, the newly generated candidate is removed or pruned.

[0162] Inheritance-Based Derivation Method for Affine Constructed Merge Candidates For each affine inherited candidate, all motion information is inherited from one selected spatial neighboring block coded in affine mode. The inherited information includes CPMV, reference index, prediction direction, affine model type, etc. On the other hand, for each affine constructed candidate, all motion information is constructed from two or three selected spatial or temporal neighboring blocks, and the selected neighboring blocks are not coded in affine mode, and only translational motion information is needed from the selected neighboring blocks.

[0163] In this section, a new candidate derivation method is disclosed that combines features of inherited and constructed candidates.

[0164] In some embodiments, the combination of inheritance and construction may be achieved by separating the affine model parameters into different groups, with one group of affine parameters inherited from one neighboring block and another group of affine parameters inherited from another neighboring block.

[0165] In one embodiment, the parameters of an affine model are constructed from two groups. As shown in Equation (3), an affine model may include six parameters, including a, b, c, d, e, and f. The translational parameters {a, b} may represent one group, and the non-translational parameters {c, d, e, f} may represent another group. With this grouping method, the two groups of parameters may be inherited independently from two different neighboring blocks in a first step, and then concatenated / constructed into a complete affine model in a second step. In this case, the group with non-translational parameters must be inherited from one affine-coded neighboring block, and the group with translational parameters may be from any inter-coded neighboring block that may or may not be coded in affine mode. It should be noted that the affine-coded neighboring blocks may be selected from the near affine neighboring blocks or the non-neighboring affine neighboring blocks based on the scanning method previously proposed for the affine-inherited candidate, such as the method shown in Figure 17A, which is the scanning method / rule including the scanning area and distance, scanning order, and scanning completion used in the section "Non-neighborhood-based Derivation Process for Affine-Inherited Merge Candidates," and the scanning method may be performed for both the near and non-neighboring neighboring blocks. Alternatively, the affine-coded neighboring blocks may not physically exist but may be virtually constructed from the regular inter-coded neighboring blocks, such as the method shown in Figure 17B, which is the scanning method / rule including the scanning area and distance, scanning order, and scanning completion used in the section "Non-neighborhood-based Derivation Process for Affine-Constructed Merge Candidates."

[0166] In some embodiments, the neighboring blocks associated with each group can be determined in different ways. In one method, the neighboring blocks for different groups of parameters can be all from non-proximate neighborhood areas, and the scanning method can be designed similar to the previously proposed method for non-proximate neighborhood-based derivation processing. In another method, the neighboring blocks for different groups of parameters can be all from proximate neighborhood areas, and the scanning method can be the same as the current VVC or AVS video standard. In another method, the neighboring blocks for different groups of parameters can be all from proximate neighborhood areas, and the scanning method can be the same as the current VVC or AVS video standard. Nearby It may be partially from an area, or partially from a nearby non-proximate area.

[0167] When several groups of affine parameters are combined to construct a new candidate, there may be several rules to be observed. The first is eligibility criteria. In one embodiment, it may be checked whether the related neighboring block or blocks for each group use the same reference picture for at least one direction or both directions. In another embodiment, it may be checked whether the related neighboring block or blocks for each group use the same precision / resolution for motion vectors.

[0168] The second is a construction formula. In one embodiment, the CPMV of a new candidate can be derived in the following formula:

number

[0169] In another embodiment, the CPMV of the new candidate may be derived in the following manner:

number

[0170] FIG. 18 shows an example of an inheritance-based derivation method for deriving an affine constructed candidate. In FIG. 18, there are three steps for deriving an affine constructed candidate. In step 1, according to a specific grouping strategy, the encoder or decoder may perform scanning of nearby and non-neighboring neighboring blocks for each group. In the case of FIG. 18, two groups are defined: Neighbor 1 is coded in affine mode and provides non-translational affine parameters, and Neighbor 2 provides translational affine parameters. Neighbor 1 may be obtained according to the process in the "Non-neighborhood-based Derivation Process for Affine Inherited Merge Candidates" section shown in FIGS. 15A-15D and 17A, and Neighbor 1 may be a nearby or non-neighboring neighbor of the current block. Furthermore, Neighbor 2 may be obtained according to the process shown in FIGS. 16 and 17B.

[0171] In step 2, the parameters and positions determined in step 1 may define a specific affine model that can derive different CPMVs according to the coordinates (x, y) of the CPMVs. For example, as shown in FIG. 18, the non-translational parameters {c, d, e, f} may be obtained based on the neighborhood 1 obtained in step 1, and the translational parameters {a, b} may be obtained based on the neighborhood 2 obtained in step 1. Furthermore, the distance parameters Δw and Δh may be obtained based on the position (x1, y1) of the current block and the position (x2, y2) of neighborhood 2. The distance parameters Δw and Δh may indicate the horizontal and vertical distances between the current block and neighborhood 1 or neighborhood 2, respectively. For example, the distance parameters Δw and Δh may indicate the horizontal distance (x1-x2) between the current block and neighborhood 2, and the vertical distance (y1-y2) between the current block and neighborhood 2, respectively. In particular, Δw=x1-x2 and Δh=y1-y2.

[0172] In step 3, two or three CPMVs are derived for the current coding block, and the current coding block can be constructed to form a new affine candidate.

[0173] In some embodiments, other prediction information may be further constructed. If neighboring blocks are checked to have the same direction and / or reference picture, the prediction direction (e.g., bi-predicted or uni-predicted) and reference picture index may be the same as that of the associated neighboring block. Alternatively, the prediction information is determined by reusing the least overlapping information among the associated neighboring blocks from different groups. For example, if only one reference index in one direction from one neighboring block is the same as the reference index in the same direction of another neighboring block, the prediction direction of the new candidate is determined as uni-predictive, and the same reference index and direction are reused.

[0174] HMVP-Based Deriving Method for Affine Constructed Merge Candidates In the case of the near-neighborhood-based derivation process, already defined in the current video standards VVC and AVS and described in the above section and in FIG. 7, a fixed order of scanning for near neighborhoods is performed to identify two or three near neighbor blocks. In the case of the non-neighborhood-based derivation process, as proposed in the previous section and in FIG. 17B, two non-neighbor neighborhoods are identified during another fixed order of scanning. In other words, for both the near-neighborhood-based derivation method and the non-neighborhood-based derivation method, a certain depth of local scanning is necessary to identify the number of neighbors. This scanning process relies on local buffering around each current block and also incurs a certain amount of computational complexity.

[0175] On the other hand, as explained in the introduction section, the HMVP merge mode is already adopted in current VVC and AVS, and the translational motion information from neighboring blocks is already stored in the history table. In this case, the scanning process can be replaced by searching the HMVP table.

[0176] Therefore, for the previously proposed non-adjacent neighborhood-based derivation process and inheritance-based derivation process, instead of the scanning method shown in Figures 17B and 18, translational motion information can be obtained from the HMVP table. However, to subsequently derive affine constructed candidates, position information, width, height, and reference information are also required, which can be accessible if the current HMVP table can be modified. Therefore, it is proposed to extend the HMVP table to store additional information in addition to the motion information of each historical neighborhood. In one embodiment, the additional information can include the position of affine or non-affine neighboring blocks, or affine motion information such as CPMV or equivalent normal motion derived from CPMV (e.g., this normal motion can be from an internal sub-block of the affine-coded neighboring block), a reference index, etc.

[0177] Candidate Derivation Methods for Affine AMVP and Regular Merge Modes As explained in the above section, for affine AMVP mode, an affine candidate list is also required to derive the CPMV predictor. As a result, all the above proposed derivation methods can be applied to the affine AMVP mode as well. The only difference is that when the above proposed derivation methods are applied in AMVP, the selected neighboring block should have the same reference picture index as the current coding block.

[0178] For the regular merge mode, the candidate list is also constructed only with translational candidate MVs, not CPMVs. In this case, all the derivation methods proposed above can still be applied by adding an additional derivation step. This additional derivation step is to derive a translational MV for the current block, which can be achieved by selecting a specific rotation position (x,y) within the current block and observing the same equation (3). In other words, to derive the CPMV of an affine block, the positions of the three corners of the block can be used as the rotation position (x,y) in equation (3), and to derive the translational MV of a regular inter-coded block, the center position of the block can be used as the rotation position (x,y) in equation (3). Once the translational MV is derived for the current block, it can be inserted into the candidate list as another candidate.

[0179] Reordering the affine merge candidate list In one embodiment, non-neighboring spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. subblock-based temporal motion vector prediction (SbTMVP) candidates, if available; 2. inherited from close neighbors; 3. inherited from non-neighbors; 4. constructed from close neighbors; 5. constructed from non-neighbors; 6. zero MV.

[0180] In another embodiment, non-neighboring spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from close neighbors, 3. constructed from close neighbors, 4. inherited from non-neighbors, 5. constructed from non-neighbors, 6. zero MV.

[0181] In another embodiment, non-neighboring spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from nearby neighbors, 3. constructed from nearby neighbors, 4. one set of zero MVs, 5. inherited from non-neighbors, 6. constructed from non-neighbors, 7. remaining zero MVs if the list is still not full.

[0182] In another embodiment, non-close spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available; 2. inherited from close neighbors, inherited from neighbors not closer than X; 4. constructed from close neighbors; 5. constructed from neighbors not closer than Y; 6. inherited from neighbors not closer than X; 7. constructed from neighbors not closer than Y; 8. zero MV. In this embodiment, the values ​​X and Y may be predefined fixed values, such as a value of 2, or signaled values ​​determined by the encoder, or configurable values ​​in the encoder or decoder. In one embodiment, the value of X may be the same as the value of Y. In another embodiment, X The value of Y may differ from the value of

[0183] 19 illustrates a computing environment (or computing device) 1910 coupled to a user interface 1960. The computing environment 1910 can be part of a data processing server. In some embodiments, the computing device 1910 can perform any of the various methods or processes (encoding / decoding methods or processes) described below according to various embodiments of the present disclosure. The computing environment 1910 can include a processor 1920, a memory 1940, and an I / O interface 1950.

[0184] The processor 1920 typically controls the overall operation of the computing environment 1910, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1920 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Additionally, the processor 1920 may include one or more modules that facilitate interaction between the processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.

[0185] Memory 1940 is configured to store various types of data to support the operation of computing environment 1910. Memory 1940 may include predetermined software 1942. Examples of such data include instructions for any applications or methods running on computing environment 1910, video data sets, image data, etc. Memory 1940 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.

[0186] The I / O interface 1950 provides an interface between the processor 1920 and peripheral interface modules, such as a keyboard, click wheel, and buttons. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1950 may be coupled to an encoder and a decoder.

[0187] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as contained in memory 1940, executable by processor 1920 in computing environment 1910 for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0188] A non-transitory computer-readable storage medium has stored thereon a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the above-described method for motion prediction.

[0189] In some embodiments, the computing environment 1910 may be implemented with one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0190] FIG. 20 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0191] In step 2001, the processor 1920 may obtain one or more first parameters based on a first neighboring block of the current block.

[0192] In some embodiments, the one or more first parameters may include non-translational parameters associated with an affine model. For example, as shown in Figure 18, the one or more first parameters may include non-translational parameters c, d, e, and f inherited from affine-coded first neighboring blocks.

[0193] In some embodiments, the first neighboring block may be obtained from a plurality of adjacent neighboring blocks and a plurality of non-adjacent neighboring blocks, i.e., the first neighboring block may be an adjacent neighboring block or a non-adjacent neighboring block, where the adjacent neighboring blocks are adjacent to the current block and the non-adjacent neighboring blocks are each located a number of blocks away from one side of the current block.

[0194] In some embodiments, the first neighboring block may be obtained from a plurality of inter-coded neighboring blocks of the current block, and the plurality of inter-coded neighboring blocks may include an affine-coded block.

[0195] In step 2002, the processor 1920 may obtain one or more second parameters based on a first neighboring block and / or a second neighboring block of the current block.

[0196] In particular, the processor 1920 may obtain one or more second parameters based on the first neighboring block, the second neighboring block, or the first neighboring block and the second neighboring block.

[0197] In some embodiments, the one or more second parameters may include translation parameters associated with an affine model. For example, as shown in Figure 18, the one or more second parameters may include translation parameters a and b constructed based on second neighboring blocks.

[0198] In some embodiments, the second neighboring block may be obtained from a plurality of inter-coded neighboring blocks of the current block, and the plurality of inter-coded neighboring blocks may include affine-coded blocks and non-affine-coded blocks.

[0199] In some embodiments, the first neighboring block may be obtained from multiple non-adjacent neighboring blocks based on a first scanning rule, where each of the multiple non-adjacent neighboring blocks is located a number of blocks away from one side of the current block. For example, the first scanning rule may be the scanning rule including the scanning area and distance, scanning order, and scanning completion used in the "Non-adjacent Neighborhood-Based Derivation Process for Affine-Inherited Merge Candidates" section, and the scanning rule may be performed on both the adjacent neighboring blocks or the non-adjacent neighboring blocks, as shown in FIGS. 8, 13A-13B, 14A-14B, 15A-15D, and 17A.

[0200] In some embodiments, the second neighboring block may be obtained from multiple non-adjacent neighboring blocks based on a second scanning rule, which may be completely or partially identical to the first scanning rule. For example, the second scanning rule may be the scanning rule including the scanning area and distance, scanning order, and scanning completion used in the "Non-adjacent Neighborhood-Based Derivation Process for Affine Constructed Merge Candidates" section, and the scanning rule may be performed on both the adjacent neighboring blocks or the non-adjacent neighboring blocks, as shown in FIGS. 9-12, 16, and 17B.

[0201] In step 2003, the processor 1920 may construct one or more affine models by using the one or more first parameters and the one or more second parameters.

[0202] In some embodiments, one or more first parameters and one or more second parameters may be combined or concatenated to construct one or more affine models.

[0203] In step 2004 , the processor 1920 may obtain one or more CPMVs for the current block based on the one or more affine models constructed in step 2003 .

[0204] In some embodiments, the processor 1920 may determine that a first neighboring block and a second neighboring block are valid for constructing one or more affine models under certain preconditions. In one embodiment, the processor 1920 may determine that a first neighboring block and a second neighboring block are valid for constructing an affine model in response to determining that the first neighboring block and the second neighboring block use the same reference picture for at least one motion direction. Furthermore, in response to determining that the first neighboring block and the second neighboring block use the same reference picture for one motion direction, the processor 1920 may determine that the prediction direction of motion vector candidates formed based on one or more CPMVs is uni-predictive and that the same reference picture is used for the motion vector candidates for one motion direction. The processor 1920 may also determine that the prediction direction and reference picture of the current block are the same as the prediction direction and reference picture of the first and second neighboring blocks, respectively, in response to determining that the first neighboring block and the second neighboring block use the same reference picture for both motion directions. Here, one or more CPMVs for the current block obtained in step 2004 may be constructed to form motion vector candidates, which are not limited to affine candidates but may also include regular merge candidates, AMVP candidates, etc.

[0205] In another embodiment, the processor 1920 may determine that the first neighboring block and the second neighboring block are valid for constructing an affine model in response to determining that the first neighboring block and the second neighboring block use the same resolution for the motion vectors.

[0206] In some embodiments, the processor 1920 may construct one or more affine models based on one or more first parameters, one or more second parameters, a first position of the current block, and a second neighboring block or a second position of the first neighboring block. For example, as shown in step 2 in FIG. 18, the affine models may be constructed based on the non-translational parameters c, d, e, and f, the translational parameters a and b, and a difference between the current block and the second neighboring block. For example, the difference may include a corresponding coordinate difference as shown in FIG. 18. The positions of the current block, the first neighboring block, and the second neighboring block may be determined in different ways.

[0207] In some embodiments, the first position of the current block may be determined according to the top left corner of the current block, and the second position of the first or second neighboring block may be determined according to the top left corner of the first or second neighboring block.

[0208] In some embodiments, the one or more first parameters may include a plurality of parameters associated with an affine model, and the one or more second parameters may include a plurality of distance parameters. For example, as shown in FIG. 18, the one or more first parameters may include affine model parameters a, b, c, d, e, and f, and the one or more second parameters may include distance parameters Δw and Δh.

[0209] In some embodiments, distance parameters may be predefined as fixed values, for example, the values ​​of (Δw, Δh) may be predefined as fixed values ​​such as (0, 0) or any constant value.

[0210] In some embodiments, the distance parameters may each indicate a distance between the current block and a first neighboring block or a second neighboring block, for example, the distance parameters may include a first distance parameter Δw indicating a horizontal distance between the current block and the first or second neighboring block, and a second distance parameter Δh indicating a vertical distance between the current block and the first or second neighboring block.

[0211] FIG. 21 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0212] In step 2101, the processor 1920 may obtain a plurality of motion vector candidates from the HMVP table, and the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate.

[0213] In some embodiments, the multiple motion vector candidates are not limited to affine candidates, but may include regular merge candidates, AMVP candidates, and the like.

[0214] In some embodiments, the HMVP table may be extended by storing additional information in the HMVP table in addition to the motion information of each historical neighboring block. The additional information may include at least one, or more, of the following information: the position of each historical neighboring block, the affine motion information of each historical neighboring block, or the reference index of each historical neighboring block.

[0215] In step 2102, the processor 1920 may obtain a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, as shown in FIG.

[0216] In step 2103, the processor 1920 may obtain multiple CPMVs for the current block based on the multiple CPMVs of the virtual block.

[0217] In some embodiments, the processor 1920 may determine a third motion vector constructed candidate based on the first and second motion vector constructed candidates and the virtual block, may obtain multiple CPMVs for the virtual block based on the translational MVs of the first, second, and third motion vector constructed candidates, and may obtain multiple CPMVs for the current block based on the multiple CPMVs of the virtual block by using the same projection process used for the inherited candidate derivation.

[0218] FIG. 22 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.

[0219] In step 2201, the processor 1920 may obtain one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning distance, one of the at least one scanning distance indicating the number of blocks away from one side of the current block.

[0220] In step 2202, the processor 1920 may obtain one or more CPMVs for the current block based on one or more candidate motion vectors.

[0221] In some embodiments, One or more Motion vector candidates are not limited to affine candidates, but may include regular merge candidates, AMVP candidates, and the like.

[0222] In some embodiments, the processor 1920 may add one or more motion vector candidates to an affine candidate list for affine AMVP mode in response to determining that one or more motion vector candidates have the same reference picture index as the current block.

[0223] In some embodiments, the processor 1920 may obtain at least one translational motion vector for the current block based on one or more CPMVs and may add the at least one translational motion vector to a regular merge candidate list for a regular merge mode.

[0224] In some embodiments, the processor 1920 may obtain at least one translational motion vector for the current block based on one or more CPMVs by selecting a particular rotational position within the current block.

[0225] FIG. 23 is a flow chart illustrating a method for video encoding corresponding to the method as illustrated in FIG.

[0226] In step 2301, the processor 1920 may determine one or more first parameters based on a first neighboring block of the current block.

[0227] In step 2302, the processor 1920 may determine one or more second parameters based on a first neighboring block and / or a second neighboring block of the current block.

[0228] In particular, the processor 1920 may determine one or more second parameters based on the first neighboring block, the second neighboring block, or the first neighboring block and the second neighboring block.

[0229] In step 2303, the processor 1920 may construct one or more affine models by using the one or more first parameters and the one or more second parameters.

[0230] In some embodiments, one or more first parameters and one or more second parameters may be combined or concatenated to construct one or more affine models.

[0231] In step 2304 , the processor 1920 may obtain one or more CPMVs for the current block based on the one or more affine models constructed in step 2303 .

[0232] FIG. 24 is a flow chart illustrating a method for video encoding corresponding to the method as illustrated in FIG.

[0233] In step 2401, the processor 1920 may determine a plurality of motion vector candidates from the HMVP table, and the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate.

[0234] In step 2402, the processor 1920 may obtain a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, as shown in FIG.

[0235] In step 2403, the processor 1920 may obtain multiple CPMVs for the current block based on the multiple CPMVs of the virtual block.

[0236] FIG. 25 is a flow chart illustrating a method for video encoding corresponding to the method as illustrated in FIG.

[0237] In step 2501, the processor 1920 may determine one or more candidate motion vectors from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning distance, one of the at least one scanning distance indicating the number of blocks away from one side of the current block.

[0238] In step 2502, the processor 1920 may obtain one or more CPMVs for the current block based on one or more candidate motion vectors.

[0239] FIG. 26 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0240] In step 2601, an inheritance-based derivation method may be used to obtain one or more first parameters.

[0241] In some embodiments, the processor 1920 may use an inheritance-based derivation method to obtain a first neighboring block from a plurality of inter-coded neighboring blocks of the current block, and may obtain one or more first parameters based on the first neighboring block, where the plurality of inter-coded neighboring blocks may include an affine-coded block.

[0242] In some embodiments, the inheritance-based derivation method may be the derivation process for affine-inherited merge candidates described in the "Non-proximal Neighborhood-Based Derivation Process for Affine-Inherited Merge Candidates" section. In the inheritance-based derivation method, neighboring blocks of the current block may be scanned using the scanning methods / rules, including the scanning area and distance, scanning order, and scanning completion, used in the "Non-proximal Neighborhood-Based Derivation Process for Affine-Inherited Merge Candidates" section, and the scanning rules may be performed for both proximal and non-proximal neighboring blocks, as shown in FIGS. 8, 13A-13B, 14A-14B, 15A-15D, and 17A.

[0243] In some embodiments, the one or more first parameters may include a plurality of parameters associated with an affine model, and the one or more second parameters may include a plurality of distance parameters, where the plurality of distance parameters may include a first distance parameter indicating a horizontal distance between the current block and a first neighboring block and a second distance parameter indicating a vertical distance between the current block and the first neighboring block. The plurality of parameters associated with the affine model may include parameters {a, b, c, d, e, f} associated with the affine model. The first distance parameter and the second distance parameter may be distance parameters Δw and Δh, respectively.

[0244] In step 2602, the processor 1920 may obtain one or more second parameters using a construction-based derivation method.

[0245] In some embodiments, the processor 1920 may use a construction-based derivation method to obtain a second neighboring block from a plurality of inter-coded neighboring blocks of the current block, and may obtain one or more second parameters based on the second neighboring block, where the plurality of inter-coded neighboring blocks may include an affine-coded block and a non-affine-coded block.

[0246] In some embodiments, the construction-based derivation method may be the derivation process for affine constructed merge candidates described in the "Non-proximal Neighborhood-Based Derivation Process for Affine Constructed Merge Candidates" section. In the construction-based derivation method, neighboring blocks of the current block may be scanned using the scanning methods / rules, including the scanning area and distance, scanning order, and scanning completion, used in the "Non-proximal Neighborhood-Based Derivation Process for Affine Constructed Merge Candidates" section, and the scanning rules may be performed for both proximal and non-proximal neighboring blocks, as shown in FIGS.

[0247] In some embodiments, the one or more first parameters may include a plurality of parameters associated with an affine model, and the one or more second parameters may include a plurality of distance parameters, where the plurality of distance parameters may include a first distance parameter indicating a horizontal distance between the current block and a first neighboring block and a second distance parameter indicating a vertical distance between the current block and the first neighboring block. The plurality of parameters associated with the affine model may include parameters {a, b, c, d, e, f} associated with the affine model. The first distance parameter and the second distance parameter may be distance parameters Δw and Δh, respectively.

[0248] In some embodiments, the one or more first parameters may include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters may include a plurality of translational parameters associated with an affine model.

[0249] In some embodiments, the one or more first parameters may include a plurality of parameters associated with an affine model, and the one or more second parameters may include a plurality of distance parameters.

[0250] In some embodiments, the distance parameters may be predefined as fixed values.

[0251] In step 2603, the processor 1920 may construct one or more affine models by using the one or more first parameters and the one or more second parameters.

[0252] In step 2604, the processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models.

[0253] FIG. 27 is a flow chart illustrating a method for video encoding corresponding to the method as illustrated in FIG.

[0254] In step 2701, the processor 1920 may determine one or more first parameters using an inheritance-based derivation method.

[0255] In step 2702, the processor 1920 may determine one or more second parameters using a construction-based derivation method.

[0256] In step 2703, the processor 1920 may construct one or more affine models using the one or more first parameters and the one or more second parameters.

[0257] In step 2704, the processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models.

[0258] In some embodiments, an apparatus for video coding is provided, the apparatus including a processor 1920 and a memory 1940 configured to store instructions executable by the processor, the instructions, when executed, configured to perform any of the methods illustrated in Figures 20-27.

[0259] In some other embodiments, a non-transitory computer-readable storage medium is provided having instructions stored thereon that, when executed by a processor 1920, cause the processor to perform any of the methods illustrated in Figures 20-25.

[0260] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the disclosure that adhere to its general principles, including such departures from the disclosure, as come within known or customary practice in the art. It is intended that the specification and examples be considered as illustrative only.

[0261] It will be appreciated that the present disclosure is not limited to the exact embodiments described above and illustrated in the accompanying drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. 1. A method of video decoding, comprising: Obtaining one or more first parameters based on a first neighbor block of a current block; obtaining one or more second parameters based on the first neighboring block and / or second neighboring block of the current block; determining that the first neighboring block and the second neighboring block are valid in response to determining that the first neighboring block and the second neighboring block use the same reference picture for at least one motion direction or use the same resolution for a motion vector; constructing one or more affine models by using the one or more first parameters and the one or more second parameters; obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models; A method comprising:

2. 2. The method of claim 1, further comprising obtaining the first neighboring block from a plurality of adjacent neighboring blocks and a plurality of non-adjacent neighboring blocks, wherein the plurality of adjacent neighboring blocks are adjacent to the current block and the plurality of non-adjacent neighboring blocks are each located a number of blocks away from one side of the current block.

3. 2. The method of claim 1, further comprising obtaining the second neighboring block from a plurality of inter-coded neighboring blocks of the current block, the plurality of inter-coded neighboring blocks including affine-coded blocks and non-affine-coded blocks.

4. 2. The method of claim 1, further comprising obtaining the first neighboring block from a plurality of inter-coded neighboring blocks of the current block, the plurality of inter-coded neighboring blocks comprising an affine-coded block.

5. Obtaining the first neighboring block from a plurality of non-adjacent neighboring blocks based on a first scanning rule, wherein each of the plurality of non-adjacent neighboring blocks is located a number of blocks away from one side of the current block, and the first scanning rule includes a first scanning area and distance, a first scanning order, and a first scanning completion for deriving affine inherited merge candidates; obtaining the second neighboring block from the plurality of non-adjacent neighboring blocks based on a second scanning rule, the second scanning rule being completely or partially identical to the first scanning rule, the second scanning rule including a second scanning area and distance, a second scanning order, and a second scanning completion for deriving an affine constructed merge candidate; The method of claim 1 further comprising:

6. 2. The method of claim 1 , wherein the one or more first parameters include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters include a plurality of translational parameters associated with the affine model.

7. 2. The method of claim 1, further comprising: in response to determining that the first neighboring block and the second neighboring block use the same reference picture for one motion direction, determining that a prediction direction of a motion vector candidate formed based on the one or more CPMVs is uni-prediction and that the same reference picture is used for the motion vector candidate for the one motion direction; or in response to determining that the first neighboring block and the second neighboring block use the same reference picture for both motion directions, determining that a prediction direction and a reference picture of the current block are the same as those of the first and second neighboring blocks, respectively.

8. 2. The method of claim 1, further comprising constructing the one or more affine models based on the one or more first parameters, the one or more second parameters, a first position of the current block, and a second position of the second neighboring block or the first neighboring block.

9. 9. The method of claim 8, wherein the first position comprises a top left corner of the current block and the second position comprises a top left corner of the first or second neighboring block.

10. The method of claim 1 , wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters.

11. the plurality of distance parameters include a first distance parameter indicating a horizontal distance between the current block and the second neighboring block, and a second distance parameter indicating a vertical distance between the current block and the second neighboring block; or 11. The method of claim 10, wherein the plurality of distance parameters includes a first distance parameter indicating a horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating a vertical distance between the current block and the first neighboring block.

12. 1. An apparatus for video decoding, comprising: one or more processors; a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors; The one or more processors, when the instructions are executed, are configured to perform the method of any one of claims 1 to 11. Device.

13. A method for storing a bitstream, comprising storing a bitstream decoded by a video decoding method according to any one of claims 1 to 11.

14. A method for receiving a bitstream, the method comprising receiving a bitstream decoded by a video decoding method according to any one of claims 1 to 11.

15. 12. A computer program stored on a non-transitory computer readable medium for execution by a computing device having one or more processors, the computer program, when executed by the one or more processors, causing the computing device to perform the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • History-based motion vector prediction for affine mode

    US20200099951A1

  • Method and apparatus for video coding

    WO2021206992A1