Candidate derivation for affine merge modes in video coding
By deriving motion vector candidates using affine models from neighboring blocks, the method enhances video coding efficiency and quality by improving the accuracy of motion prediction in video compression.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video coding standards face challenges in accurately deriving motion vector candidates for affine motion prediction modes, leading to inefficiencies in video compression and quality preservation.
The method involves obtaining parameters from neighboring blocks to construct affine models and derive control point motion vectors (CPMVs) using inheritance-based and construction-based derivation methods, enhancing the accuracy of motion vector candidates.
Improves the derivation of motion vector candidates, leading to more efficient video coding and better compression performance while maintaining video quality.
Smart Images

Figure 2026065069000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 277,148, filed on November 8, 2021, entitled "Candidate Derivation for Affine Merge Mode in Video Coding", which is incorporated herein by reference in its entirety for all purposes.
[0002] The present disclosure relates to video coding and compression, and more particularly, but not limited thereto, to methods and apparatuses for improving affine - merge candidate derivation for affine motion prediction modes in video encoding or decoding processes.
Background Art
[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some well-known video coding standards in recent years include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMediaVideo1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its preceding standard VP9. Audio Video Coding (AVS), referring to digital audio and digital video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup of China. Most existing video coding standards are built on well-known hybrid video coding frameworks, that is, using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in video images or sequences, and using transform coding to reduce the energy of prediction errors. A key goal of video coding technology is to compress video data into a format that uses lower bit rates while avoiding or minimizing degradation of video quality.
[0004] The first generation of AVS standards includes the Chinese national standards "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+). It can offer approximately 50% bitrate savings at the same perceived quality compared to the MPEG-2 standard. The AVS1 standard video part was published as a Chinese national standard in February 2006. The second generation of AVS standards includes the Chinese national standard series "Information Technology, Efficient Multimedia Coding" (known as AVS2), primarily targeted at the transmission of extra HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was published as a Chinese national standard. Meanwhile, the AVS2 standard video part was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for applications. The AVS3 standard is a next-generation video coding standard for UHD video applications, aiming to surpass the coding efficiency of the latest international standard, HEVC. In March 2019, the AVS3-P2 baseline was completed at the 68th AVS meeting, which offers approximately 30% bit rate savings compared to the HEVC standard. Currently, there is one reference software called the High Performance Model (HPM), which is maintained by the AVS Group to demonstrate the reference implementation of the AVS3 standard. [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] This disclosure provides an embodiment of the technique for improving the derivation of motion vector candidates for motion prediction modes in video coding or decoding processes. [Means for solving the problem]
[0006] A method for video decoding is provided according to a first aspect of the present disclosure. The method includes obtaining one or more first parameters based on a first neighbor block of the current block, and obtaining one or more second parameters based on a first neighbor block and / or a second neighbor block of the current block. The method may further include constructing one or more affine models using one or more first parameters and one or more second parameters, and obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models.
[0007] A second aspect of the present disclosure provides a method for video decoding. The method may include obtaining a plurality of motion vector candidates from a history-based motion vector prediction (HMVP) table, the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate. The method may further include obtaining a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, and obtaining a plurality of control point motion vectors (CPMVs) for the current block based on a plurality of CPMVs of the virtual block.
[0008] A third aspect of the present disclosure provides a video decoding method, which may include: obtaining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning distance, where at least one of the scanning distances may indicate a number of blocks away from one side of the current block; and obtaining one or more CPMVs for the current block based on the one or more motion vector candidates.
[0009] A fourth aspect of this disclosure provides a method for video coding. The method may include determining one or more first parameters based on a first neighboring block of the current block, and determining one or more second parameters based on a first neighboring block and / or a second neighboring block of the current block. Furthermore, the method may include constructing one or more affine models using one or more first parameters and one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.
[0010] A fifth aspect of this disclosure provides a method for video encoding. The method may include determining a plurality of motion vector candidates from a history-based motion vector prediction (HMVP) table, the plurality of motion vector candidates may include a first motion vector constructed candidate and a second motion vector constructed candidate. Furthermore, the method may include determining a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, and obtaining a plurality of CPMVs for the current block based on a plurality of CPMVs of the virtual block.
[0011] A sixth aspect of the present disclosure provides a method for video encoding. The method may include determining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning distance, where one of the at least one scanning distances indicates the number of blocks away from one side of the current block. Furthermore, the method may include obtaining one or more CPMVs for the current block based on the one or more motion vector candidates.
[0012] A seventh aspect of this disclosure provides a method for video decoding. The method may include obtaining one or more first parameters using an inheritance-based derivation method, and obtaining one or more second parameters using a construction-based derivation method. Furthermore, the method may include constructing one or more affine models using one or more first parameters and one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.
[0013] A method for video coding is provided according to an eighth aspect of the present disclosure. The method may include determining one or more first parameters using an inheritance-based derivation method, and determining one or more second parameters using a construction-based derivation method. Furthermore, the method may include constructing one or more affine models using one or more first parameters and one or more second parameters, and obtaining one or more CPMVs for the current block based on the one or more affine models.
[0014] A ninth aspect of this disclosure provides an apparatus for video decoding. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to execute a method according to the first, second, third, or seventh aspect when an instruction is executed.
[0015] A tenth aspect of this disclosure provides an apparatus for video encoding. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to execute a method according to the fourth, fifth, sixth, or eighth aspect when an instruction is executed.
[0016] According to an eleventh aspect of the present disclosure, a non-temporary computer-readable storage medium storing computer executable instructions is provided, and when the computer executable instructions are executed by one or more computer processors, the computer executable instructions cause one or more computer processors to perform a method according to any one of the above aspects.
[0017] Further specific descriptions of embodiments of this disclosure are made by reference to certain embodiments illustrated in the accompanying drawings. Assuming that these drawings represent only a few embodiments and are therefore not considered to limit the scope, embodiments are described and explained through the use of the accompanying drawings, along with additional specificities and details. [Brief explanation of the drawing]
[0018] [Figure 1A] This is a block diagram illustrating a system for encoding and decoding video blocks according to some embodiments of the present disclosure. [Figure 1B] This is a block diagram of an encoder according to some embodiments of the present disclosure. [Figure 1C] This block diagram illustrates how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1D] This block diagram illustrates how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1E]A block diagram illustrating how a frame is recursively partitioned into a plurality of video blocks of different sizes and shapes according to some embodiments of the present disclosure. [Figure 1F] A block diagram illustrating how a frame is recursively partitioned into a plurality of video blocks of different sizes and shapes according to some embodiments of the present disclosure. [Figure 2] A block diagram of a decoder according to some embodiments of the present disclosure. [Figure 3A] A diagram illustrating block partitioning within a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3B] A diagram illustrating block partitioning within a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3C] A diagram illustrating block partitioning within a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3D] A diagram illustrating block partitioning within a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3E] A diagram illustrating block partitioning within a multi-type tree structure according to some embodiments of the present disclosure. [Figure 4A] A diagram illustrating a 4-parameter affine model according to some embodiments of the present disclosure. [Figure 4B] A diagram illustrating a 4-parameter affine model according to some embodiments of the present disclosure. [Figure 5] A diagram illustrating a 6-parameter affine model according to some embodiments of the present disclosure. [Figure 6] A diagram illustrating neighboring blocks for an inherited affine merge candidate according to some embodiments of the present disclosure. [Figure 7] A diagram illustrating neighboring blocks for a constructed affine merge candidate according to some embodiments of the present disclosure. [Figure 8]This figure illustrates non-adjacent neighbor blocks to inherited affine merge candidates according to some embodiments of the present disclosure. [Figure 9] This figure illustrates the derivation of a constructed affine merge candidate using neighbor blocks, according to some embodiments of the present disclosure. [Figure 10] This figure illustrates vertical scanning of non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 11] This figure illustrates horizontal scanning of non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 12] This figure illustrates vertical and horizontal scanning of non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 13A] This figure illustrates a neighboring block having the same size as the current block, according to some embodiments of the present disclosure. [Figure 13B] This figure illustrates neighboring blocks having a different size from the current block, according to some embodiments of the present disclosure. [Figure 14A] This figure illustrates an embodiment of some embodiments of the present disclosure in which the bottom-left or top-right block of the bottom-right or top-left block in the previous distance is used as the bottom-right or top-right block in the current distance. [Figure 14B] This figure illustrates an embodiment of some embodiments of the present disclosure in which the leftmost or topmost block of the lowest or rightmost block in the previous distance is used as the lowest or rightmost block in the current distance. [Figure 15A] This figure illustrates scanning positions at the bottom-left and top-right positions used for non-adjacent neighboring blocks above and to the left, according to some embodiments of the present disclosure. [Figure 15B] This figure illustrates a scanning position in the lower-right position used for both upper and left non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 15C]This figure illustrates a scanning position in the lower-left position used for both upper and left non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 15D] This figure illustrates the scanning position at the top right position used for both upper and left non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 16] This figure illustrates a simplified scanning process for deriving pre-built merge candidates, according to some embodiments of the present disclosure. [Figure 17A] This figure illustrates the spatial neighborhoods from which inherited affine merge candidates are derived, according to some embodiments of the present disclosure. [Figure 17B] This figure illustrates the spatial neighborhoods from which constructed affine merge candidates are derived, according to some embodiments of the present disclosure. [Figure 18] This figure illustrates an example of an inheritance-based derivation method for deriving an affine-constructed candidate according to some embodiments of the present disclosure. [Figure 19] This figure illustrates a computing environment coupled with a user interface, according to some embodiments of the present disclosure. [Figure 20] This flowchart illustrates a method for video decoding according to some embodiments of the present disclosure. [Figure 21] This flowchart illustrates a method for video decoding according to some embodiments of the present disclosure. [Figure 22] This flowchart illustrates a method for video decoding according to some embodiments of the present disclosure. [Figure 23] This flowchart illustrates a method for video encoding according to some embodiments of the present disclosure. [Figure 24] This flowchart illustrates a method for video encoding according to some embodiments of the present disclosure. [Figure 25] This flowchart illustrates a method for video encoding according to some embodiments of the present disclosure. [Figure 26] This flowchart illustrates a method for video decoding according to some embodiments of the present disclosure. [Figure 27] This flowchart illustrates a method for video encoding according to some embodiments of the present disclosure. [Modes for carrying out the invention]
[0019] Here, specific implementations are referenced in detail, and their embodiments are illustrated in the accompanying drawings. The following detailed description provides numerous non-limiting specific details to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices having digital video capabilities.
[0020] The terms used in the disclosure are adopted solely for the purpose of describing specific embodiments and are not intended to limit the disclosure. The singular forms “a / an,” “said,” and “the” in the disclosure and the attached claims are intended to also include the plural form unless otherwise explicitly indicated throughout the disclosure. Furthermore, the terms “and / or” used in the disclosure are to be understood to refer to, and include, any or all possible combinations of the multiple related items described.
[0021] Throughout this specification, references to “one embodiment,” “an embodiment,” “an example,” “some embodiments,” “some examples,” or similar language mean that the particular features, structures, or characteristics described are included in at least one embodiment or example. Features, structures, elements, or characteristics described in relation to one or more embodiments are applicable to other embodiments unless otherwise expressly specified.
[0022] Throughout the disclosure, terms such as “first,” “second,” “third,” etc., are used solely as nomenclature for reference to relevant elements, such as devices, components, compositions, steps, etc., without any spatial or chronological implied meaning, unless otherwise explicitly specified. For example, “first device” and “second device” may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be arbitrarily named.
[0023] The terms “module,” “submodule,” “circuit,” “sub-circuit,” “circuitry,” “sub-circuitry,” “unit,” or “sub-unit” may include memory (shared, dedicated, or grouped) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits that have stored code or instructions, or that do not have stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected to each other. These components may or may not be physically attached to or located in close proximity to each other.
[0024] As used herein, the terms “if” or “when” may be understood, depending on the context, to mean “upon” or “in response to.” Where these terms appear in a claim, they may not indicate that any related limitation or feature is conditional or optional. For example, a method may include steps such that i) a function or action X' is performed when or if condition X exists, and ii) a function or action Y' is performed when or if condition Y exists. The method may be implemented by both the ability to perform the function or action X' and the ability to perform the function or action Y'. Thus, both functions X' and Y' may be performed for multiple executions of the method at different times.
[0025] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a purely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a particular function.
[0026] Figure 1A is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, relating to several implementations of the present disclosure. As shown in Figure 1A, the system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. The source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, or video streaming devices. In some implementations, the source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0027] In some implementations, the destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one embodiment, link 16 may include a communication medium to enable the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, switch, base station, or any other device that may be useful in facilitating communication from the source device 12 to the destination device 14.
[0028] In some other implementations, encoded video data can be transmitted from the output interface 22 to the storage device 32. The encoded video data in the storage device 32 can then be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, digital multipurpose disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In further embodiments, the storage device 32 may correspond to a file server or another intermediate storage device capable of holding the encoded video data generated by the source device 12. The destination device 14 can access the stored video data from the storage device 32 via streaming or download. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. An exemplary file server includes a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both appropriate for accessing the encoded video data stored on the file server. Transmission of the encoded video data from the storage device 32 may be streaming transmission, download transmission, or a combination of both.
[0029] As shown in Figure 1A, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as video supplementary devices, e.g., a video camera, a video archive containing previously captured video, a video feeding interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. In one embodiment, if the video source 18 is a video camera of a security surveillance system, the source device 12 and the destination device 14 may form a camera phone or video phone. However, the implementations described herein may be applicable to video coding in general and may be applicable to wireless and / or wired applications.
[0030] Captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data can also (or instead) be stored in the storage device 32 for decoding and / or playback for later access by the destination device 14 or other devices. The output interface 22 may further include a modem and / or transmitter.
[0031] The destination device 14 includes an input interface 28, a video decoder 30, and a display device. The input interface 28 may include a receiver and / or modem and may receive encoded video data through link 16. The encoded video data communicated to link 16 and provided on the storage device 32 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 when decoding the video data. Such syntax elements may be included in the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0032] In some implementations, the destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to the user and may include any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0033] The video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. The video encoder 20 of the source device 12 can be configured to encode video data in accordance with their current or later standards. Similarly, the video decoder 30 of the destination device 14 can also be configured to decode video data in accordance with their current or later standards.
[0034] The video encoder 20 and video decoder 30 can each be implemented as one or more suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or a combination thereof. When partially implemented in software, the electronic device may store instructions for the software in a suitable non-temporary computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed herein. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of an encoder / decoder (codec) combined with the respective device.
[0035] Like HEVC, VVC is built upon a block-based hybrid video coding framework. Figure 1B is a block diagram illustrating a block-based video encoder according to several implementations of the present disclosure. In encoder 100, the input video signal is processed block by block, called coding units (CUs). Encoder 100 may be a video encoder 20 as shown in Figure 1A. In VTM-1.0, a CU can be up to 128 × 128 pixels. However, unlike HEVC, which partitions blocks only on a quadtree basis, in VVC, a single coding tree unit (CTU) is divided into CUs to fit variable local characteristics based on a quadtree / binary / ternary tree. In addition, the concept of multiple partitioning unit types in HEVC is eliminated; i.e., the separation of CUs, predictive units (PUs), and transform units (TUs) no longer exists in VVC, and instead, each CU is always used as the basic unit for both predictive and transformive without further partitioning. In a multi-type tree structure, a single CTU is initially partitioned by a quadtree structure. Each quadtree leaf node can then be further partitioned by binary and ternary tree structures.
[0036] Figures 3A to 3E are schematic diagrams illustrating multi-type tree partitioning modes according to several implementations of the present disclosure. Figures 3A to 3E each show five partitioning types, including quadripartitioning (Figure 3A), vertical bipartitioning (Figure 3B), horizontal bipartitioning (Figure 3C), vertically extended tertiary partitioning (Figure 3D), and horizontally extended tertiary partitioning (Figure 3E).
[0037] For each given video block, spatial and / or temporal prediction may be performed. Spatial prediction (or "intra-prediction") uses pixels from samples of already coded neighboring blocks (called reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is transmitted, which is used to identify which reference picture in the reference picture store the temporal prediction signal is coming from.
[0038] After spatial and / or temporal prediction, the intra / inter-mode determination circuit 121 in encoder 100 selects the best prediction mode, for example, based on a rate-distortion optimization method. The block predictor 120 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using the transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are dequantized by the inverse quantization circuit 116 and inversely transformed by the inverse transform circuit 118 to form the reconstructed residual, and the reconstructed residual is then added back to the prediction block to form the reconstructed signal of the CU. Furthermore, in-loop filtering 115 such as a deblocking filter, sample-adaptive offset (SAO), and / or adaptive in-loop filter (ALF) may be applied to the reconstructed CU before it is placed in the reference picture store of picture buffer 117 and used to code subsequent video blocks. To form the output video bitstream 114, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all further compressed and packed and sent to the entropy coding unit 106 to form the bitstream.
[0039] For example, deblocking filters are available in the current version of VVC, along with AVC and HEVC. In HEVC, an additional in-loop filter called SAO is defined to further improve coding efficiency. In the current version of the VVC standard, yet another in-loop filter called ALF is being actively investigated and has a good chance of being included in the final standard.
[0040] These in-loop filter operations are optional. Performing them helps improve coding efficiency and visual quality. They can also be turned off as a decision made by encoder 100 to save computational complexity.
[0041] It should be noted that intra-predictions are typically based on unfiltered reconstructed pixels, and if those filtering options are turned on by encoder 100, intra-predictions are based on filtered reconstructed pixels.
[0042] Figure 2 is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section present in the encoder 100 in Figure 1B. The block-based video decoder 200 may be a video decoder 30 as shown in Figure 1A. In the decoder 200, the incoming video bitstream 201 is first decoded through entropy decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 204 and inverse transform 206 to obtain reconstructed prediction residuals. A block predictor mechanism, implemented in the intra / inter-mode selector 212, is configured to perform either intra-prediction 208 or motion compensation 210 based on the decoded prediction information. The unfiltered reconstructed pixel set is obtained by summing the reconstructed predicted residuals from the inverse transform 206 and the predicted outputs generated by the block predictor mechanism using an adder 214.
[0043] The reconstructed blocks may further pass through an in-loop filter 209 before being stored in a picture buffer 213, which functions as a reference picture store. The reconstructed video in the picture buffer 213 may be sent to drive a display device and may also be used to predict subsequent video blocks. When the in-loop filter 209 is turned on, filtering operations are performed on those reconstructed pixels to derive the final reconstructed video output 222.
[0044] In the current VVC and AVS3 standards, motion information for the current coding block is either replicated from spatially or temporally neighboring blocks specified by merge candidate indexes, or obtained through explicit signaling of motion estimation. The focus of this disclosure is to improve the accuracy of motion vectors for affine merge modes by improving the method of deriving affine merge candidates. To facilitate the explanation of this disclosure, existing affine merge mode designs in the VVC standard are used as examples to illustrate the proposed idea. While existing affine mode designs in the VVC standard are used as examples throughout this disclosure, it should be noted to those skilled in the art of modern video coding techniques that the proposed technique may also be applicable to different designs of affine motion prediction modes or other coding tools having the same or similar design spirit.
[0045] In a typical video coding process, a video sequence typically contains an ordered set of frames or pictures. Each frame may contain three sample arrays, represented as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame may be monochrome and therefore may contain only one two-dimensional array of luma samples.
[0046] As shown in Figure 1C, the video encoder 20 (or, more specifically, the partitioning unit in the predictive processing unit of the video encoder 20) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may contain integer CTUs that are sequentially ordered in the raster scan order from left to right and from top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set, so that all CTUs in a video sequence have the same size, which is one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not necessarily limited to a specific size. As shown in Figure 1D, each CTU may contain one CTB of a chroma sample, two corresponding coding tree blocks of a chroma sample, and syntax elements used to code the samples in the coding tree blocks. The syntax elements describe how the video sequence can be reconstructed in the video decoder 30, including the properties of different types of units of the coded blocks of pixels, inter- or intra-prediction, intra-prediction mode, motion vector, and other parameters. For monochrome pictures or pictures with three separate color planes, the CTU may include a single coding tree block and syntax elements used to code samples of the coding tree block. The coding tree block may be an N×N block of samples.
[0047] To achieve better performance, the video encoder 20 may recursively perform tree partitioning on the coding tree blocks of the CTU, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof, to divide the CTU into smaller CUs. As shown in Figure 1E, the 64x64 CTU400 is initially partitioned into four smaller CUs, each having a block size of 32x32. Among the four smaller CUs, CU410 and CU420 are each partitioned into four 16x16 CUs by their block size. The two 16x16 CUs, CU430 and CU440, are each further partitioned into four 8x8 CUs by their block size. Figure 1F shows a quadtree data structure illustrating the final result of the partitioning process of the CTU400 as shown in Figure 1E, where each leaf node of the quadtree corresponds to one CU of each size ranging from 32x32 to 8x8. Similar to the CTU shown in Figure 1D, each CU may contain a CB of a chroma sample, two corresponding coding blocks of chroma samples in frames of the same size, and syntax elements used to code the samples in the coding blocks. In monochrome pictures or pictures with three separate color planes, a CU may contain a single coding block and the syntax structure used to code the samples in the coding block. The quadtree partitioning shown in Figures 1E-1F is for illustrative purposes only, and it should be noted that a single CTU can be partitioned into CUs based on quadtree / ternary / binary tree partitioning to suit variable local characteristics. In a multi-type tree structure, a single CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned into binary and ternary tree structures. As shown in Figures 3A to 3E, there are five possible partitioning types for coding blocks having width W and height H: quadtree partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.
[0048] In some implementations, the video encoder 20 may further partition the coding block of the CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of a sample to which the same prediction, inter, or intra is applied. The PU of the CU may include a PB for a luma sample, two corresponding PBs for a chroma sample, and syntax elements used to predict the PBs. For a monochrome picture or a picture with three distinct color planes, the PU may include a single PB and syntax elements used to predict the PB. The video encoder 20 may generate predictive luma, Cb and Cr blocks for the luma, and Cb and Cr PBs for each PU of the CU.
[0049] The video encoder 20 may use intra-prediction or inter-prediction to generate predictive blocks for the PU. If the video encoder 20 uses intra-prediction to generate predictive blocks for the PU, the video encoder 20 may generate predictive blocks for the PU based on decoded samples of frames associated with the PU. If the video encoder 20 uses inter-prediction to generate predictive blocks for the PU, the video encoder 20 may generate predictive blocks for the PU based on decoded samples of one or more frames other than the frames associated with the PU.
[0050] After the video encoder 20 generates predictive luma, Cb, and Cr blocks for one or more PUs of the CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predictive luma block of the CU from its original luma coding block, so that each sample in the luma residual block of the CU represents the difference between a luma sample in one of the predictive luma blocks of the CU and the corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, so that each sample in the Cb residual block of the CU represents the difference between a Cb sample in one of the predictive Cb blocks of the CU and the corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU represents the difference between a Cr sample in one of the predictive Cr blocks of the CU and the corresponding sample in the original Cr coding block of the CU.
[0051] Furthermore, as illustrated in Figure 1E, the video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of the CU into one or more luma, Cb, and Cr transformation blocks, respectively. A transformation block is a rectangular (square or non-square) block of samples to which the same transformation is applied. The TU of the CU may include a transformation block for luma samples, two corresponding transformation blocks for chroma samples, and syntax elements used to transform the transformation block samples. Thus, each TU of the CU may be associated with a luma transformation block, a Cb transformation block, and a Cr transformation block. In some embodiments, a luma transformation block associated with a TU may be a sub-block of the luma residual block of the CU. A Cb transformation block may be a sub-block of the Cb residual block of the CU. A Cr transformation block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, the TU may include a single transformation block and a syntax structure used to transform the samples of the transformation block.
[0052] The video encoder 20 may apply one or more transformations to the Luma transformation block of TU in order to generate a Luma coefficient block for TU. The coefficient block may be a two-dimensional array of transformation coefficients. The transformation coefficients may be scalar quantities. The video encoder 20 may apply one or more transformations to the Cb transformation block of TU in order to generate a Cb coefficient block for TU. The video encoder 20 may apply one or more transformations to the Cr transformation block of TU in order to generate a Cr coefficient block for TU.
[0053] After generating a coefficient block (e.g., a Luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing the transformation coefficients to potentially reduce the amount of data used to represent the transformation coefficients and provide further compression. After the video encoder 20 has quantized the coefficient block, the video encoder 20 may entropi encode the syntax elements representing the quantized transformation coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements representing the quantized transformation coefficients. Finally, the video encoder 20 may output a bitstream containing a sequence of bits that form a representation of the coded frame and associated data, which is either stored in the storage device 32 or transmitted to the destination device 14.
[0054] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from it. The video decoder 30 may reconstruct frames of video data based at least partially on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally interactive with the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding a sample of the predictive block for the PU of the current CU to the corresponding sample of the transform block for the TU of the current CU. After reconstructing the coding block for each CU of the frame, the video decoder 30 may reconstruct the frame.
[0055] As mentioned above, video coding archives video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). It will be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding efficiency more than intra-frame prediction due to the use of motion vectors to predict the current video block from a reference video block.
[0056] However, with improvements in video data capture technology and more refined video block sizes for storing content within video data, the amount of data required to represent the motion vector for the current frame also increases substantially. One way to overcome this challenge is to take advantage of the fact that groups of neighboring CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but the motion vectors between those neighboring CUs are also similar. Thus, by utilizing their spatial and temporal correlations, also referred to as the "motion vector predictor (MVP)" of the current CU, it is possible to use the motion information of spatially neighboring CUs and / or CUs at the same temporal location as an approximation of the motion information (e.g., motion vector) of the current CU.
[0057] Instead of encoding the actual motion vector of the current CU, determined by the motion estimation unit as described above in relation to Figure 1B, into the video bitstream, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to produce a motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.
[0058] Similar to the process of selecting a predictive block in a reference frame during inter-frame prediction of a code block, a set of rules is required to construct a motion vector candidate list (also known as a "merge list") for the current CU using the potential candidate motion vectors associated with spatially neighboring CUs and / or CUs at the same temporal location as the current CU, and then to select one member from the motion vector candidate list as the motion vector predictor for the current CU, which is then employed by both the video encoder 20 and the video decoder 30. By doing so, there is no need to transmit the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor in the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU.
[0059] Affine Model In HEVC, only translational motion models are applied for motion compensation prediction. In the real world, many types of motion exist, such as zoom in / out, rotation, depth motion, and other irregular motions. In VVC and AVS3, affine motion compensation prediction is applied by signaling a single flag for each intercoding block to indicate whether a translational motion model or an affine motion model is applied for intercoding prediction. In the current VVC and AVS3 designs, two affine modes, including 4-parameter affine mode and 6-parameter affine mode, are supported for one affine coding block.
[0060] The 4-parameter affine model has the following parameters: two parameters for translational movement in the horizontal and vertical directions, one parameter for zoom motion, and one parameter for rotational motion in both directions. In this model, the horizontal zoom parameter is equal to the vertical zoom parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. To achieve a better fit of motion vectors and affine parameters, these affine parameters are derived from two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. As shown in Figures 4A-4B, the affine motion field of the block is described by two CPMVs (V0,V1). Based on the control point motion, the motion field (v) of one affine-coded block is described. x ,v y )teeth,
number
[0061] A 6-parameter affine mode has the following parameters: two parameters for translational movement in the horizontal and vertical directions, two parameters each for zoom and rotational movement in the horizontal direction, and two additional parameters each for zoom and rotational movement in the vertical direction. The 6-parameter affine motion model is coded by three CPMVs. As shown in Figure 5, the three control points of a 6-parameter affine block are located at the top left, top right, and bottom left corners of the block. The movement at the top left control point relates to translational movement, the movement at the top right control point relates to rotational and zoom movement in the horizontal direction, and the movement at the bottom left control point relates to rotational and zoom movement in the vertical direction. Compared to a 4-parameter affine motion model, the rotational and zoom movements in the horizontal direction of the 6-parameter model cannot be identical to their movements in the vertical direction. Assuming that (V0, V1, V2) are the MVs of the top left, top right, and bottom left corners of the current block in Figure 5, the motion vectors (v) of each sub-block are as follows: x ,v y )teeth,
number
[0062] Affine Merge Mode In affine merge mode, the CPMV for the current block is not explicitly signaled but is derived from neighboring blocks. In particular, in this mode, motion information of spatially neighboring blocks is used to generate the CPMV for the current block. The affine merge mode candidate list has a limited size. For example, in the current VVC design, there may be up to five candidates. The encoder can evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. The affine merge candidate can be determined in three ways. In the first way, the affine merge candidate can be inherited from neighboring affine-coded blocks. In the second way, the affine merge candidate can be constructed from translational MVs from neighboring blocks. In the third way, zero MV is used as the affine merge candidate.
[0063] There can be up to two candidates for inherited methods. If available, the candidates are taken from the neighboring block located to the lower left of the current block (for example, with a scanning order from A0 to A1, as shown in Figure 6) and from the neighboring block located to the upper right of the current block (for example, with a scanning order from B0 to B2, as shown in Figure 6).
[0064] Regarding the constructed method, the candidates are combinations of neighboring translational MVs that can be generated in two steps.
[0065] Step 1: Obtain four translational MVs from the available neighborhoods, including MV1, MV2, MV3, and MV4. MV1: The MV from one of the three neighboring blocks closest to the top left corner of the current block. The scanning order is B2, B3, and A2, as shown in Figure 7. MV2: The two neighboring blocks closest to the top right corner of the current block home MV from one. As shown in Figure 7, the scanning order is B1 and B0. MV3: The two neighboring blocks closest to the bottom left corner of the current block home MV from one. As shown in Figure 7, the scanning order is A1 and A0. MV4: The MV from a neighboring block near the lower right corner of the current block, which is at the same temporal location as the current block. As shown in the diagram, the neighboring block is T.
[0066] Step 2: Derive combinations based on the four translational MVs from Step 1. Combination 1: MV1, MV2, MV3, Combination 2: MV1, MV2, MV4, Combination 3: MV1, MV3, MV4, Combination 4: MV2, MV3, MV4, Combination 5: MV1, MV2, Combination 6: MV1, MV3.
[0067] If the merge candidate list is not full after being filled with inherited and constructed candidates, zero MVs are inserted at the end of the list.
[0068] Affine AVP Mode The Affine Advanced Motion Vector Prediction (AMVP) mode can be applied to CUs having both width and height greater than 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predictor CPMVP is signaled in the bitstream. The affine AMVP candidate list size is 2, and the affine AMVP candidate list is signaled by using the following four types of CPMV candidates in the following order: - Inherited affine AMVP candidates extrapolated from CPMV of neighboring CUs, - Constructed affine AMVP candidate CPMVP derived using translational MV of neighboring CUs, - Translational MV from nearby CU, - Zero MV.
[0069] The checking order for inherited affine AMVP candidates is the same as that for inherited affine merge candidates. The only difference is that for AMVP candidates, only affine CUs that have the same reference picture as the one in the current block are considered. Pruning is not applied when inserting inherited affine motion predictors into the candidate list.
[0070] The constructed AMVP candidate is derived from the same spatial neighborhood as the affine merge mode. The same checking order used for constructing the affine merge candidate is used. In addition, the reference picture indices of the neighborhood blocks are also checked. The first block in the checking order that has the same reference picture as the intercoded block in the current CU is used. If the current CU is coded using the 4-parameter affine mode and both mv0 and mv1 are available, mv0 and mv1 are added to the affine AMVP candidate list as a single candidate. If the current CU is coded using the 6-parameter affine mode and all three CPMVs are available, they are added to the affine AMVP candidate list as a single candidate. Otherwise, the constructed AMVP candidate is set as unavailable.
[0071] After valid inherited affine AMVP candidates and constructed AMVP candidates are inserted, the affine AMVP candidate list If it is still less than 2, then, when available, mv0, mv1, and mv2 are added in order to predict all control point MVs of the current CU as translational MVs. Finally, if it is still not full, zero MVs are used to fill the affine AMVP list.
[0072] History-based merge candidate derivation History-based MVP (HMVP) merge candidates are added to the merge list after spatial MVP and time-motion vector prediction (TMVP). In this method, motion information from previously coded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the coding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is an inter-coded CU for a non-subblock, the associated motion information is added to the last entry in the table as a new HMVP candidate.
[0073] The HMVP table size S can be set to 6, indicating that a maximum of 5 history-based MVP (HMVP) candidates can be added to the table. When inserting a new move candidate into the table, a constrained first-in, first-out (FIFO) rule is used, and a redundancy check is applied first to find out if a duplicate HMVP already exists in the table. If found, the duplicate HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the duplicate HMVP is inserted as the last entry in the table.
[0074] HMVP candidates are used in the merge candidate list construction process. The most recent HMVP candidates in the table are checked in order and inserted into the candidate list after TMVP candidates. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.
[0075] To reduce the number of calculations required for redundancy checks, the following simplifications are introduced: First, the last two entries in the table are checked for redundancy for each of the A1 and B1 spatial candidates. Second, the process of building the merge candidate list from HMVP is completed when the total number of available merge candidates reaches the maximum allowed merge candidate minus 1.
[0076] For current video standards VVC and AVS, only adjacent neighbor blocks are used to derive affine merge candidates for the current block, as shown in Figures 6 and 7 for inherited and constructed candidates, respectively. To increase the diversity of merge candidates and to further utilize spatial correlations, it is straightforward to extend neighbor block coverage from adjacent areas to non-adjacent areas.
[0077] In the current video standards VVC and AVS, each affine inherited candidate is derived from one neighborhood block using affine motion information. On the other hand, each affine constructed candidate is derived from two or three neighborhood blocks using translational motion information. To further utilize spatial correlations, new candidate derivation methods combining affine and translational motion may be investigated.
[0078] The proposed candidate derivation method for affine merge mode can be extended to other coding modes, such as affine AMVP mode and regular merge mode.
[0079] In this disclosure, the candidate derivation process for affine merge mode can be extended by using not only adjacent neighbor blocks but also non-adjacent neighbor blocks. The detailed methods can be summarized in the following embodiments, which include affine merge candidate pruning, non-adjacent neighbor-based derivation process for affine inherited merge candidates, non-adjacent neighbor-based derivation process for affine constructed merge candidates, inheritance-based derivation method for affine constructed merge candidates, HMVP-based derivation method for affine constructed merge candidates, and candidate derivation method for affine AMVP mode and regular merge mode.
[0080] Affine Marge candidate pruning Just as affine merge candidate lists in typical video coding standards usually have a limited size, candidate pruning is a necessary process to remove redundant ones. This pruning process is required for both affine merge inherited candidates and constructed candidates. As explained in the introduction, the CPMV of the current block is not used directly for affine motion compensation. Instead, the CPMV needs to be converted to a translational MV at the location of each subblock within the current block. The conversion process is performed by adhering to a general affine model, as shown below.
number
[0081] For a 6-parameter affine model, three CPMVs are available, denoted as V0, V1, and V2. Then, the six model parameters a, b, c, d, e, and f are:
number
[0082] For a 4-parameter affine model, if the leftmost top corner CPMV and the rightmost top corner CPMV, denoted as V0 and V1, are available, then the six parameters a, b, c, d, e, and f are:
number
[0083] For a 4-parameter affine model, if the top-left corner CPMV and bottom-left corner CPMV, denoted as V0 and V2, are available, then the six parameters a, b, c, d, e, and f are:
number
[0084] In equations (4), (5), and (6) above, w and h represent the width and height of the current block, respectively.
[0085] When two merge candidate sets of CPMV are compared for redundancy checking, it is proposed to check the similarity of the six affine model parameters. Therefore, the candidate pruning process can be performed in two steps.
[0086] Step 1 assumes two candidate sets of CPMV and derives the corresponding affine model parameters for each candidate set. More specifically, the two candidate sets of CPMV can be represented by two sets of affine model parameters, e.g., (a1,b1,c1,d1,e1,f1) and (a2,b2,c2,d2,e2,f2).
[0087] In step 2, a similarity check is performed between two sets of affine model parameters based on one or more predefined thresholds. In one embodiment, two candidates are considered similar when the absolute values of (a1-a2), (b1-b2), (c1-c2), (d1-d2), (e1-e2), and (f1-f2) are all below a positive threshold, such as a value of 1, and one of them can be pruned / removed and not placed in the merge candidate list.
[0088] In some embodiments, the splitting or right-shift operation in step 1 may be omitted to simplify calculations in the CPMV pruning process.
[0089] In particular, the model parameters c, d, e, and f can be calculated without being divided by the width w and height h of the current block. For example, taking equation (4) above as an example, the approximate model parameters c', d', e', and f' can be calculated as shown in equation (7) below.
number
[0090] In the case where only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters that depend on the width or height of the current block. In this case, the model parameters can be transformed to take into account the effects of width and height. For example, in the case of equation (5), the approximate model parameters c', d', e', and f' can be calculated based on equation (8) below. In the case of equation (6), the approximate model parameters c', d', e', and f' can be calculated based on equation (9) below.
number
[0091] In step 2 above, a threshold is needed to evaluate the similarity between two candidate sets of CPMVs. Several methods can exist for defining the threshold. In one embodiment, the threshold may be defined for each comparable parameter. Table 1 shows one example in this embodiment, illustrating the thresholds defined for each comparable model parameter. In another embodiment, the threshold may be defined by considering the size of the current coding block. Table 2 shows one example in this embodiment, illustrating the threshold defined by the size of the current coding block.
[0092] [Table 1]
[0093] [Table 2]
[0094] In another embodiment, the threshold is the current block width Alternatively, it may be defined by taking height into consideration. Tables 3 and 4 show examples in this embodiment. Table 3 shows thresholds defined by the width of the current coding block, and Table 4 shows thresholds defined by the height of the current coding block.
[0095] [Table 3]
[0096] [Table 4]
[0097] In another embodiment, the threshold may be defined as a group of fixed values. In another embodiment, the threshold may be defined by any combination of the embodiments described above. In one embodiment, the threshold is defined by different parameters as well as the current block. widthAnd it can be defined by taking height into consideration. Table 5 shows the current coding block width and This is one embodiment of this design that shows a threshold defined by height. In any of the embodiments proposed above, the comparable parameter may, if necessary, represent any parameter defined in any of the equations from equations (4) to (9).
[0098] [Table 5]
[0099] The advantage of using converted affine model parameters for candidate redundancy checking is that it results in a unified similarity checking process for candidates with different affine model types, for example, one merge candidate having a 6-parameter affine model with three CPMVs. use Another possible approach is to use a four-parameter affine model with two CPMVs, which takes into account the different effects of each CPMV in the merge candidate when deriving the target MV in each sub-block, and which provides similarity significance of the two affine merge candidates relative to the width and height of the current block.
[0100] Non-neighbor-based derivation for affine inherited merge candidates For inherited merge candidates, non-adjacent neighbor-based derivation can be performed in three steps: Step 1 is for candidate scanning; Step 2 is for CPMV projection; Step 3 is for candidate pruning.
[0101] In Step 1, non-adjacent neighboring blocks are scanned and selected in the following manner:
[0102] Scanning area and distance In some embodiments, non-adjacent neighboring blocks may be scanned from the area to the left and above the current coding block. The scanning distance may be defined as the number of coding blocks from the scanning position to the left or top of the current coding block.
[0103] As shown in Figure 8, multiple lines of non-adjacent neighboring blocks may be scanned above or to the left of the current coding block. The distances shown in Figure 8 represent the number of coding blocks from each candidate position to the left or top of the current block. For example, an area with a "distance of 2" above and to the left of the current block indicates that candidate neighboring blocks located within this area are 2 blocks away from the current block. Similar indications may be applied to other scanning areas with different distances.
[0104] In one or more embodiments, non-adjacent neighbor blocks at each distance may have the same block size as the current coding block, as shown in Figure 13A. As shown in Figure 13A, the upper left non-adjacent neighbor block 1301 and the upper upper non-adjacent neighbor block 1302 have the same size as the current block 1303. In some embodiments, non-adjacent neighbor blocks at each distance may have a different block size than the current coding block, as shown in Figure 13B. Neighbor block 1304 is an adjacent neighbor block to the current block 1303. As shown in Figure 13B, the upper left non-adjacent neighbor block 1305 and the upper upper non-adjacent neighbor block 1306 are the same size as the current block 1307. different It has a size. Neighboring block 1308 is a neighboring block that is adjacent to current block 1307.
[0105] It should be noted that when non-adjacent neighboring blocks at each distance have the same block size as the current coding block, the block size value will adaptively change according to the partitioning granularity in each different area of the image. When non-adjacent neighboring blocks at each distance have a different block size than the current coding block, the block size value can be predefined as a constant value, such as 4×4, 8×8, or 16×16. The 4×4 non-adjacent motion fields shown in Figures 10 and 12 are embodiments of this case, and motion fields can be considered as special cases of sub-blocks, although they are not limited to these.
[0106] Similarly, non-adjacent coding blocks shown in Figure 11 may also have different sizes. In one embodiment, a non-adjacent coding block may have a size as the current coding block, which is adaptively changed. In another embodiment, a non-adjacent coding block may have a predefined size with a fixed value, such as 4x4, 8x8, or 16x16.
[0107] Based on the defined scanning distance, current coding... block The total size of the scanning area on either the left or top of the bitstream can be determined by a configurable distance value. In one or more embodiments, the maximum scanning distances on the left and top may be the same or different values. Figure 13 shows an embodiment in which the maximum distances on both the left and top share the same value of 2. The maximum scanning distance value(s) may be determined by the encoder and signaled in the bitstream. Alternatively, the maximum scanning distance value(s) may be predefined as a fixed value(s), such as 2 or 4. When the maximum scanning distance is predefined as a value of 4, it indicates that the scanning process is completed when the candidate list is full, or that all non-adjacent neighboring blocks with a maximum distance of 4 have been scanned, whichever comes first.
[0108] In one or more embodiments, within each scanning area at a specific distance, the start and end neighborhood blocks may be position-dependent.
[0109] In some embodiments, with respect to a left-side scanning area, the start neighborhood block may be the adjacent lower-left neighborhood block of the start neighborhood block of a nearby scanning area having a shorter distance. For example, as shown in Figure 8, the start neighborhood block of the “distance 2” scanning area above left of the current block is the adjacent lower-left neighborhood block of the start neighborhood block of the “distance 1” scanning area. The end neighborhood block may be the adjacent left neighborhood block of the end neighborhood block of an upper scanning area having a shorter distance. For example, as shown in Figure 8, the end neighborhood block of the “distance 2” scanning area above left of the current block is the adjacent left neighborhood block of the end neighborhood block of the “distance 1” scanning area above the current block.
[0110] Similarly, with respect to the upper scanning area, the start neighborhood block may be the upper right block adjacent to the start neighborhood block of a nearby scanning area that has a shorter distance between it and the start neighborhood block. The end neighborhood block may be the upper left block adjacent to the end neighborhood block of a nearby scanning area that has a shorter distance between it and the end neighborhood block.
[0111] Scanning order When neighboring blocks are scanned in an area where they are not adjacent, a certain order and / or rule may be followed to determine the selection of scanned neighboring blocks.
[0112] In some embodiments, the left area may be scanned first, followed by scanning the upper area. As shown in Figure 8, three non-adjacent lines in the upper left area (e.g., from distance 1 to distance 3) may be scanned first, followed by scanning three non-adjacent lines in the upper current block.
[0113] In some embodiments, the left area and the upper area may be scanned alternately. For example, as shown in Figure 8, the left scanning area with distance 1 is scanned first, followed by the upper area with distance 1.
[0114] For scanning areas located on the same side (e.g., the left or top area), the scanning order is from the area with the shortest distance to the area with the longest distance. This order can be flexibly combined with other embodiments of the scanning order. For example, the left and top areas may be scanned alternately, and the order for areas on the same side may be scheduled from the shortest distance to the longest distance.
[0115] Within each scanning area at a given distance, a scanning order can be defined. In one embodiment, for the left scanning area, scanning may begin from the lower neighbor block to the uppermost neighbor block. For the upper scanning area, scanning may begin from the right block to the left block.
[0116] Scanning complete For inherited merge candidates, neighboring blocks coded in affine mode are defined as eligible candidates. In some embodiments, the scanning process can be performed iteratively. For example, scanning performed within a specific area at a certain distance may be stopped in the instance when the first X eligible candidates are identified, where X is a predefined positive value. For example, as shown in Figure 8, scanning within a left scanning area at a distance of 1 may be stopped when the first one or more eligible candidates are identified. The next iteration of the scanning process is then started by targeting another scanning area, which is governed by a predefined scanning order / rule.
[0117] In some embodiments, the scanning process may be performed continuously. For example, scanning performed within a specific area at a given distance may be stopped in instances when all covered neighboring blocks have been scanned, no more qualified candidates have been identified, or the maximum allowed number of candidates has been reached.
[0118] During the candidate scanning process, non-adjacent neighbor blocks of each candidate are determined and scanned by adhering to the scanning method proposed above. For easier implementation, non-adjacent neighbor blocks of each candidate may be indicated and identified by a specific scanning position. Once a specific scanning area and distance are determined by adhering to the method proposed above, the scanning position may be determined accordingly based on the following method.
[0119] In one method, the bottom-left and top-right positions are used for non-adjacent neighboring blocks above and to the left, as shown in Figure 15A.
[0120] Alternatively, the lower-right position is used for both the upper and left non-adjacent neighboring blocks, as shown in Figure 15B.
[0121] Alternatively, the lower left position is used for both the upper and left non-adjacent neighboring blocks, as shown in Figure 15C.
[0122] Alternatively, the top-right position is used for non-adjacent neighboring blocks both above and to the left, as shown in Figure 15D.
[0123] For a simpler illustration, in Figures 15A–15D, each non-adjacent neighboring block is assumed to have the same block size as the current block. Without loss of versatility, this illustration can be easily extended to non-adjacent neighboring blocks having different block sizes.
[0124] Furthermore, in step 2, the same processing of CPMV projection as used in the current AVS and VVC standards may be utilized. In this CPMV projection processing, the current block shares the same affine model with the selected neighboring block, and then the coordinates of two or three corner pixels (for example, two coordinates (top-left pixel / sample position and top-right pixel / sample position) are used if the current block uses a 4-parameter model, and three coordinates (top-left pixel / sample position, top-right pixel / sample position, and bottom-left pixel / sample position) are used if the current block uses a 6-parameter model) are plugged into equation (1) or (2), which is assumed to depend on whether the neighboring block is coded by a 4-parameter or 6-parameter affine model to generate two or three CPMVs.
[0125] In Step 3, any eligible candidate identified in Step 1 and converted in Step 2 may undergo a similarity check against all existing candidates already in the merge candidate list. Details of the similarity check have already been explained in the "Affine Merge Candidate Pruning" section above. If a newly eligible candidate is found to be similar to any existing candidate in the candidate list, this newly eligible candidate will be removed / prune.
[0126] Non-proximate neighbor-based derivation for affine-constructed merge candidates In the case of deriving inherited merge candidates, one neighborhood block is identified at a time, and this single neighborhood block must be coded in affine mode and may contain two or three CPMVs. In the case of deriving constructed merge candidates, two or three neighborhood blocks are identified at a time, and each identified neighborhood block does not need to be coded in affine mode, and only one translational MV is taken from this block.
[0127] Figure 9 presents an example in which a constructed affine merge candidate can be derived by using non-adjacent neighbor blocks. In Figure 9, A, B, and C are the geographical positions of three non-adjacent neighbor blocks. A virtual coding block is formed using position A as the top-left corner, position B as the top-right corner, and position C as the bottom-left corner. When considering a virtual CU as an affine-coded block, the MVs at positions A', B', and C' can be derived by adhering to equation (3), and the model parameters (a, b, c, d, e, f) can be calculated by the translational MVs at positions A, B, and C. Once derived, the MVs at positions A', B', and C' can be used as three CPMVs for the current block, and existing processes for generating constructed affine merge candidates (one used in the AVS and VVC standards) can be used.
[0128] For constructed merge candidates, non-adjacent neighbor-based derivation can be performed in five steps. Non-adjacent neighbor-based derivation can be performed in five steps in a device such as an encoder or decoder. Step 1 is for candidate scanning. Step 2 is for affine model determination. Step 3 is for CPMV projection. Step 4 is for candidate generation. Step 5 is for candidate pruning. In Step 1, non-adjacent neighbor blocks can be scanned and selected by the following method.
[0129] Scanning area and distance In some embodiments, in order to maintain a rectangular coding block, the scanning process is performed only on two non-adjacent neighbor blocks. A third non-adjacent neighbor block may depend on the horizontal and vertical positions of the first and second non-adjacent neighbor blocks.
[0130] In some embodiments, as shown in Figure 9, the scanning process is performed only on positions B and C. Position A may be uniquely determined by the horizontal position C and the vertical position B. In this case, the scanning area and distance may be defined according to a specific scanning direction.
[0131] In some embodiments, the scanning direction may be perpendicular to the side of the current block. One embodiment is shown in Figure 10, where the scanning area is defined as a single line of continuous motion fields to the left or above the current block. The scanning distance is defined as the number of motion fields from the scanning position to the side of the current block. It will be noted that the size of the motion fields may depend on the maximum granularity of the applicable video coding standard. In the embodiment shown in Figure 10, the size of the motion fields is assumed to be set to 4x4, consistent with the current VVC standard.
[0132] In some embodiments, the scanning direction may be parallel to the side of the current block. One embodiment is shown in Figure 11, where the scanning area is defined as a single line of consecutive coding blocks to the left or above the current block.
[0133] In some embodiments, the scanning direction may be a combination of vertical and horizontal scanning toward the side of the current block. One embodiment is shown in Figure 12. As shown in Figure 12, the scanning direction may also be a combination of parallel and diagonal. Scanning at position B begins from left to right, then diagonally toward the left and top blocks. Scanning at position B is repeated as shown in Figure 12. Similarly, scanning at position C begins from top to bottom, then diagonally toward the left and top blocks. Scanning at position C is repeated as shown in Figure 12.
[0134] Scanning order In some embodiments, the scanning order may be defined as going from a position with a shorter distance to the current coding block to a position with a longer distance. This order may apply to the case of vertical scanning.
[0135] In some embodiments, the scanning sequence may be defined as a fixed pattern. This fixed pattern scanning sequence can be used for candidate positions having similar distances. One embodiment is the case of horizontal scanning. In one embodiment, as in the embodiment shown in Figure 11, the scanning sequence may be defined as top-to-bottom for the left scanning area and as left-to-right for the upper scanning area.
[0136] In the case of the combined scanning method, the scanning sequence may be a combination of a fixed pattern and a distance-dependent sequence, similar to the embodiment shown in Figure 12.
[0137] Scanning complete For pre-built merge candidates, eligible candidates only require translational MVs and therefore do not need to be affine coded.
[0138] Depending on the number of candidates required, the scanning process may be completed when the first X qualified candidates are identified, where X is a positive value.
[0139] As shown in Figure 9, three corners named A, B, and C are required to form a virtual coding block. For a simpler implementation, the scanning process in step 1 may be performed only to identify non-adjacent neighbor blocks located at corners B and C, and the coordinates of A may be precisely determined by taking the horizontal coordinate of C and the vertical coordinate of B. In this way, the formed virtual coding block is constrained to be rectangular. In cases where either point B or C is unavailable, for example, outside the boundary, or where motion information for the non-adjacent neighbor block corresponding to B or C is unavailable, the horizontal or vertical coordinate of C may be defined as the horizontal or vertical coordinate of the top-leftmost point of the current block, respectively.
[0140] In another embodiment, when corners B and / or C are initially determined from the scanning process in step 1, non-adjacent neighbor blocks located at corners B and / or C may be identified accordingly. Secondly, the position(s) of corners B and / or C may be reset to a pivot point within the corresponding non-adjacent neighbor block, such as the center of mass of each non-adjacent neighbor block. For example, the center of mass may be defined as the geometric center of each neighbor block.
[0141] For the purpose of unification, the method for defining the scanning area and distance, scanning order, and scanning completion proposed for deriving inherited merge candidates may be reused, in whole or in part, for deriving constructed merge candidates. In one or more embodiments, but not limited to them, the same method defined for inherited merge candidate scanning, including the scanning area and distance, scanning order, and scanning completion, may be reused in whole for constructed merge candidate scanning.
[0142] In some embodiments, the same method defined for inherited merge candidate scanning may be partially reused for constructed merge candidate scanning. Figure 16 shows an example in this case. In Figure 16, the block size of each non-adjacent neighboring block is the same as the current block, and it is similarly defined as inherited candidate scanning, but the overall process is a simplified version as scanning at each distance is limited to only one block.
[0143] Figures 17A-17B illustrate another embodiment of this case. In Figures 17A-17B, both non-adjacent inherited merge candidates and non-adjacent constructed merge candidates are defined by having the same block size as the current coding block, while the scanning order, scanning area, and scanning completion conditions may be defined differently.
[0144] In Figure 17A, the maximum distance for non-adjacent neighbors on the left is 4 coding blocks, and the maximum distance for non-adjacent neighbors on the upper side is 5 coding blocks. Furthermore, at each distance, the scanning direction is down-up for the left side and right-left for the upper side. In Figure 17B, the maximum distance for non-adjacent neighbors is 4 for both the left and upper sides. In addition, scanning at a particular distance is unavailable because only one block exists at each distance. In Figure 17A, if M eligible candidates are identified, a scanning operation can be completed within each distance. The value of M can be a predefined fixed value such as 1 or any other positive integer, a signaled value determined by the encoder, or a configurable value in the encoder or decoder. In one embodiment, the value of M may be the same as the merge candidate list size.
[0145] In Figures 17A-17B, scanning operations can be completed at different distances when N qualified candidates are identified. The value of N can be a predefined fixed value such as 1 or any other positive integer, a signaled value determined by the encoder, or a configurable value in the encoder or decoder. In one embodiment, the value of N may be the same as the merge candidate list size. In another embodiment, the value of N may be the same as the value of M.
[0146] In both Figures 17A and 17B, non-adjacent spatial neighbors having a closer distance to the current block may be prioritized, indicating that a non-adjacent spatial neighbor with distance i is scanned or checked before a neighbor with distance i+1, where i can be a non-negative integer representing a specific distance.
[0147] At a given distance, a maximum of two non-adjacent spatial neighbors are used, meaning that, if available, a maximum of one neighbor from one side of the current block, e.g., left and top, is selected for inherited or constructed candidate derivation. As shown in Figure 17A, the checking order of the left and top neighbors is bottom-top and right-left, respectively. For Figure 17B, this rule may also apply, the difference being that at any given distance, there may be only one option for each side of the current block.
[0148] For the constructed candidate, as shown in Figure 17B, left The positions of the non-adjacent spatial neighbors above are determined independently first. Then, the position of the top-left neighbor, which together with the left and top non-adjacent neighbors can enclose the rectangular virtual block, can be determined accordingly. Subsequently, as shown in Figure 9, the motion information of the three non-adjacent neighbors is used to form the CPMV at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which are ultimately projected onto the current CU to generate the corresponding constructed candidates.
[0149] In Step 2, the translational MV at the selected candidate position after Step 1 is evaluated, and a suitable affine model can be determined. For easier illustration and without loss of generality, Figure 9 is used again as an example.
[0150] Due to factors such as hardware constraints, implementation complexity, and different reference indices, the scanning process may be completed before a sufficient number of candidates are identified. For example, motion information for the motion field may not be available for one or more of the selected candidates after step 1.
[0151] If motion information is available for all three candidates, the corresponding virtual coding block represents a 6-parameter affine model. If motion information is unavailable for one of the three candidates, the corresponding virtual coding block represents a 4-parameter affine model. If motion information is unavailable for more than one of the three candidates, the corresponding virtual coding block may be unavailable to represent a valid affine model.
[0152] In some embodiments, if motion information is unavailable for the top left corner of a virtual coding block, for example, corner A in Figure 9, or if motion information is unavailable for both the top right corner, for example, corner B in Figure 9, and the bottom left corner, for example, corner C in Figure 9, the virtual block may be set as invalid and unavailable to represent a valid model, and steps 3 and 4 may then be skipped during the current iteration.
[0153] In some embodiments, while not both are unavailable, if either the top right corner, e.g., corner B in Figure 9, or the bottom left corner, e.g., corner C in Figure 9, is unavailable, the virtual block may represent a valid four-parameter affine model.
[0154] In step 3, if the virtual coding block can represent a valid affine model, the same projection process used for inherited merge candidates may be applied.
[0155] In one or more embodiments, the same projection process used for inherited merge candidates may be used. In this case, a 4-parameter model represented by a virtual coding block from step 2 is projected onto the current block as a 4-parameter model, and a 6-parameter model represented by a virtual coding block from step 2 is projected onto the current block as a 6-parameter model.
[0156] In some embodiments, the affine model represented by the virtual coding block from step 2 is always projected onto a 4-parameter model or a 6-parameter model with respect to the current block.
[0157] According to equations (5) and (6), there may be two types of 4-parameter affine models: Type A is in which the top left corner CPMV and top right corner CPMV, referred to as V0 and V1, are available; and Type B is in which the top left corner CPMV and bottom left corner CPMV, referred to as V0 and V2, are available.
[0158] In one or more embodiments, the type of the projected 4-parameter affine model is the same type as the 4-parameter affine model represented by the virtual coding block. For example, the affine model represented by the virtual coding block from step 2 is a 4-parameter affine model of type A or B, and then the projected affine model for the current block is also type A or B, respectively.
[0159] In some embodiments, a 4-parameter affine model represented by a virtual coding block from step 2 is always projected onto the same type of 4-parameter model for the current block. For example, type A or B of a 4-parameter affine model represented by a virtual coding block is always projected onto type A of the 4-parameter affine model.
[0160] In step 4, based on the projected CPMV after step 3, in one embodiment, the same candidate generation process used in the current VVC or AVS standard may be used. In another embodiment, the time-motion vectors used in the candidate generation process used in the current VVC or AVS standard may not be used for the non-adjacent neighbor block-based derivation method. When time-motion vectors are not used, it indicates that the generated combination does not include any time-motion vectors.
[0161] In Step 5, any newly generated candidate after Step 4 may undergo a similarity check against all existing candidates already in the merge candidate list. Details of the similarity check have already been explained in the "Affine Merge Candidate Pruning" section. If a newly generated candidate is found to be similar to any existing candidate in the candidate list, the newly generated candidate is removed or pruned.
[0162] Inheritance base derivation method for affine-constructed merge candidates For each affine inherited candidate, all motion information is inherited from one selected spatial neighbor block coded in affine mode. This inherited information includes CPMV, reference index, prediction direction, affine model type, etc. On the other hand, for each affine constructed candidate, all motion information is constructed from two or three selected spatial or temporal neighbor blocks, and the selected neighbor blocks are not coded in affine mode; only translational motion information is required from the selected neighbor blocks.
[0163] This section discloses a novel candidate derivation method that combines the features of inherited and constructed candidates.
[0164] In some embodiments, a combination of inheritance and construction may be achieved by separating the affine model parameters into different groups, where one group of affine parameters is inherited from one neighborhood block and the other group of affine parameters is inherited from another neighborhood block.
[0165] In one embodiment, the parameters of an affine model are constructed from two groups. As shown in equation (3), the affine model may encompass six parameters, including a, b, c, d, e, and f. The translational parameters {a,b} may represent one group, while the non-translational parameters {c,d,e,f} may represent another group. This grouping method allows the two groups of parameters to be inherited independently from two different neighborhood blocks in a first step, and then concatenated / constructed to form a complete affine model in a second step. In this case, the group with non-translational parameters must be inherited from one affine-coded neighborhood block, while the group with translational parameters may be from any inter-coded neighborhood block that may or may not be coded in affine mode. Affine-coded neighbor blocks can be selected from adjacent or non-adjacent affine neighbor blocks based on a scanning method previously proposed for affine inherited candidates, such as the method shown in Figure 17A, which includes the scanning area and distance, scanning order, and scanning completion used in the section "Non-Adjacent Neighbor-Based Derivation Process for Affine Inherited Merge Candidates," and it should be noted that the scanning method can be performed on both adjacent and non-adjacent neighbor blocks. Alternatively, affine-coded neighbor blocks can be virtually constructed from regular inter-coded neighbor blocks, such as the method shown in Figure 17B, which includes the scanning area and distance, scanning order, and scanning completion used in the section "Non-Adjacent Neighbor-Based Derivation Process for Affine-Constructed Merge Candidates."
[0166] In some embodiments, the neighborhood blocks associated with each group may be determined in different ways. In one method, the neighborhood blocks for groups with different parameters may be all from non-adjacent neighborhood areas, and the scanning method may be designed similarly to previously proposed methods for non-adjacent neighborhood-based derivation. In another method, the neighborhood blocks for groups with different parameters may be all from adjacent neighborhood areas, and the scanning method may be the same as that of current VVC or AVS video standards. In yet another method, the neighborhood blocks for groups with different parameters may be adjacent neighbor It can be partial from an area, or partial from a neighboring area that is not adjacent.
[0167] When several groups of affine parameters are combined to construct a new candidate, there may be several rules that must be followed. The first is the eligibility criterion. In one embodiment, it may be checked whether the associated neighborhood blocks or blocks(s) within a group use the same reference picture for at least one or both directions. In another embodiment, it may be checked whether the associated neighborhood blocks or blocks(s) within a group use the same precision / resolution for the motion vector.
[0168] The second is the construction formula. In one embodiment, a new candidate CPMV can be derived in the following formula:
number
[0169] In another embodiment, a new candidate CPMV may be derived in the following formula:
number
[0170] Figure 18 shows an example of an inheritance-based derivation method for deriving an affine-constructed candidate. In Figure 18, there are three steps for deriving an affine-constructed candidate. In step 1, according to a specific grouping strategy, the encoder or decoder may perform scanning of adjacent and non-adjacent neighbor blocks for each group. In the case of Figure 18, two groups are defined, with neighbor 1 coded in affine mode and providing non-translational affine parameters, and neighbor 2 providing translational affine parameters. Neighbor 1 can be obtained according to the processing in the section “Non-adjacent Neighbor-Based Derivation Processing for Affine Inherited Merge Candidates” as shown in Figures 15A-15D and 17A, and neighbor 1 can be adjacent or non-adjacent neighbor blocks of the current block. Furthermore, neighbor 2 can be obtained according to the processing shown in Figures 16 and 17B.
[0171] In step 2, a specific affine model can be defined, from which different CPMVs can be derived according to the coordinates (x,y) of the CPMV, based on the parameters and positions determined in step 1. For example, as shown in Figure 18, the non-translational parameters {c,d,e,f} can be obtained based on neighborhood 1 obtained in step 1, and the translational parameters {a,b} can be obtained based on neighborhood 2 obtained in step 1. Furthermore, the distance parameters Δw and Δh can thus be obtained based on the position of the current block (x1,y1) and the position of neighborhood 2 (x2,y2). The distance parameters Δw and Δh can represent the horizontal and vertical distances between the current block and neighborhood 1 or neighborhood 2, respectively. For example, the distance parameters Δw and Δh can represent the horizontal distance (x1-x2) between the current block and neighborhood 2, and the vertical distance (y1-y2) between the current block and neighborhood 2, respectively. In particular, Δw = x1-x2 and Δh = y1-y2.
[0172] In step 3, two or three CPMVs are derived for the current coding block, and the current coding block may be constructed to form a new affine candidate.
[0173] In some embodiments, additional prediction information may be constructed. If it is checked that neighboring blocks have the same direction and / or reference picture, the prediction direction (e.g., bi-predicted or uni-predicted) and the index of the reference picture may be identical to those of the related neighboring blocks. Alternatively, the prediction information is determined by reusing the least overlapping information from related neighboring blocks from different groups. For example, if only the reference index for one direction from one neighboring block is identical to the reference index for the same direction in another neighboring block, the prediction direction for a new candidate is determined as a uni-predicted, and the same reference index and direction are reused.
[0174] HMVP-based derivation method for affine-constructed merge candidates In the case of near-neighborhood-based derivation, which is already defined in the current video standards VVC and AVS and described in the above section and Figure 7, the fixed order of scanning for near-neighborhoods is performed to identify two or three near-neighborhood blocks. In the case of non-neighborhood-based derivation, as proposed in the previous section and Figure 17B, two non-neighborhoods are identified during scanning in a different fixed order. In other words, for both near-neighborhood-based and non-neighborhood-based derivation methods, a certain depth of local scanning is necessary to identify the number of neighborhoods. This scanning process relies on local buffering around each current block and also incurs a certain amount of computational complexity.
[0175] On the other hand, as explained in the introductory section, the HMVP merge mode is already employed in current VVC and AVS systems, and translational motion information from neighboring blocks is already stored in a history table. In this case, the scanning process can be replaced by searching the HMVP table.
[0176] Therefore, for the previously proposed non-adjacent neighbor-based derivation and inheritance-based derivation processes, translational motion information can be obtained from the HMVP table instead of the scanning method shown in Figures 17B and 18. However, position information, width, height, and reference information are also required to derive affine-constructed candidates, which may be accessible if the current HMVP table is modified. Therefore, it is proposed to extend the HMVP table to store additional information in addition to the motion information of each historical neighbor. In one embodiment, the additional information may include the position of an affine or non-affine neighbor block, or affine motion information such as CPMV or equivalent normal motion derived from CPMV (for example, this normal motion may be from an internal sub-block of an affine-coded neighbor block), a reference index, and so on.
[0177] Candidate derivation method for affine AMVP and normal merge mode As explained in the section above, for affine AMVP mode, an affine candidate list is also required to derive the CPMV predictor. Consequently, all the derivation methods proposed above can be applied similarly to affine AMVP mode. The only difference is that when the derivation methods proposed above are applied in AMVP, the selected neighboring block should have the same reference picture index as the current coding block.
[0178] For the normal merge mode, the candidate list is also constructed using only translational candidate MVs, not CPMVs. In this case, all the derivation methods proposed above can still be applied by adding an additional derivation step. This additional derivation step is to derive the translational MV for the current block, which can be achieved by selecting a specific pivot position (x,y) within the current block and adhering to the same equation (3). In other words, to derive the CPMV for an affine block, the positions of the three corners of the block are used as pivot positions (x,y) in equation (3), and to derive the translational MV for a normal intercoded block, the center position of the block can be used as a pivot position (x,y) in equation (3). Once the translational MV is derived for the current block, it can be inserted into the candidate list as another candidate.
[0179] Reordering the list of Affine Merge candidates In one embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. Subblock-based temporal motion vector prediction (SbTMVP) candidate, if available; 2. Inherited from adjacent neighbors; 3. Inherited from non-adjacent neighbors; 4. Constructed from adjacent neighbors; 5. Constructed from non-adjacent neighbors; 6. Zero MV.
[0180] In another embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available; 2. Inherited from adjacent neighbors; 3. Constructed from adjacent neighbors; 4. Inherited from non-adjacent neighbors; 5. Constructed from non-adjacent neighbors; 6. Zero MV.
[0181] In another embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available; 2. inherited from adjacent neighbors; 3. constructed from adjacent neighbors; 4. one set of zero MVs; 5. inherited from non-adjacent neighbors; 6. constructed from non-adjacent neighbors; 7. remaining zero MVs if the list is still not full.
[0182] In another embodiment, non-proximate spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available; 2. Inherited from a proximate neighbor; 3. Inherited from a non-proximate neighbor by a distance shorter than X; 4. Constructed from a proximate neighbor; 5. Constructed from a non-proximate neighbor by a distance shorter than Y; 6. Inherited from a non-proximate neighbor by a distance longer than X; 7. Constructed from a non-proximate neighbor by a distance longer than Y; 8. Zero MV. In this embodiment, the values X and Y may be predefined fixed values such as the values of 2, or signaled values determined by the encoder, or configurable values in the encoder or decoder. In one embodiment, the value of X may be identical to the value of Y. In another embodiment, X The value of is Y The value may differ from that of [the other value].
[0183] Figure 19 shows a computing environment (or computing device) 1910 coupled to a user interface 1960. The computing environment 1910 may be part of a data processing server. In some embodiments, the computing device 1910 may perform any of the various methods or processes (encoding / decoding methods or processes) described below, relating to various embodiments of this disclosure. The computing environment 1910 may include a processor 1920, memory 1940, and an I / O interface 1950.
[0184] Processor 1920 typically controls the overall computation of the computing environment 1910, including operations associated with display, data acquisition, data communication, and image processing. Processor 1920 may include one or more processors for executing instructions that perform all or part of the steps in the manner described above. Furthermore, processor 1920 may include one or more modules that facilitate interaction between processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, or a GPU, etc.
[0185] Memory 1940 is configured to store various types of data to support the computations of the computing environment 1910. Memory 1940 may include predetermined software 1942. Examples of such data include instructions for any application or method operating on the computing environment 1910, video datasets, image data, etc. Memory 1940 may be implemented using any type of volatile or non-volatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disks, etc.
[0186] The I / O interface 1950 provides an interface between the processor 1920 and peripheral interface modules, such as a keyboard, click wheel, and buttons. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1950 can be coupled with an encoder and a decoder.
[0187] In some embodiments, a non-temporary computer-readable storage medium containing multiple programs, such as those contained in memory 1940, is also provided, which can be executed by a processor 1920 in a computing environment 1910 for performing the methods described above. For example, the non-temporary computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0188] A non-temporary computer-readable storage medium stores therein multiple programs for execution by a computing device having one or more processors, and when the multiple programs are executed by one or more processors, they cause the computing device to perform the methods for motion prediction described above.
[0189] In some embodiments, the computing environment 1910 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic circuits (PLDs), field-programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0190] Figure 20 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.
[0191] In step 2001, processor 1920 may obtain one or more first parameters based on the first neighboring block of the current block.
[0192] In some embodiments, one or more first parameters may include multiple non-translational parameters associated with the affine model. For example, as shown in Figure 18, one or more first parameters may include non-translational parameters c, d, e, and f inherited from an affine-coded first neighborhood block.
[0193] In some embodiments, the first neighborhood block may be obtained from a plurality of adjacent neighborhood blocks and a plurality of non-adjacent neighborhood blocks. That is, the first neighborhood block may be an adjacent neighborhood block or a non-adjacent neighborhood block. The plurality of adjacent neighborhood blocks are adjacent to the current block, and the plurality of non-adjacent neighborhood blocks are each located a number of blocks away from one side of the current block.
[0194] In some embodiments, the first neighbor block may be obtained from multiple intercoded neighbor blocks of the current block, and the multiple intercoded neighbor blocks may include affine-coded blocks.
[0195] In step 2002, processor 1920 may obtain one or more second parameters based on the first and / or second neighboring blocks of the current block.
[0196] In particular, the processor 1920 may obtain one or more second parameters based on a first neighboring block, a second neighboring block, or the first and second neighboring blocks.
[0197] In some embodiments, one or more second parameters may include multiple translation parameters associated with the affine model. For example, as shown in Figure 18, one or more second parameters may include translation parameters a, b constructed based on a second neighborhood block.
[0198] In some embodiments, a second neighborhood block may be obtained from multiple intercoded neighborhood blocks of the current block, and the multiple intercoded neighborhood blocks may include affine-coded blocks and non-affine-coded blocks.
[0199] In some embodiments, the first neighbor block may be obtained from a plurality of non-proximate neighbor blocks based on a first scanning rule, where each of the non-proximate neighbor blocks is located a number of blocks away from one side of the current block. For example, the first scanning rule may be a scanning rule that includes a scanning area and distance, a scanning order, and a scanning completion, as used in the section “Non-proximate Neighbor-Based Derivation Process for Affine Inherited Merge Candidates”, and the scanning rule may be applied to both proximate and non-proximate neighbor blocks, as shown in Figures 8, 13A–13B, 14A–14B, 15A–15D, and 17A.
[0200] In some embodiments, a second neighborhood block may be obtained from multiple non-adjacent neighborhood blocks based on a second scanning rule, and the second scanning rule may be entirely or partially identical to the first scanning rule. For example, the second scanning rule may be a scanning rule that includes the scanning area and distance, scanning order, and scanning completion used in the section “Non-adjacent Neighborhood-Based Derivation Process for Affine-Constructed Merge Candidates”, and the scanning rule may be applied to both adjacent and non-adjacent neighborhood blocks, as shown in Figures 9-12, 16, and 17B.
[0201] In step 2003, the processor 1920 may construct one or more affine models by using one or more first parameters and one or more second parameters.
[0202] In some embodiments, one or more first parameters and one or more second parameters may be combined or linked to construct one or more affine models.
[0203] In step 2004, processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models constructed in step 2003.
[0204] In some embodiments, the processor 1920 may determine that a first neighborhood block and a second neighborhood block are valid for constructing one or more affine models under certain preconditions. In one embodiment, the processor 1920 may determine that a first neighborhood block and a second neighborhood block are valid for constructing an affine model in response to determining that the first neighborhood block and the second neighborhood block use the same reference picture for at least one motion direction. Furthermore, in response to determining that the first neighborhood block and the second neighborhood block use the same reference picture for one motion direction, the processor 1920 may determine that the predicted direction of one or more motion vector candidates formed based on CPMVs is a single prediction and that the same reference picture is used for one motion vector candidate. The processor 1920 may also determine that the predicted direction and reference picture of the current block are identical to the predicted direction and reference picture of the first and second neighborhood blocks, respectively, in response to determining that the first and second neighborhood blocks use the same reference picture for both motion directions. Here, one or more CPMVs for the current block obtained in step 2004 may be constructed to form motion vector candidates. Motion vector candidates are not limited to affine candidates and may include normal merge candidates, AMVP candidates, and so on.
[0205] In another embodiment, the processor 1920 may determine that the first and second neighbor blocks are valid for constructing an affine model, in response to determining that the first and second neighbor blocks use the same resolution for motion vectors.
[0206] In some embodiments, the processor 1920 may construct one or more affine models based on one or more first parameters, one or more second parameters, a first position of the current block, and a second neighboring block or a second position of the first neighboring block. For example, as shown in step 2 in Figure 18, the affine model may be constructed based on non-translational parameters c, d, e, f, translational parameters a, b, and the difference between the current block and the second neighboring block. For example, the difference may include a corresponding coordinate difference as shown in Figure 18. The position of the current block, the first and second neighboring blocks may be determined in different ways.
[0207] In some embodiments, the first position of the current block may be determined according to the upper left corner of the current block, and the second position of the first or second neighboring block may be determined according to the upper left corner of the first or second neighboring block.
[0208] In some embodiments, one or more first parameters may include multiple parameters associated with the affine model, and one or more second parameters may include multiple distance parameters. For example, as shown in Figure 18, one or more first parameters may include affine model parameters a, b, c, d, e, f, and one or more second parameters may include distance parameters Δw and Δh.
[0209] In some embodiments, multiple distance parameters may be predefined as fixed values. For example, the value of (Δw, Δh) may be predefined as a fixed value such as (0, 0) or any constant value.
[0210] In some embodiments, each of the distance parameters may represent the distance between the current block and a first or second neighboring block. For example, the distance parameters may include a first distance parameter Δw representing the horizontal distance between the current block and the first or second neighboring block, and a second distance parameter Δh representing the vertical distance between the current block and the first or second neighboring block.
[0211] Figure 21 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.
[0212] In step 2101, the processor 1920 may obtain multiple motion vector candidates from the HMVP table, which may include a first motion vector constructed candidate and a second motion vector constructed candidate.
[0213] In some embodiments, the multiple motion vector candidates may not be limited to affine candidates, but may include normal merge candidates, AMVP candidates, and so on.
[0214] In some embodiments, the HMVP table may be extended by storing additional information in the HMVP table in addition to the motion information of each historical neighbor block. This additional information may include: the position of each historical neighbor block, the affine motion information of each historical neighbor block, or at least one or more reference indices of each historical neighbor block.
[0215] In step 2102, the processor 1920 may acquire a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, as shown in Figure 9.
[0216] In step 2103, processor 1920 may obtain multiple CPMVs for the current block based on multiple CPMVs of the virtual block.
[0217] In some embodiments, the processor 1920 may determine a third motion vector constructed candidate based on first and second motion vector constructed candidates and virtual blocks, obtain multiple CPMVs for virtual blocks based on the translational MVs of the first, second, and third motion vector constructed candidates, and obtain multiple CPMVs for the current block based on the multiple CPMVs of virtual blocks by using the same projection process used for inherited candidate derivation.
[0218] Figure 22 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.
[0219] In step 2201, the processor 1920 may obtain one or more motion vector candidates from multiple non-adjacent neighboring blocks to the current block based on at least one scanning distance, where at least one of the scanning distances indicates the number of blocks away from one side of the current block.
[0220] In step 2202, processor 1920 may obtain one or more CPMVs for the current block based on one or more motion vector candidates.
[0221] In some embodiments, One or more Motion vector candidates are not limited to affine candidates, but may also include normal merge candidates, AMVP candidates, and others.
[0222] In some embodiments, the processor 1920 may add one or more motion vector candidates to the affine candidate list for the affine AMVP mode in response to determining that one or more motion vector candidates have the same reference picture index as the current block.
[0223] In some embodiments, the processor 1920 may obtain at least one translational motion vector for the current block based on one or more CPMVs, and may add the at least one translational motion vector to a normal merge candidate list for the normal merge mode.
[0224] In some embodiments, the processor 1920 may obtain at least one translational motion vector for the current block based on one or more CPMVs by selecting a specific rotation position within the current block.
[0225] FIG. 23 is a flowchart illustrating a method for video encoding corresponding to the method as illustrated in FIG. 20.
[0226] In step 2301, the processor 1920 may determine one or more first parameters based on the first neighboring block of the current block.
[0227] In step 2302, the processor 1920 may determine one or more second parameters based on the first neighboring block and / or the second neighboring block of the current block.
[0228] In particular, the processor 1920 may determine one or more second parameters based on the first neighboring block, the second neighboring block, or the first neighboring block and the second neighboring block.
[0229] In step 2303, the processor 1920 may construct one or more affine models by using the one or more first parameters and the one or more second parameters.
[0230] In some embodiments, the one or more first parameters and the one or more second parameters may be combined or concatenated to construct one or more affine models.
[0231] In step 2304, the processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models constructed in step 2303.
[0232] Figure 24 is a flowchart illustrating a method for video encoding corresponding to the method illustrated in Figure 21.
[0233] In step 2401, the processor 1920 may determine multiple motion vector candidates from the HMVP table, which may include a first motion vector constructed candidate and a second motion vector constructed candidate.
[0234] In step 2402, the processor 1920 may acquire a virtual block based on the first motion vector constructed candidate and the second motion vector constructed candidate, as shown in Figure 9.
[0235] In step 2403, processor 1920 may obtain multiple CPMVs for the current block based on multiple CPMVs of the virtual block.
[0236] Figure 25 is a flowchart illustrating a method for video encoding corresponding to the method illustrated in Figure 22.
[0237] In step 2501, the processor 1920 may determine one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning distance, where at least one of the scanning distances indicates the number of blocks away from one side of the current block.
[0238] In step 2502, the processor 1920 may obtain one or more CPMVs for the current block based on one or more motion vector candidates.
[0239] Figure 26 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.
[0240] In step 2601, one or more first parameters can be obtained using an inheritance-based derivation method.
[0241] In some embodiments, the processor 1920 may use an inheritance-based derivation method to obtain a first neighbor block from a plurality of intercoded neighbor blocks of the current block, and based on the first neighbor block, obtain one or more first parameters, the plurality of intercoded neighbor blocks may include affine-coded blocks.
[0242] In some embodiments, the inheritance-based derivation method may be a derivation process for an affine inherited merge candidate as described in the section “Non-proximate Neighbor-Based Derivation Process for Affine Inherited Merge Candidates”. In the inheritance-based derivation method, neighboring blocks of the current block may be scanned using a scanning method / rule that includes the scanning area and distance, scanning order, and scanning completion used in the section “Non-proximate Neighbor-Based Derivation Process for Affine Inherited Merge Candidates”, and the scanning rule may be applied to both adjacent neighbor blocks or non-adjacent neighbor blocks, as shown in Figures 8, 13A-13B, 14A-14B, 15A-15D, and 17A.
[0243] In some embodiments, one or more first parameters may include a plurality of parameters associated with an affine model, and one or more second parameters may include a plurality of distance parameters, where the plurality of distance parameters may include a first distance parameter indicating a horizontal distance between a current block and a first neighboring block, and a second distance parameter indicating a vertical distance between the current block and the first neighboring block. The plurality of parameters associated with the affine model may include the parameters {a, b, c, d, e, f} associated with the affine model. The first distance parameter and the second distance parameter may be the distance parameters Δw and Δh, respectively.
[0244] In step 2602, the processor 1920 may obtain one or more second parameters using a construction-based derivation method.
[0245] In some embodiments, the processor 1920 may obtain a second neighboring block from a plurality of inter-coded neighboring blocks of a current block using a construction-based derivation method, and based on the second neighboring block, obtain one or more second parameters, where the plurality of inter-coded neighboring blocks may include affine-coded blocks and non-affine-coded blocks.
[0246] In some embodiments, the construction-based derivation method may be a derivation process for an affine-constructed merge candidate described in the section "Proximity-free Neighbor-based Derivation Process for Affine-Constructed Merge Candidates". In the construction-based derivation method, the neighboring blocks of the current block may be scanned using a scanning method / rule that includes a scanning area and distance, a scanning order, and a scanning completion used in the section "Proximity-free Neighbor-based Derivation Process for Affine-Constructed Merge Candidates", and the scanning rule may be executed for both neighboring blocks that are adjacent or not adjacent, as shown in FIGS. 9-12, 16, and 17B.
[0247] In some embodiments, one or more first parameters may include a plurality of parameters associated with an affine model, and one or more second parameters may include a plurality of distance parameters, the plurality of distance parameters may include a first distance parameter indicating the horizontal distance between the current block and a first neighboring block, and a second distance parameter indicating the vertical distance between the current block and a first neighboring block. The plurality of parameters associated with an affine model may include parameters {a, b, c, d, e, f} associated with an affine model. The first and second distance parameters may be distance parameters Δw and Δh, respectively.
[0248] In some embodiments, one or more first parameters may include a plurality of non-translational parameters associated with the affine model, and one or more second parameters may include a plurality of translational parameters associated with the affine model.
[0249] In some embodiments, one or more first parameters may include multiple parameters associated with the affine model, and one or more second parameters may include multiple distance parameters.
[0250] In some embodiments, multiple distance parameters can be predefined as fixed values.
[0251] In step 2603, the processor 1920 may construct one or more affine models by using one or more first parameters and one or more second parameters.
[0252] In step 2604, processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models.
[0253] Figure 27 is a flowchart illustrating a method for video encoding corresponding to the method illustrated in Figure 26.
[0254] In step 2701, the processor 1920 may determine one or more first parameters using an inheritance-based derivation method.
[0255] In step 2702, the processor 1920 may determine one or more second parameters using a build-based derivation method.
[0256] In step 2703, the processor 1920 may construct one or more affine models using one or more first parameters and one or more second parameters.
[0257] In step 2704, processor 1920 may obtain one or more CPMVs for the current block based on one or more affine models.
[0258] In some embodiments, an apparatus for video coding is provided. The apparatus includes a processor 1920 and a memory 1940 configured to store instructions that can be executed by the processor, and the processor is configured to perform one of the methods illustrated in Figures 20-27 when an instruction is executed.
[0259] In some other embodiments, a non-temporary computer-readable storage medium having instructions stored therein is provided. When an instruction is executed by the processor 1920, the instruction causes the processor to perform one of the methods illustrated in Figures 20-25.
[0260] Other embodiments of the disclosure will be apparent to those skilled in the art from the specification and practice considerations of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the disclosure, including such deviations, in accordance with the general principles therein, as is within known or common practice in the art. The specification and embodiments are intended to be considered illustrative only.
[0261] This disclosure is not limited to the exact embodiments described above and illustrated in the accompanying drawings, and it should be acknowledged that various modifications and changes may be made without departing from its scope.
Claims
1. A method for video decoding, Based on the first neighbor block of the current block, obtain one or more first parameters, Obtaining one or more second parameters based on the first neighboring block and / or second neighboring block of the current block, By using the one or more first parameters and the one or more second parameters, one or more affine models are constructed. Based on the one or more affine models, one or more control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
2. The method according to claim 1, further comprising obtaining the first neighborhood block from a plurality of adjacent neighborhood blocks and a plurality of non-adjacent neighborhood blocks, wherein the plurality of adjacent neighborhood blocks are adjacent to the current block and the plurality of non-adjacent neighborhood blocks are each located in a number of blocks away from one side of the current block.
3. The method according to claim 1, further comprising obtaining the second neighbor block from a plurality of inter-coded neighbor blocks of the current block, wherein the plurality of inter-coded neighbor blocks include affine-coded blocks and non-affine-coded blocks.
4. The method according to claim 1, further comprising obtaining the first neighbor block from a plurality of intercoded neighbor blocks of the current block, wherein the plurality of intercoded neighbor blocks include affine-coded blocks.
5. Acquiring a first neighboring block from a plurality of non-adjacent neighboring blocks based on a first scanning rule, wherein each of the plurality of non-adjacent neighboring blocks is located a number of blocks away from one side of the current block. Acquiring a second neighboring block from a plurality of non-adjacent neighboring blocks based on a second scanning rule, wherein the second scanning rule is entirely or partially identical to the first scanning rule. The method according to claim 1, further comprising:
6. The method according to claim 1, wherein the one or more first parameters include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters include a plurality of translational parameters associated with the affine model.
7. The method of claim 2, further comprising determining that the first neighborhood block and the second neighborhood block are valid in response to the determination that the first neighborhood block and the second neighborhood block use the same reference picture for at least one direction of motion.
8. The method according to claim 7, further comprising determining, in response to the first neighborhood block and the second neighborhood block deciding to use the same reference picture for one direction of motion, that the predicted direction of the motion vector candidates formed based on the one or more CPMVs is a uni-prediction and that the same reference picture is used for the motion vector candidates for the one direction of motion.
9. The method according to claim 7, further comprising determining that the predicted direction and reference picture of the current block are identical to those of the first and second neighboring blocks, respectively, in response to the determination that the first and second neighboring blocks use the same reference picture for both directions of movement.
10. The method of claim 2, further comprising determining that the first neighborhood block and the second neighborhood block are valid in response to the determination that the first neighborhood block and the second neighborhood block use the same resolution for motion vectors.
11. The method according to claim 1, further comprising constructing one or more affine models based on one or more first parameters, one or more second parameters, a first position of the current block, and a second neighborhood block or a second position of the first neighborhood block.
12. The method according to claim 11, wherein the first position includes the upper left corner of the current block, and the second position includes the upper left corner of the first or second neighboring block.
13. The method according to claim 1, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters.
14. The method according to claim 13, wherein the plurality of distance parameters include a first distance parameter indicating the horizontal distance between the current block and the second neighboring block, and a second distance parameter indicating the vertical distance between the current block and the second neighboring block.
15. The method according to claim 13, wherein the plurality of distance parameters include a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
16. A method for video decoding, The method involves obtaining multiple motion vector candidates from a history-based motion vector prediction (HMVP) table, wherein the multiple motion vector candidates include a first motion vector construction candidate and a second motion vector construction candidate. Based on the first motion vector constructed candidate and the second motion vector constructed candidate, a virtual block is obtained, Based on the multiple CPMVs of the aforementioned virtual block, multiple control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
17. The method according to claim 16, further comprising extending the HMVP table by storing additional information in the HMVP table in addition to the movement information of each history neighbor block.
18. The aforementioned additional information is as follows: The position of each history neighboring block, Affine motion information for each history neighboring block, or Reference index of each history neighbor block The method according to claim 17, comprising at least one of the following.
19. Based on the first and second motion vector construction candidates and the virtual block, a third motion vector construction candidate is determined. Based on the translational MVs of the first, second, and third motion vector constructed candidates, the plurality of CPMVs for the virtual block are obtained, By using the same projection process used for inherited candidate derivation, the multiple CPMVs for the current block are obtained based on the multiple CPMVs of the virtual block, The method according to claim 16, further comprising:
20. A method for video decoding, Obtaining one or more motion vector candidates from multiple non-adjacent neighboring blocks to the current block based on at least one scanning distance, wherein one of the at least one scanning distances indicates the number of blocks away from one side of the current block. Based on the one or more motion vector candidates, one or more control point motion vectors (CPMV) for the current block are obtained, Methods that include...
21. The method according to claim 20, further comprising adding the one or more motion vector candidates to an affine candidate list for an affine advanced motion vector prediction (AMVP) mode in response to the determination that the one or more motion vector candidates have the same reference picture index as the current block.
22. Based on the one or more CPMVs, obtain at least one translational motion vector for the current block, Adding the aforementioned at least one translational motion vector to the list of normal merge candidates for the normal merge mode, The method according to claim 21, further comprising:
23. The method according to claim 22, wherein obtaining the at least one translational motion vector for the current block based on the one or more CPMVs includes obtaining the at least one translational motion vector for the current block by selecting a specific pivot position within the current block.
24. A method of video encoding, Determine one or more first parameters based on the first neighboring blocks of the current block, Determining one or more second parameters based on the first neighboring block and / or second neighboring block of the current block, By using the one or more first parameters and the one or more second parameters, one or more affine models are constructed. Based on the one or more affine models, one or more control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
25. The method according to claim 24, further comprising obtaining a first neighborhood block from a plurality of adjacent neighborhood blocks and a plurality of non-adjacent neighborhood blocks, wherein the plurality of adjacent neighborhood blocks are adjacent to the current block and the plurality of non-adjacent neighborhood blocks are each located in a number of blocks away from one side of the current block.
26. The method of claim 24, further comprising obtaining the second neighbor block from a plurality of intercoded neighbor blocks of the current block, wherein the plurality of intercoded neighbor blocks include affine-coded blocks and non-affine-coded blocks.
27. The method of claim 24, further comprising obtaining the first neighbor block from a plurality of intercoded neighbor blocks of the current block, wherein the plurality of intercoded neighbor blocks include affine-coded blocks.
28. Acquiring a first neighboring block from a plurality of non-adjacent neighboring blocks based on a first scanning rule, wherein each of the plurality of non-adjacent neighboring blocks is located a number of blocks away from one side of the current block. Acquiring a second neighboring block from a plurality of non-adjacent neighboring blocks based on a second scanning rule, wherein the second scanning rule is entirely or partially identical to the first scanning rule. The method according to claim 24, further comprising:
29. The method according to claim 24, wherein the one or more first parameters include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters include a plurality of translational parameters associated with the affine model.
30. The method of claim 25, further comprising determining that the first neighborhood block and the second neighborhood block are valid in response to the determination that the first neighborhood block and the second neighborhood block use the same reference picture for at least one direction of motion.
31. The method according to claim 30, further comprising determining, in response to the first neighborhood block and the second neighborhood block deciding to use the same reference picture for one direction of motion, that the predicted direction of the motion vector candidates formed based on the one or more CPMVs is a single prediction and that the same reference picture is used for the motion vector candidates for the one direction of motion.
32. The method according to claim 30, further comprising determining that the predicted direction and reference picture of the current block are identical to those of the first and second neighboring blocks, respectively, in response to the determination that the first and second neighboring blocks use the same reference picture for both directions of movement.
33. The method of claim 25, further comprising determining that the first neighborhood block and the second neighborhood block are valid in response to the determination that the first neighborhood block and the second neighborhood block use the same resolution for motion vectors.
34. The method according to claim 24, further comprising constructing one or more CPMVs based on one or more first parameters, one or more second parameters, a first position of the current block, and a second neighboring block or a second position of the first neighboring block.
35. The method according to claim 34, wherein the first position includes the upper left corner of the current block, and the second position includes the upper left corner of the first or second neighboring block.
36. The method according to claim 24, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters.
37. The method according to claim 36, wherein the plurality of distance parameters include a first distance parameter indicating the horizontal distance between the current block and the second neighboring block, and a second distance parameter indicating the vertical distance between the current block and the second neighboring block.
38. The method according to claim 36, wherein the plurality of distance parameters include a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
39. A method of video encoding, The process involves determining a plurality of motion vector candidates from a history-based motion vector prediction (HMVP) table, wherein the plurality of motion vector candidates include a first motion vector construction candidate and a second motion vector construction candidate. Based on the first motion vector constructed candidate and the second motion vector constructed candidate, a virtual block is determined, Based on the multiple CPMVs of the aforementioned virtual block, multiple control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
40. The method according to claim 39, further comprising extending the HMVP table by storing additional information in the HMVP table in addition to the movement information of each history neighbor block.
41. The aforementioned additional information is as follows: The position of each history neighboring block, Affine motion information for each history neighboring block, or Reference index of each history neighbor block The method according to claim 40, comprising at least one of the following.
42. Based on the first and second motion vector construction candidates and virtual blocks, a third motion vector construction candidate is determined. Based on the translational MVs of the first, second, and third motion vector constructed candidates, the plurality of CPMVs for the virtual block are obtained, By using the same projection process used for inherited candidate derivation, the multiple CPMVs for the current block are obtained based on the multiple CPMVs of the virtual block, The method according to claim 41, further comprising:
43. A method of video encoding, Determining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning distance, wherein one of the at least one scanning distances indicates the number of blocks away from one side of the current block. Based on the one or more motion vector candidates, one or more control point motion vectors (CPMV) for the current block are obtained, Methods that include...
44. The method according to claim 43, further comprising adding the one or more motion vector candidates to the affine candidate list for the affine advanced motion vector prediction (AMVP) mode in response to the determination that the one or more motion vector candidates have the same reference picture index as the current block.
45. Based on the one or more CPMVs, obtain at least one translational motion vector for the current block, Adding the aforementioned at least one translational motion vector to the list of normal merge candidates for the normal merge mode, The method according to claim 44, further comprising:
46. The method according to claim 45, wherein obtaining the at least one translational motion vector for the current block based on the one or more CPMVs includes obtaining the at least one translational motion vector for the current block by selecting a specific pivot position within the current block.
47. A method for video decoding, Using inheritance-based derivation methods, obtain one or more first parameters, Using a construct-based derivation method, obtain one or more second parameters, Using the one or more first parameters and the one or more second parameters, one or more affine models are constructed. Based on the aforementioned one or more affine models, one or more control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
48. The method according to claim 47, wherein the one or more first parameters include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters include a plurality of translational parameters associated with the affine model.
49. The method according to claim 48, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters.
50. The method according to claim 49, wherein the plurality of distance parameters are defined in advance as fixed values.
51. Obtaining a first neighboring block from a plurality of intercoded neighboring blocks of the current block using the inheritance-based derivation method, wherein the plurality of intercoded neighboring blocks include affine-coded blocks. Based on the first neighboring block, one or more first parameters are obtained, The method according to claim 47, further comprising:
52. The method according to claim 51, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters, the plurality of distance parameters including a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
53. Obtaining a second neighboring block from a plurality of intercoded neighboring blocks of the current block using the construction base derivation method, wherein the plurality of intercoded neighboring blocks include affine-coded blocks and non-affine-coded blocks. Based on the second neighborhood block, one or more second parameters are obtained, The method according to claim 47, further comprising:
54. The method according to claim 53, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters, the plurality of distance parameters including a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
55. A method of video encoding, Using an inheritance-based derivation method, determine one or more first parameters, Using a construction-based derivation method, determine one or more second parameters, Using the one or more first parameters and the one or more second parameters, one or more affine models are constructed. Based on the aforementioned one or more affine models, one or more control point motion vectors (CPMVs) for the current block are obtained, Methods that include...
56. The method according to claim 55, wherein the one or more first parameters include a plurality of non-translational parameters associated with an affine model, and the one or more second parameters include a plurality of translational parameters associated with the affine model.
57. The method according to claim 56, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters.
58. The method according to claim 57, wherein the plurality of distance parameters are defined in advance as fixed values.
59. Obtaining a first neighboring block from a plurality of intercoded neighboring blocks of the current block using the inheritance-based derivation method, wherein the plurality of intercoded neighboring blocks include affine-coded blocks. Based on the first neighboring block, one or more first parameters are obtained, The method according to claim 55, further comprising:
60. The method according to claim 59, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters, the plurality of distance parameters including a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
61. Obtaining a second neighboring block from a plurality of intercoded neighboring blocks of the current block using the construction base derivation method, wherein the plurality of intercoded neighboring blocks include affine-coded blocks and non-affine-coded blocks. Based on the second neighborhood block, one or more second parameters are obtained, The method according to claim 55, further comprising:
62. The method according to claim 61, wherein the one or more first parameters include a plurality of parameters associated with an affine model, and the one or more second parameters include a plurality of distance parameters, the plurality of distance parameters including a first distance parameter indicating the horizontal distance between the current block and the first neighboring block, and a second distance parameter indicating the vertical distance between the current block and the first neighboring block.
63. A device for video decoding, One or more processors, Includes a memory coupled to one or more processors and configured to store instructions executable by the one or more processors, The one or more processors are configured to perform the method according to any one of claims 1 to 23 and 47 to 54 when the instruction is executed. Device.
64. A device for video encoding, One or more processors, Includes a memory coupled to one or more processors and configured to store instructions executable by the one or more processors, The one or more processors are configured to perform the method described in any one of claims 24 to 46 and 55 to 62 when the instruction is executed. Device.
65. A non-temporary computer-readable storage medium storing computer-executable instructions, wherein, when the computer-executable instructions are executed by one or more computer processors, the non-temporary computer-readable storage medium causes the one or more computer processors to execute the method according to any one of claims 1 to 62.