Candidate Derivation of Affine Merge Modes in Video Coding and Decoding

By deriving affine merge candidates from both adjacent and non-adjacent neighboring blocks using scanning distances and CPMVs, the method addresses the challenge of inefficient affine motion prediction in video encoding and decoding, resulting in improved compression efficiency and quality.

JP7720482B2Active Publication Date: 2025-08-07BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024519344
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-29
Filing Date
2022-09-29
Publication Date
2025-08-07
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

Existing video encoding and decoding standards, such as VVC and AVS3, face challenges in accurately deriving affine merge candidates for affine motion prediction modes, limiting the efficiency of video compression and quality preservation.

Method used

The method involves obtaining motion vector candidates from both adjacent and non-adjacent neighboring blocks based on scanning distances and using control point motion vectors (CPMVs) to enhance the derivation process, with candidate pruning and non-adjacent block-based derivation processes to improve candidate diversity and accuracy.

Benefits of technology

This approach increases the diversity and accuracy of affine merge candidates, leading to improved video compression efficiency and quality by leveraging spatial correlations beyond adjacent blocks, thus enhancing the coding performance of video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007720482000018
    Figure 0007720482000018
  • Figure 0007720482000019
    Figure 0007720482000019
  • Figure 0007720482000020
    Figure 0007720482000020
Patent Text Reader

Abstract

A method, an apparatus, and a non-transitory computer-readable storage medium thereof for video encoding are provided. The method includes obtaining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from a side of the current block, and the number may be a positive integer. Furthermore, the method may include obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more motion vector candidates.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 250,184, filed September 29, 2021, entitled "Candidate Derivation for Affine Merge Mode in Video Coding," the entire contents of which are incorporated herein by reference. [Technical Field]

[0002] This disclosure relates to video encoding / decoding and compression, and more particularly, but not exclusively, to methods and apparatus for improving affine merge candidate derivation for affine motion prediction modes in a video encoding or decoding process. [Background technology]

[0003] Various video encoding / decoding techniques can be used to compress video data. Video encoding / decoding is performed according to one or more video encoding / decoding standards. For example, some well-known video encoding / decoding standards today include the Universal Video Coding / Decoding (VVC), jointly developed by ISO / IEC MPEG and ITU-T VECG, High Efficiency Video Coding / Decoding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding / Decoding (AVC, also known as H.264 or MPEG-4 Part 10). AO Media Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the predecessor standard VP9. Audio-Video Coding / Decoding (AVS), referring to digital audio and digital video compression standards, is another video compression standard series developed by the China Audio and Video Coding Standards Workgroup. Many of the existing video encoding and decoding standards are built on well-known hybrid video encoding and decoding frameworks, such as using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy present in a video image or sequence, and using transform coding to compress the energy of prediction errors, etc. One important goal of video encoding and decoding technology is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation of video quality.

[0004] The first generation of AVS standards includes the Chinese national standards "Information Technology, Advanced Audio-Video Coding / Decoding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio-Video Coding / Decoding, Part 16: Radio-Television Video" (known as AVS+). Compared to the MPEG-2 standard, it can save approximately 50% of the bitrate at the same perceptual quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards includes a series of Chinese national standards, "Information Technology, Efficient Multimedia Coding" (also known as AVS2), primarily targeted at the transmission of additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was published as a Chinese national standard. Meanwhile, the video portion of the AVS2 standard has been submitted by the Institute of Electrical and Electronics Engineers (IEEE) as one of the international standards for applications. The AVS3 standard is a new generation of video coding / decoding standard for UHD video applications, aiming to exceed the coding efficiency of the latest international standard, HEVC. In March 2019, the 68th AVS Conference finalized the AVS3-P2 baseline, which offers approximately 30% bitrate savings over the HEVC standard. Currently, there is one reference software, called the High Performance Model (HPM), which is maintained by the AVS group and demonstrates a reference implementation of the AVS3 standard. Summary of the Invention

[0005] This disclosure provides example techniques for improving affine merge candidate derivation for affine motion prediction modes in a video encoding or decoding process.

[0006] In a first aspect of the present disclosure, a method for video decoding is provided. The method may include obtaining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks of a current block based on at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from one side of the current block, and the number is a positive integer. The method may further include obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more motion vector candidates.

[0007] In a second aspect of the present disclosure, a video encoding method is provided, which may include determining one or more motion vector candidates from a plurality of non-adjacent neighboring blocks of a current block based on at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from one side of the current block, and the number may be a positive integer. Further, the method may include obtaining one or more CPMVs for the current block based on the one or more motion vector candidates.

[0008] In a third aspect of the present disclosure, there is provided an apparatus for video decoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors further configured, upon execution of the instructions, to perform a method according to the first aspect.

[0009] In a fourth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors further configured to, upon execution of the instructions, perform a method according to the second aspect.

[0010] In a fifth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform a method according to the first aspect.

[0011] In a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform a method according to the second aspect. [Brief explanation of the drawings]

[0012] Embodiments of the present disclosure will be more particularly described with reference to specific embodiments illustrated in the accompanying drawings, which are to be considered illustrative of some embodiments only and not limiting in scope, and which will be used to describe those embodiments with additional specificity and detail.

[0013] [Figure 1] FIG. 2 is a block diagram of an encoder according to some embodiments of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to some embodiments of the present disclosure. [Figure 3A] FIG. 1 illustrates a block partition of a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3B] FIG. 1 illustrates a block partition of a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3C] FIG. 1 illustrates a block partition of a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3D] FIG. 1 illustrates a block partition of a multi-type tree structure according to some embodiments of the present disclosure. [Figure 3E] FIG. 1 illustrates a block partition of a multi-type tree structure according to some embodiments of the present disclosure. [Figure 4A]1 illustrates a four-parameter affine model according to some embodiments of the present disclosure. [Figure 4B] 1 illustrates a four-parameter affine model according to some embodiments of the present disclosure. [Figure 5] 1 illustrates a six-parameter affine model according to some embodiments of the present disclosure. [Figure 6] 10 illustrates adjacent neighboring blocks of inherited affine merge candidates, according to some embodiments of the present disclosure. [Figure 7] 10 illustrates adjacent neighboring blocks of constructed affine merge candidates, according to some embodiments of the present disclosure. [Figure 8] 10 illustrates non-adjacent neighboring blocks of inherited affine merge candidates, according to some embodiments of the present disclosure. [Figure 9] FIG. 10 illustrates the derivation of constructed affine merge candidates using non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 10] 10 illustrates vertical scanning of non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 11] 1 illustrates parallel scanning of non-adjacent neighboring blocks according to some embodiments of the present disclosure. [Figure 12] 1 illustrates a combination of vertical and parallel scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 13A] 1 illustrates neighboring blocks having the same size as the current block, according to some embodiments of the present disclosure. [Figure 13B] 10 illustrates neighboring blocks having a different size than the current block, according to some embodiments of the present disclosure. [Figure 14A] 10 illustrates an example in which the block to the lower left or upper right of the bottom-most or right-most block at the previous distance is used as the bottom-most or right-most block at the current distance, according to some embodiments of the present disclosure. [Figure 14B] 10 illustrates an example in which the block to the left or top of the bottom-most or right-most block at the previous distance is used as the bottom-most or right-most block at the current distance, according to some embodiments of the present disclosure. [Figure 15A] 10 illustrates scan positions at the bottom-left and top-right positions used for the top non-adjacent neighboring block and the left non-adjacent neighboring block, according to some embodiments of the present disclosure. [Figure 15B] 10 illustrates a scan position at the bottom right position used for both the top non-adjacent neighboring block and the left non-adjacent neighboring block, according to some embodiments of the present disclosure. [Figure 15C] 10 illustrates a scan position at the bottom left position used for both the upper non-adjacent neighboring block and the left non-adjacent neighboring block, according to some embodiments of the present disclosure. [Figure 15D] 10 illustrates a scan position at the top right position used for both the upper non-adjacent neighboring block and the left non-adjacent neighboring block, according to some embodiments of the present disclosure. [Figure 16] 1 illustrates a simplified scanning process for deriving constructed merge candidates, according to some embodiments of the present disclosure. [Figure 17A] 10 illustrates spatial neighboring blocks from which inherited affine merge candidates are derived, according to some embodiments of the present disclosure. [Figure 17B] 10 illustrates spatial neighboring blocks from which constructed affine merge candidates are derived, according to some embodiments of the present disclosure. [Figure 18] FIG. 1 illustrates a computing environment coupled to a user interface, according to some embodiments of the present disclosure. [Figure 19] 1 is a flowchart illustrating a method for video encoding and decoding, according to some embodiments of the present disclosure. [Figure 20] FIG. 1 is a block diagram illustrating a system for encoding and decoding video blocks in accordance with some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to provide a thorough understanding of the subject matter described herein. However, it will be apparent to those skilled in the art that various alternative embodiments may be used. For example, it will be apparent to those skilled in the art that the subject matter described herein may be implemented in multiple types of electronic devices having digital video capabilities.

[0015] Throughout this specification, references to "one embodiment," "one embodiment," "one example," "some embodiments," "some examples," or similar terms mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. A feature, structure, element, or characteristic described in connection with one or some embodiments may also be applicable to other embodiments, unless expressly stated otherwise.

[0016] Throughout this disclosure, all terms such as "first," "second," "third," etc. are used solely as nomenclature to refer to related elements, e.g., devices, components, compositions, steps, etc., without implying any spatial or chronological order, unless otherwise specified. For example, "first device" and "second device" may refer to two separately formed devices or two parts, components, or operating states of the same device, which may be arbitrarily named.

[0017] The terms "module," "sub-module," "circuit," "sub-circuit," "circuit system," "sub-circuit system," "unit," or "sub-unit" may include memory (shared, dedicated, or group) that stores code or instructions executable by one or more processors. A module may include one or more circuits, with or without stored code or instructions. A module or circuit may include one or more components directly or indirectly connected. These components may or may not be physically attached to each other or located adjacent to each other.

[0018] As used herein, the terms "if" or "if" may be understood to mean "when" or "depending on," depending on the context. When these terms appear in the claims, they may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include: i) performing function or action X' if or if condition X exists; and ii) performing function or action Y' if or if condition Y exists. The method may be implemented by both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be realized in multiple executions of the method at different times.

[0019] A unit or module may be implemented entirely in software, entirely in hardware, or a combination of hardware and software. In an entirely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked to perform a particular function.

[0020] 20 is a block diagram illustrating an exemplary system 10 for concurrently encoding and decoding video blocks according to some embodiments of the present disclosure. As shown in FIG. 20, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0021] In some embodiments, destination device 14 may receive the encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to destination device 14. In one example, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.

[0022] In some other embodiments, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or other intermediate storage device that may store the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network-attached storage (NAS) device, or a local disk drive. Destination device 14 may access the encoded video data over any standard data connection, including a wireless channel (e.g., a Wi-Fi (Wireless Fidelity) connection), a wired connection (e.g., a DSL (Digital Subscriber Line), a cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0023] 20 , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include a video capture device such as a video camera, a video archive containing previously captured video, a video feed interface receiving video from a video content provider, and / or a source such as a computer graphics system generating computer graphics data as source video, or a combination of these sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, the embodiments described herein may be applicable to video encoding and decoding generally, and may also be applicable to wireless and / or wired applications.

[0024] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access by destination device 14 or another device for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0025] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem to receive encoded video data over link 16. The encoded video data communicated over link 16 or provided to storage device 32 may include various syntax elements generated by video encoder 20 that are used by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communications medium, stored on a storage medium, or stored on a file server.

[0026] In some embodiments, destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0027] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as, for example, VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of those standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is contemplated that video encoder 20 of source device 12 may generally be configured to encode video data in accordance with any of these current or future standards. Similarly, it is contemplated that video decoder 30 of destination device 14 may generally be configured to decode video data in accordance with any of these current or future standards.

[0028] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If implemented partially in software, an electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, or either may be integrated within the respective device as part of a combined encoder / decoder (CODEC).

[0029] Like HEVC, VVC is built on a block-based hybrid video encoding / decoding framework. FIG. 1 is a block diagram illustrating a block-based video encoder according to some embodiments of the present disclosure. In encoder 100, an input video signal is processed by blocks called coding units (CUs). Encoder 100 may also be a video encoder 20, as shown in FIG. 20. In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which divides blocks based only on a quadtree, in VVC, a coding tree unit (CTU) is divided into CUs based on a quadtree / binary / ternary tree to adapt to various local characteristics. Furthermore, the concept of multiple division unit types in HEVC has been eliminated. That is, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as a basic unit for both prediction and transformation without further division. In the multi-type tree structure, a CTU is first divided by a quadtree structure. Each quadtree leaf node can then be further divided by binary and ternary tree structures.

[0030] 3A-3E are schematic diagrams illustrating a multi-type tree partition mode according to some embodiments of the present disclosure. Each of Figures 3A-3E illustrates five partition types, including 4-way partition (Figure 3A), vertical 2-way partition (Figure 3B), horizontal 2-way partition (Figure 3C), vertically expanded 3-way partition (Figure 3D), and horizontally expanded 3-way partition (Figure 3E).

[0031] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples (called reference samples) of previously coded neighboring blocks within the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion-compensated prediction") predicts the current video block using reconstructed pixels from previously coded video pictures. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. If multiple reference pictures are supported, a reference picture index is additionally transmitted and is used to identify which reference picture in the reference picture store the temporal prediction signal is from.

[0032] After spatial prediction and / or temporal prediction, an intra / inter mode decision circuit 121 of the encoder 100 selects an optimal prediction mode, for example, based on a rate-distortion optimization method. The block predictor 120 is then subtracted from the current video block, and the obtained prediction residual is decorrelated using the transform circuit 102 and the quantization circuit 104. The obtained quantized residual coefficients are inverse quantized by the inverse quantization circuit 116 and inverse transformed by the inverse transform circuit 118 to form a reconstructed residual, which is then added to the prediction block to form a reconstructed signal for the CU. Furthermore, in-loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), can be applied to the reconstructed CU before it is placed in a reference picture store in a picture buffer 117 and used to encode future video blocks. To form the output video bitstream 114, Forecast related information 110( Coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients etc.) are all sent to the entropy coding unit 106 where they are further compressed and packed to form a bitstream.

[0033] For example, deblocking filters are available in the latest versions of AVC, HEVC, and VVC. HEVC defines an additional in-loop filter called SAO to further improve coding efficiency. For the latest version of the VVC standard, another in-loop filter called ALF is being actively researched and is likely to be included in the final standard.

[0034] These in-loop filter operations are selectable. Performing these operations helps improve coding efficiency and visual quality. These operations may be turned off as a decision made by the encoder 100 to reduce computational complexity.

[0035] Note that when these filter options are turned on by the encoder 100, intra prediction is typically based on unfiltered reconstructed pixels, whereas inter prediction is based on filtered reconstructed pixels.

[0036] FIG. 2 is a block diagram illustrating a block-based video decoder 200 that can be used in combination with many video encoding and decoding standards. The decoder 200 is similar to the reconstruction-related parts present in the encoder 100 of FIG. 1. The block-based video decoder 200 may be a video decoder 30, as shown in FIG. 20. In the decoder 200, an input video bitstream 201 is first decoded by entropy decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by inverse quantization 204 and inverse transform 206 to obtain a reconstructed prediction residual. A block predictor mechanism implemented in an intra / inter mode selection 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by adding the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block predictor mechanism using an adder 214.

[0037] The reconstructed blocks may further pass through an in-loop filter 209 before being stored in a picture buffer 213, which acts as a reference picture store. The reconstructed video in the picture buffer 213 may be used to predict future video blocks as well as transmitted to drive a display device. When the in-loop filter 209 is turned on, it performs a filtering operation on these reconstructed pixels to derive the final reconstructed video output 222.

[0038] In the current VVC and AVS3 standards, the motion information of the current coding block is copied from spatial or temporal neighboring blocks identified by merge candidate indexes or obtained by explicit signaling of motion estimation. The focus of this disclosure is to improve the accuracy of motion vectors in affine merge mode by improving the method for deriving affine merge candidates. To facilitate the explanation of this disclosure, the existing affine merge mode design in the VVC standard is used as an example to explain the proposed idea. It should be noted that although the existing affine mode design in the VVC standard is used as an example throughout this disclosure, those skilled in the art will recognize that the proposed technology can also be applied to different designs of affine motion prediction modes or other coding tools with the same or similar design concepts.

[0039] Affine Mode In HEVC, only the translational motion model is applied to motion compensated prediction. In the real world, various motions exist, such as zoom-in / zoom-out, rotation, viewpoint movement, and other irregular motions. In VVC and AVS3, affine motion compensated prediction is applied by signaling one flag for each inter-coded block to indicate whether the translational motion model or the affine motion model is applied to the inter prediction. In the current VVC and AVS3 designs, two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode, are supported for one affine-coded block.

[0040] The four-parameter affine model has two parameters for horizontal and vertical translation, one for zoom motion, and one for rotation in both directions. In this model, the horizontal zoom parameter is equal to the vertical zoom parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. To better align the motion vectors and affine parameters, these affine parameters are derived from two MVs (also called control point motion vectors (CPMVs)) located at the top-left and top-right corners of the current block. As shown in Figures 4A and 4B, the affine motion field of a block is described by two CPMVs (v0, v1). Based on the control point motion, the motion field (v x ,v y ) is written as follows:

number

[0041] The six-parameter affine mode has two parameters used for horizontal and vertical translation, two parameters used for horizontal zoom and rotation, and two parameters used for vertical zoom and rotation. The six-parameter affine motion model is coded with three CPMVs. As shown in Figure 5, the three control points of one six-parameter affine block are located at the top-left, top-right, and bottom-left corners of the block. The top-left control point's motion is related to translation, the top-right control point's motion is related to horizontal rotation and zoom, and the bottom-left control point's motion is related to vertical rotation and zoom. Compared with the four-parameter affine motion model, the six-parameter horizontal rotation and zoom motion may not be the same as the vertical motion. Assuming (v0, v1, v2) are the CPMVs of the top-left, top-right, and bottom-left corners of the current block in Figure 5, the motion vectors (v x ,v y ) is derived using the three MVs of the control point as follows:

number

[0042] Affine Merge Mode In affine merge mode, the CPMV of the current block is not explicitly signaled but is derived from neighboring blocks. Specifically, in this mode, motion information of spatially neighboring blocks is used to generate the CPMV of the current block. The candidate list for affine merge mode has a finite size. For example, in the current VVC design, there can be up to five candidates. The encoder may evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. The affine merge candidate can be determined in three ways. In the first method, the affine merge candidate may be inherited from a neighboring affine coded block. In the second method, the affine merge candidate is constructed from the translational MV from the neighboring block. In the third method, the zero MV is used as the affine merge candidate.

[0043] There can be up to two candidates for the inheritance method, which, if available, are taken from a neighboring block located to the bottom left of the current block (e.g., as shown in FIG. 6, the scanning order is A0 to A1) and a neighboring block located to the top right of the current block (e.g., as shown in FIG. 6, the scanning order is B0 to B2).

[0044] In the construction method, the candidates are combinations of translational MVs of neighboring blocks and may be generated in the following two steps.

[0045] In step 1, four translational MVs including MV1, MV2, MV3 and MV4 are obtained from the available neighboring blocks. For MV1, it is the MV from one of the three neighboring blocks close to the top left corner of the current block. As shown in Figure 7, the scanning order is B2, B3, A2. For MV2, it is the MV from one of the two neighboring blocks close to the top right corner of the current block. As shown in Figure 7, the scanning order is B1, B0. For MV3, it is the MV from one of the two neighboring blocks close to the bottom left corner of the current block. As shown in Figure 7, the scanning order is A1, A0. For MV4, it is the MV from the temporally aligned block of the neighboring block that is close to the bottom right corner of the current block. As shown in the figure, the neighboring block is T.

[0046] In step 2, combinations are derived based on the four translational MVs from step 1. Combination 1 is MV1, MV2, and MV3. Combination 2 is MV1, MV2, and MV4. Combination 3 is MV1, MV3, and MV4. Combination 4 is MV2, MV3, and MV4. Combination 5 is MV1 and MV2. Combination 6 is MV1, MV3.

[0047] If the merge candidate list is not full after being filled with inheritance and construction candidates, insert zero MV at the end of the list.

[0048] In current video standards VVC and AVS, only adjacent neighboring blocks are used to derive affine merge candidates for the current block, as shown in Figures 6 and 7 for inheritance and construction candidates, respectively. To increase the diversity of merge candidates and further explore spatial correlations, the coverage of neighboring blocks is extended directly from adjacent to non-adjacent regions.

[0049] In this disclosure, the candidate derivation process for the affine merge mode is extended by using not only adjacent neighboring blocks but also non-adjacent neighboring blocks. The detailed method may be summarized into three aspects, including pruning of affine merge candidates, a non-adjacent neighboring block-based derivation process for inherited affine merge candidates, and a non-adjacent neighboring block-based derivation process for constructed affine merge candidates.

[0050] Pruning Affine Merge Candidates Because affine merge candidate lists in common video coding standards usually have a finite size, candidate pruning is an essential process for removing redundant candidates. This pruning process is necessary for both inherited and constructed affine merge candidates. As explained in the introduction section, the CPMV of the current block is directly used for affine motion compensation. Instead, the CPMV needs to be transformed into a translational MV at the position of each sub-block within the current block. The transformation process is performed according to a general affine model, as shown below.

number

[0051] where (a, b) are delta translation parameters, (c, d) are horizontal delta zoom and rotation parameters, (e, f) are vertical delta zoom and rotation parameters, (x, y) are the horizontal and vertical distances (e.g., coordinates (x, y) shown in FIG. 5) between the pivot position (e.g., center or top-left corner) of the sub-block and the top-left corner of the current block, and (v x ,v y ) is the target translation MV of the sub-block.

[0052] For a six-parameter affine model, three CPMVs are available, called V0, V1, and V2. Then, the six model parameters, a, b, c, d, e, and f, can be calculated as follows:

number

[0053] For a four-parameter affine model, if the top-left and top-right corner CPMVs, called V0 and V1, are available, the six parameters a, b, c, d, e, and f can be calculated as follows:

number

[0054] For a four-parameter affine model, if the top-left corner CPMV and bottom-left corner CPMV, called V0 and V2, are available, the six parameters a, b, c, d, e, and f can be calculated as follows:

number

[0055] In the above equations (4), (5), and (6), w and h represent the width and height of the current block, respectively.

[0056] When two merge candidate sets of CPMV are compared for redundancy check, it is proposed to check the similarity of six affine model parameters. Therefore, the candidate pruning process can be performed in two steps.

[0057] In step 1, given two candidate sets of CPMVs, affine model parameters corresponding to each candidate set are derived. More specifically, the two candidate sets of CPMVs may be represented by two sets of affine model parameters, for example, (a1, b1, c1, d1, e1, f1) and (a2, b2, c2, d2, e2, f2).

[0058] In step 2, a similarity check is performed between the two sets of affine model parameters based on one or more predetermined thresholds. In one embodiment, if the absolute values of (a1-a2), (b1-b2), (c1-c2), (d1-d2), (e1-e2) and (f1-f2) are all less than a positive threshold, such as value 1, then the two candidates are considered similar and one of them can be pruned / removed from the merge candidate list.

[0059] In some embodiments, the division or right shift operation in step 1 may be removed to simplify the calculations in the CPMV pruning process.

[0060] JPEG0007720482000007.jpg21165

number

[0061] When only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters that depend on the width or height of the current block. In this case, the model parameters may be transformed to take into account the influence of the width and height. For example, in the case of Equation (5), the approximate model parameters of c', d', e', and f' may be calculated based on the following Equation (8). In the case of Equation (6), the approximate model parameters of c', d', e', and f' may be calculated based on the following Equation (9).

number

number

[0062] JPEG0007720482000011.jpg25166

[0063] In step 2 above, a threshold is required to evaluate the similarity between two candidate sets of CPMV. There may be multiple ways to define the threshold. In one embodiment, a threshold may be defined for each comparable parameter. Table 1 is an example of this embodiment, showing the threshold defined for each comparable model parameter. In another embodiment, the threshold may be defined taking into account the size of the current coding block. Table 2 is an example of this embodiment, showing the threshold defined by the size of the current coding block.

[0064] [Table 1]

[0065] [Table 2]

[0066] In another embodiment, the thresholds may be defined taking into account the weight or height of the current block. Tables 3 and 4 are examples of this embodiment. Table 3 shows thresholds defined by the width of the current coding block, and Table 4 shows thresholds defined by the height of the current coding block.

[0067] [Table 3]

[0068] [Table 4]

[0069] In another embodiment, the threshold may be defined as a set of fixed values. In another embodiment, the threshold may be defined by any combination of the above embodiments. In one example, the threshold is determined by various parameters, such as the weight of the current block and Width andThe threshold may be defined by taking into account the height of the current coding block. Table 5 shows an example of this embodiment, showing the threshold defined by the height of the current coding block. Note that in any of the above proposed embodiments, the comparable parameter may represent any parameter defined by any of Equations (4) to (9), as needed.

[0070] [Table 5]

[0071] Advantages of using the transformed affine model parameters for candidate redundancy checking include creating a unified similarity checking process for candidates with different affine model types, for example, one merge candidate may use a 6-parameter affine model with 3 CPMVs and another candidate may use a 4-parameter affine model with 2 CPMVs; considering the different influences of each CPMV in the merge candidate when deriving the target MV for each sub-block; and providing the importance of the similarity of the two affine merge candidates relative to the width and height of the current block.

[0072] Inherited non-adjacent neighbor block-based derivation process of affine merge candidates For inherited merge candidates, the non-adjacent neighborhood-based derivation process may be performed in three steps: Step 1 is used to scan candidates; Step 2 is used to project CPMVs; and Step 3 is used to prune candidates.

[0073] In step 1, scan and select non-adjacent neighboring blocks in the following manner.

[0074] Scanning Area and Distance In some embodiments, the non-adjacent neighboring blocks may be scanned from the left and top regions of the current coding block, and the scanning distance may be defined as the number of coding blocks from the scanning position to the left or top edge of the current coding block.

[0075] As shown in Figure 8, multiple lines of non-adjacent neighboring blocks may be scanned either to the left or above the current coding block. The distances shown in Figure 8 represent the number of coding blocks from each candidate position to the left or top edge of the current block. For example, an area with a "distance of 2" on the left edge of the current block indicates that the candidate neighboring block located in this area is two blocks away from the current block. Similar instructions may be applied to other scanning areas with different distances.

[0076] In one or more embodiments, as shown in Figure 13A, the non-adjacent neighboring blocks at each distance may have the same block size as the current coding block. As shown in Figure 13A, the non-adjacent neighboring block 1301 on the left side and the non-adjacent neighboring block 1302 on the top side have the same size as the current block 1303. In some embodiments, as shown in Figure 13B, the non-adjacent neighboring blocks at each distance may have a different block size than the current coding block. Neighboring block 1304 is an adjacent neighboring block to the current block 1303. As shown in Figure 13B, the non-adjacent neighboring block 1305 on the left side and the non-adjacent neighboring block 1306 on the top side are adjacent to the current block 1307. different The neighboring blocks 1308 are adjacent neighboring blocks to the current block 1307.

[0077] If the non-adjacent neighboring blocks at each distance have the same block size as the current coding block, the block size value is adaptively changed according to the division granularity for different regions in the image. If the non-adjacent neighboring blocks at each distance have a block size different from the current coding block, the block size value may be pre-defined as a constant value such as 4x4, 8x8, or 16x16. The 4x4 non-adjacent motion field shown in Figures 10 and 12 is an example of this case, and the motion field may be considered as a special case of a sub-block, but is not limited to this.

[0078] Similarly, the non-adjacent coding blocks shown in Figure 11 may also have different sizes. In one embodiment, the non-adjacent coding blocks may have the same size as the current coding block, which may be adaptively changed. In another embodiment, the non-adjacent coding blocks may have a predefined size, including a fixed value such as 4x4, 8x8, or 16x16.

[0079] Based on the defined scan distance, the total size of the scan area either to the left or above the current encoding clock may be determined by a configurable distance value. In one or more embodiments, the maximum scan distances for the left and top edges may use the same value or different values. A-Figure 13B shows an example where the maximum distances for both the left and top edges share the same value of 2. The maximum scanning distance value may be determined by the encoder side and signaled in the bitstream. Alternatively, the maximum scanning distance value may be predefined as a fixed value, such as the value 2 or 4. If the maximum scanning distance is predefined as the value 4, it indicates that the scanning process will end on a first-come, first-served basis when the candidate list is full or when all non-adjacent neighboring blocks with a distance of at most 4 have been scanned.

[0080] In one or more embodiments, within each scan area at a particular distance, the starting and ending neighboring blocks may be positionally related.

[0081] In some embodiments, for the left side scanning area, the starting neighboring block may be the lower left neighboring block of the starting neighboring block of the neighboring scanning area with a smaller distance. For example, as shown in FIG. 8, the starting neighboring block of the "distance 2" scanning area on the left side of the current block is the lower left neighboring block of the starting neighboring block of the "distance 1" scanning area. The ending neighboring block is the lower left neighboring block of the ending neighboring block of the upper scanning area with a smaller distance. above For example, as shown in FIG. 8, the end neighboring block of the "distance 2" scanning area on the left side of the current block is the left neighboring block of the end neighboring block of the "distance 1" scanning area above the current block. above It is a side neighboring block.

[0082] Similarly, for a top-side scanning area, the starting neighboring block may be the neighboring upper right block of the starting neighboring block of the adjacent scanning area with a smaller distance, and the ending neighboring block may be the neighboring upper left block of the ending neighboring block of the adjacent scanning area with a smaller distance.

[0083] Traversal Order When neighboring blocks are scanned in a non-adjacent region, the selection of neighboring blocks to be scanned may be determined according to a certain order and / or rules.

[0084] In some embodiments, the left region may be scanned first, followed by the upper region, as shown in Figure 8. The left three-line non-adjacent region (e.g., from distance 1 to distance 3) may be scanned first, followed by the upper three-line non-adjacent region of the current block.

[0085] In some embodiments, the left region and the upper region may be scanned alternately, for example, first scanning the left scan region with "Distance 1" and then scanning the upper region with "Distance 1" as shown in FIG.

[0086] For scanning areas located on the same side (such as the left area or the upper area), the scanning order is from the area with the smallest distance to the area with the largest distance. This order may be flexibly combined with other embodiments of the scanning order. For example, the left area and the upper area may be scanned alternately, and the order for areas on the same side is planned to be from the smallest distance to the largest distance.

[0087] Within each scan region at a particular distance, a scan order may be defined. In one embodiment, for the left scan region, scanning may start from the bottom neighbor block to the top neighbor block. For the top scan region, scanning may start from the right block to the left block.

[0088] Scanning End For the inherited merge candidates, neighboring blocks coded in affine mode are defined as suitable candidates. In some embodiments, the scanning process may be performed in an alternating manner. For example, scanning performed in a particular region at a certain distance may be stopped once the first X suitable candidates have been identified, where X is a predetermined positive value. For example, as shown in FIG. 8, scanning in the left scanning region with a distance of 1 may be stopped once one or more previous suitable candidates have been identified. Then, the next iteration of the scanning process begins by targeting another scanning region controlled by a predetermined scanning order / rule.

[0089] In some embodiments, the scanning process may be performed continuously, for example, scanning performed in a particular area at a particular distance may stop when all covered neighboring blocks have been scanned and no more suitable candidates are identified, or when a maximum allowed number of candidates has been reached.

[0090] During the candidate scanning process, the non-adjacent neighboring blocks of each candidate are determined and scanned by the proposed scanning method. For easier implementation, the non-adjacent neighboring blocks of each candidate can be indicated or located by a specific scanning position. After the specific scanning area and distance are determined by the proposed method, the scanning position can be determined accordingly based on the following method:

[0091] In one method, as shown in FIG. 15A, the bottom left and top right positions are used for the top non-adjacent neighboring block and the left non-adjacent neighboring block, respectively.

[0092] Alternatively, as shown in FIG. 15B, the bottom right position is used for both the top non-adjacent neighboring block and the left non-adjacent neighboring block.

[0093] Alternatively, the bottom left position is used for both the top non-adjacent neighboring block and the left non-adjacent neighboring block, as shown in FIG. 15C.

[0094] Alternatively, the top right position is used for both the top non-adjacent neighboring block and the left non-adjacent neighboring block, as shown in FIG. 15D.

[0095] For convenience of explanation, in Figures 15A to 15D, each non-adjacent neighboring block is assumed to have the same block size as the current block. Without loss of generality, this explanation can be easily extended to non-adjacent neighboring blocks with different block sizes.

[0096] Furthermore, step 2 may utilize the same CPMV projection process as used in the current AVS and VVC standards, in which, assuming that the current block and the selected neighboring blocks share the same affine model, two or three corner pixel coordinates (e.g., if the current block uses a four-parameter model, two coordinates (the position of the top-left pixel / sample and the position of the top-right pixel / sample) are used, and if the current block uses a six-parameter model, three coordinates (the position of the top-left pixel / sample, the position of the top-right pixel / sample, and the position of the bottom-left pixel / sample) are used) are substituted into equation (1) or (2) to generate two or three CPMVs, depending on whether the neighboring blocks are coded with a four-parameter affine model or a six-parameter affine model.

[0097] In step 3, a similarity check may be performed between any suitable candidate identified in step 1 and transformed in step 2 and all existing candidates already in the merge candidate list. Details of the similarity check have already been discussed in the section on pruning affine merge candidates above. If a new suitable candidate is found to be similar to an existing candidate in the candidate list, then this new suitable candidate is removed / pruned.

[0098] Non-adjacent neighborhood-based derivation process for constructed affine merge candidates When deriving inherited merge candidates, one neighboring block is identified at a time, and this single neighboring block needs to be coded in affine mode and may contain two or three CPMVs. When deriving constructed merge candidates, two or three neighboring blocks may be identified at a time, and each identified neighboring block does not need to be coded in affine mode, and only one translational MV is retrieved from this block.

[0099] JPEG0007720482000017.jpg51165

[0100] For the constructed merge candidates, the non-adjacent neighbor block-based derivation process may be performed in five steps. The non-adjacent neighbor block-based derivation process may be performed in a device such as an encoder or decoder in five steps: Step 1 is candidate scanning; Step 2 is affine model determination; Step 3 is CPMV projection; Step 4 is candidate generation; and Step 5 is candidate pruning. In Step 1, non-adjacent neighbor blocks may be scanned and selected by the following method:

[0101] Scanning area and distance In some embodiments, to maintain a rectangular coding block, the scanning process is performed for only two non-adjacent neighboring blocks, and the third non-adjacent neighboring block may depend on the horizontal and vertical positions of the first and second non-adjacent neighboring blocks.

[0102] In some embodiments, as shown in Figure 9, the scanning process is performed only for the positions of B and C. The position of A may be uniquely determined by the horizontal position of C and the vertical position of B. In this case, the scanning area and distance may be defined according to a specific scanning direction.

[0103] In some embodiments, the scanning direction may be perpendicular to an edge of the current block. One example is shown in Figure 10, where the scanning area is defined as one line of continuous motion field to the left or above the current block. The scanning distance is defined as the number of motion field segments from the scanning position to the edge of the current block. Note that the size of the motion field may depend on the maximum granularity of the applicable video coding standard. In the example shown in Figure 10, the size of the motion field is assumed to be set to 4x4 to match the current VVC standard.

[0104] In some embodiments, the scan direction may be parallel to the edges of the current block. One example is shown in Figure 11, where the scan region is defined as one line of contiguous coded blocks to the left or above the current block.

[0105] In some embodiments, the scanning direction may be a combination of perpendicular scanning and parallel scanning relative to the edge of the current block. One example is shown in FIG. 12. As shown in FIG. 12, the scanning direction may be a combination of parallel and diagonal. The scanning at position B starts from the left to the right, then diagonally moves toward the left block and the top block. The scanning at position B repeats as shown in FIG. 12. Similarly, the scanning at position C starts from the top to the bottom, then diagonally moves toward the left block and the top block. The scanning at position C repeats as shown in FIG. 12.

[0106] Traversal Order In some embodiments, the scanning order may be defined as an order from a position with a small distance to the current coding block to a position with a large distance to the current coding block, which may be applied in the case of vertical scanning.

[0107] In some embodiments, the scan order may be defined as a fixed pattern. This fixed pattern scan order may be used for candidate locations with similar distances. One example is parallel scanning. In one example, the scan order may be defined as an order from top to bottom for the left scan area, or from left to right for the upper scan area, as shown in FIG. 11 .

[0108] In the case of a combined scanning method, the scanning order may be a combination of fixed pattern and distance dependent, as in the example shown in FIG.

[0109] Scanning End For the constructed merge candidates, only translational MVs are needed, so there is no need to affine-code the appropriate candidates.

[0110] Depending on the number of candidates desired, the scanning process may terminate when the first X suitable candidates have been identified, where X is a positive value.

[0111] As shown in FIG. 9, three corners named A, B, and C are required to form a virtual coding block. For ease of implementation, only the scanning process of step 1 can be performed to identify non-adjacent neighboring blocks located at corners B and C, and the coordinate of A can be accurately determined by obtaining the horizontal coordinate of C and the vertical coordinate of B. In this way, the formed virtual coding block is restricted to a rectangle. If either point B or point C is unavailable, for example, if it is out of bounds, or if the motion information of the non-adjacent neighboring blocks corresponding to B or C is unavailable, the horizontal or vertical coordinate of C may be defined as the horizontal or vertical coordinate of the top-left point of the current block, respectively.

[0112] In another embodiment, first, when corner B and / or corner C are determined from the scanning process of step 1, non-adjacent neighboring blocks located at corner B and / or corner C may be identified accordingly. Then, the positions of corner B and / or corner C may be reset to a pivot point within the corresponding non-adjacent neighboring block, such as the center of mass of each non-adjacent neighboring block. For example, the center of mass may be defined as the geometric center of each neighboring block.

[0113] For purposes of uniformity, the methods for defining the scan area and distance, scan order, and scan termination proposed for deriving the inherited merge candidate may be fully or partially reused for deriving the constructed merge candidate. In one or more embodiments, the same methods defined for the inherited merge candidate traversal, including but not limited to the scan area and distance, scan order, and scan termination, may be fully reused for the constructed merge candidate traversal.

[0114] In some embodiments, the same method defined for inherited merge candidate scanning may be partially reused for constructed merge candidate scanning. Figure 16 shows an example of this case. In Figure 16, the block size of each non-adjacent neighboring block is the same as the current block and is similarly defined as inherited candidate scanning, but the overall process is a simplified version because scanning at each distance is limited to only one block.

[0115] Figures 17A and 17B show another example of this case, where both the inherited non-adjacent merge candidates and the constructed non-adjacent merge candidates are defined with the same block size as the current coding block, but the scan order, scan area, and scan end condition may be defined differently.

[0116] In FIG. 17A, the maximum distance of non-adjacent neighboring blocks on the left edge is four coding blocks, while the maximum distance of non-adjacent neighboring blocks on the top edge is five coding blocks. The scanning direction at each distance is bottom-to-top on the left edge and right-to-left on the top edge. In FIG. 17B, the maximum distance of non-adjacent neighboring blocks is four on both the left and top edges. Since there is only one block at each distance, scanning at certain distances is unavailable. In FIG. 17A, the scanning operation within each distance can be terminated when M suitable candidates are identified. The value of M may be a predefined fixed value, such as the value 1 or any other positive integer, a signal value determined by the encoder, or a value configurable by the encoder or decoder. In one embodiment, the value of M may be the same as the size of the merge candidate list.

[0117] 17A and 17B, the scanning operation at different distances may be terminated if N suitable candidates are identified. The value of N may be a predefined fixed value, such as the value 1 or any other positive integer, a signal value determined by the encoder, or a value configurable by the encoder or decoder. In one embodiment, the value of N may be the same as the size of the merge candidate list. In another embodiment, the value of N may be the same as the value of M.

[0118] In both Figures 17A and 17B, non-adjacent spatial neighboring blocks at distances closer to the current block may be prioritized, which indicates that non-adjacent spatial neighboring blocks at distance i are scanned or checked before neighboring blocks at distance i+1, where i is a non-negative integer representing a particular distance.

[0119] At a particular distance, a maximum of two non-adjacent spatial neighboring blocks are used, which means that, if available, a maximum of one neighboring block from one side (e.g., left and top) of the current block is selected for deriving inheritance or construction candidates. As shown in Figure 17A, the checking order for the left-side neighboring block and the top-side neighboring block is from bottom to top and from right to left, respectively. This rule may also be applied to Figure 17B, with the difference being that at any particular distance, there may be only one choice for each side of the current block.

[0120] For the construction candidate, as shown in Figure 17B, first, the positions of one left non-adjacent spatial neighboring block and one upper non-adjacent spatial neighboring block are independently determined. Then, the position of the upper left neighboring block is determined, so that the left non-adjacent neighboring block and the upper non-adjacent neighboring block can surround a rectangular virtual block. Next, as shown in Figure 9, the motion information of the three non-adjacent neighboring blocks is used to form CPMVs at the upper left (A), upper right (B), and lower left (C) of the virtual block, and finally projected onto the current CU to generate a corresponding construction candidate.

[0121] In step 2, the translational MVs at the candidate locations selected after step 1 can be evaluated to determine an appropriate affine model. For ease of explanation without loss of generality, Figure 9 is used again as an example.

[0122] Factors such as hardware constraints, implementation complexity, various reference indices, etc. may cause the scanning process to terminate before a sufficient number of candidates are identified. For example, motion information for the motion field of one or more of the candidates selected after step 1 may not be available.

[0123] If motion information for all three candidates is available, the corresponding hypothetical coding block represents a six-parameter affine model. If motion information for one of the three candidates is unavailable, the corresponding hypothetical coding block represents a four-parameter affine model. If motion information for one or more of the three candidates is unavailable, the corresponding hypothetical coding block may not represent a valid affine model.

[0124] In some embodiments, if motion information for the upper left corner of the virtual coding block, e.g., corner A in FIG. 9, is unavailable, or if motion information for both the upper right corner, e.g., corner B in FIG. 9, and the lower left corner, e.g., corner C in FIG. 9, is unavailable, the virtual block can be set as invalid and no valid model can be represented, and steps 3 and 4 may be skipped for the current iteration.

[0125] In some embodiments, if either the top right corner, e.g., corner B in FIG. 9, or the bottom left corner, e.g., corner C in FIG. 9, is unavailable, but not both, the virtual block can represent a valid four-parameter affine model.

[0126] In step 3, if the hypothetical coding block can represent a valid affine model, the same projection process used for the inherited merge candidates may be used.

[0127] In one or more embodiments, the same projection process used for the inherited merge candidates may be used, where the four-parameter model represented by the hypothetical coding block from step 2 is projected onto the four-parameter model of the current block, and the six-parameter model represented by the hypothetical coding block from step 2 is projected onto the six-parameter model of the current block.

[0128] In some embodiments, the affine model represented by the hypothetical coding block from step 2 is always projected onto the four-parameter model or the six-parameter model of the current block.

[0129] According to equations (5) and (6), there are two types of four-parameter affine models: Type A is a model in which the CPMVs at the top left and top right corners, called V0 and V1, are available, and Type B is a model in which the CPMVs at the top left and bottom left corners, called V0 and V2, are available.

[0130] In one or more embodiments, the type of the projected four-parameter affine model is the same type as the four-parameter affine model represented by the virtual coding block. For example, if the affine model represented by the virtual coding block from step 2 is a four-parameter affine model of type A or type B, the projected affine model of the current block will also be type A or type B, respectively.

[0131] In some embodiments, the four-parameter affine model represented by the hypothetical coding block from step 2 is always projected onto a four-parameter model of the same type of the current block, e.g., a four-parameter affine model of type A or type B represented by the hypothetical coding block is always projected onto a four-parameter affine model of type A.

[0132] In step 4, based on the projected CPMV after step 3, in one embodiment, the same candidate generation process used in the current VVC or AVS standard may be used. In another embodiment, the temporal motion vectors used in the candidate generation process of the current VVC or AVS standard may not be used in the non-adjacent neighbor block-based derivation method. When a temporal motion vector is not used, it indicates that the generated combination does not include a temporal motion vector.

[0133] In step 5, a similarity check may be performed between the newly generated candidate after step 4 and all existing candidates already present in the merge candidate list. The details of the similarity check have already been discussed in the section on pruning affine merge candidates. If the newly generated candidate is found to be similar to an existing candidate in the candidate list, then this newly generated candidate is removed or pruned.

[0134] Reordering Affine Merge Candidate Lists In one embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order: 1, sub-block-based temporal motion vector prediction (SbTMVP) candidates if available, 2, inherited from adjacent neighboring blocks, 3, inherited from non-adjacent neighboring blocks, 4, constructed from adjacent neighboring blocks, 5, constructed from non-adjacent neighboring blocks, 6, zero MV.

[0135] In another embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order: 1, SbTMVP candidate if available, 2, inherited from adjacent neighboring blocks, 3, constructed from adjacent neighboring blocks, 4, inherited from non-adjacent neighboring blocks, 5, constructed from non-adjacent neighboring blocks, 6, zero MV.

[0136] In another embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order: 1, SbTMVP candidate if available, 2, inherited from adjacent neighbors, 3, constructed from adjacent neighbors, 4, one set of zero MVs, 5, inherited from non-adjacent neighbors, 6, constructed from non-adjacent neighbors, 7, else zero MVs if the list is not yet full.

[0137] In another embodiment, non-adjacent spatial merge candidates may be inserted into the affine merge candidate list by following the order: 1, SbTMVP candidate if available, 2, inherited from adjacent neighbors, 3, inherited from non-adjacent neighbors at a distance less than X, 4, constructed from adjacent neighbors, 5, constructed from non-adjacent neighbors at a distance less than Y, 6, inherited from non-adjacent neighbors at a distance greater than X, 7, constructed from non-adjacent neighbors at a distance greater than Y, 8, zero MV. In this embodiment, the values X and Y may be predefined fixed values, such as the value 2, signal values determined by the encoder, or values configurable in the encoder or decoder. In one embodiment, the value of X may be the same as the value of Y. In another embodiment, X The value of Y may differ from the value of

[0138] 18 illustrates a computing environment (or computing device) 1810 coupled to a user interface 1860. The computing environment 1810 may be part of a data processing server. In some embodiments, the computing device 1810 may perform any of the various methods or processes (e.g., encoding / decoding methods or processes) according to various embodiments of the present disclosure, as described above. The computing environment 1810 may include a processor 1820, a memory 1840, and an I / O interface 1850.

[0139] The processor 1820 typically controls the overall operation of the computing environment 1810, such as operations related to display, data acquisition, data communication, and image processing. The processor 1820 may include one or more processors for executing instructions to perform all or part of the steps of the methods described above. The processor 1820 may also include one or more modules that facilitate interaction between the processor 1820 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.

[0140] Memory 1840 is configured to store various types of data to support the operation of computing environment 1810. Memory 1840 may include predefined software 1842. Examples of such data include instructions for any applications or methods operating on computing environment 1810, video data sets, image data, etc. Memory 1840 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.

[0141] The I / O interface 1850 provides an interface between the processor 1820 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1850 may be coupled to an encoder and a decoder.

[0142] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as those contained in memory 1840, executable by processor 1820 in computing environment 1810 to perform the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0143] A non-transitory computer-readable storage medium stores a plurality of programs that are executed by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the motion prediction method described above.

[0144] In some embodiments, the computing environment 1810 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0145] FIG. 19 is a flowchart illustrating a video encoding method according to one embodiment of the present disclosure.

[0146] In step 1901, the processor 1820 can obtain or determine one or more MV candidates from multiple non-adjacent neighboring blocks to the current block based on at least one scanning distance, where one of the at least one scanning distances can indicate the number of blocks away from one side of the current block, and this number is a positive integer.

[0147] In step 1902, the processor 1820 may obtain one or more CPMVs for the current block based on one or more MV candidates.

[0148] In some embodiments, the plurality of non-adjacent neighboring blocks may include non-adjacent coding blocks such as those shown in Figures 11, 12, 13A, 13B, 14A, 14B, 15A-15D, and 16.

[0149] In some embodiments, the processor 1820 may obtain one or more MV candidates according to a scanning rule, which may include, but are not limited to, affine candidates or regular candidates.

[0150] In some embodiments, the scanning rule may be determined based on at least one scanning area, at least one scanning distance, and at least one scanning direction, i.e., the processor 1820 may obtain one or more MV candidates based on the scanning rule including at least one scanning area, at least one scanning distance, and at least one scanning direction.

[0151] In some embodiments, multiple non-adjacent neighboring blocks located at at least one scanning distance have the same size as the current block, as shown in FIG. 13A, or a different size from the current block, as shown in FIG. 13B.

[0152] In some embodiments, the at least one scanning area may include a first scanning area and a second scanning area, where the first scanning area is determined according to a first maximum scanning distance indicating the maximum number of blocks away from a first side of the current block, and the second scanning area is determined according to a second maximum scanning distance indicating the maximum number of blocks away from a second side of the current block, and further, the first maximum scanning distance may be the same as or different from the second maximum scanning distance. In some embodiments, the first maximum scanning distance or the second maximum scanning distance may be set as a fixed value. For example, the first maximum scanning distance may be 4, and the second maximum scanning distance may be 5.

[0153] For example, as shown in Figure 17A, the first scanning area may be the left side area of the current block, and the first maximum scanning distance is four blocks away from the left side of the current block, and the second scanning area may be the top side area of the current block, and the second maximum scanning distance is five blocks away from the top side of the current block.

[0154] In some embodiments, the at least one scanning direction may include a first scanning direction and a second scanning direction, and the processor 1820 may further scan the first scanning area in the first scanning direction and scan the second scanning area in the second scanning direction.

[0155] In some embodiments, as shown in FIG. 17A, the first scanning direction may be from bottom to top parallel to the first side, and the second scanning direction may be from right to left parallel to the second side.

[0156] In some embodiments, the processor 1820 may stop scanning at least one scanning area to obtain one or more motion vector candidates in response to determining that a termination condition is appropriate.

[0157] In some embodiments, the termination condition may include determining that the number of one or more motion vector candidates has reached a predetermined value, which may include at least one of a positive integer, a signal value determined by the encoder, a value configurable in the encoder or decoder, or a size value of a candidate list including one or more MV candidates.

[0158] In some embodiments, the processor 1820 may scan at least one scanning area at a first scanning distance before scanning at least one scanning area at a second scanning distance, where the second scanning distance is one block greater than the first scanning distance, i.e., non-adjacent spatial neighboring blocks at a distance closer to the current block may be prioritized.

[0159] In some embodiments, the processor 1820 may scan a first scanning area at a first scanning distance to obtain a first MV candidate, and scan a second scanning area at a second scanning distance to obtain a second MV candidate, where the number of first MV candidates and second MV candidates is two or less.

[0160] In some embodiments, the one or more MV candidates may include one or more MV inheritance candidates and one or more MV construction candidates. The processor 1820 may obtain the one or more MV inheritance candidates according to a first scanning rule and obtain the one or more MV construction candidates according to a second scanning rule, where the second scanning rule may be completely or partially the same as the first scanning rule.

[0161] In some embodiments, processor 1820 may determine a second scanning rule based on at least one second scanning area and at least one second scanning distance, where the at least one second scanning area may include a left scanning area and an top scanning area. Furthermore, processor 1820 may scan the left scanning area from right to left in a direction perpendicular to the first edge to obtain a first motion vector construction candidate at the first scanning distance, scan the top scanning area from bottom to top in a direction perpendicular to the second edge to obtain a second motion vector construction candidate at the second scanning distance, obtain a virtual block based on a first candidate position of the first motion vector construction candidate and a second candidate position of the second motion vector construction candidate, determine a third candidate position of the third motion vector construction candidate based on the first candidate position, the second candidate position, and the virtual block, and obtain two or three CPMVs of the current block based on the three CPMVs of the virtual block using the same projection process used to derive the successor candidate.

[0162] In some embodiments, the first candidate position, the second candidate position, and the third candidate position are pivot points within the first MV construction candidate, the second MV construction candidate, and the third MV construction candidate, respectively. The first candidate position and the second candidate position may be positions B and C as shown in Figure 9. The third candidate position may be position A as shown in Figure 9.

[0163] In some embodiments, the processor 1820 may insert one or more MV candidates into the candidate list according to a predetermined order, as described in the section Reordering the Affine Merge Candidate List.

[0164] In some embodiments, the one or more MV candidates may include one or more MV inheritance candidates and one or more MV construction candidates, and the processor 1820 may insert the one or more MV inheritance candidates before the one or more MV construction candidates in the candidate list.

[0165] In some embodiments, the candidate list may include one or more SbTMVP candidates and one or more zero MVs, where the one or more MV candidates are inserted after the one or more SbTMVP candidates and before the zero MVs.

[0166] In some embodiments, the candidate list may further include one or more adjacent MV candidates, including one or more adjacent MV inheritance candidates and one or more adjacent MV construction candidates, where the one or more adjacent MV candidates are from multiple adjacent neighboring blocks adjacent to the current block.

[0167] In some embodiments, the processor 1820 may insert one or more adjacent MV inheritance candidates before one or more adjacent MV construction candidates in the candidate list, and may insert one or more adjacent motion vector construction candidates before one or more motion vector inheritance candidates in the candidate list.

[0168] In some embodiments, the processor 1820 may insert one or more zero MVs in the candidate list between one or more neighboring motion vector construction candidates and one or more motion vector inheritance candidates.

[0169] In some embodiments, the one or more motion vector candidates may include one or more motion vector inheritance candidates and one or more motion vector construction candidates, the one or more motion vector inheritance candidates may include at least one first motion vector inheritance candidate having a scanning distance less than a first threshold and at least one second motion vector inheritance candidate having a scanning distance greater than the first threshold, and the one or more motion vector construction candidates may include at least one first motion vector construction candidate having a scanning distance less than a second threshold and at least one second motion vector construction candidate having a scanning distance greater than the second threshold. In some embodiments, the processor 1820 may insert the at least one first motion vector inheritance candidate before the at least one second motion vector inheritance candidate in the candidate list, and may insert the at least one first motion vector construction candidate before the at least one second motion vector construction candidate in the candidate list.

[0170] In some embodiments, the processor 1820 may insert one or more adjacent motion vector inheritance candidates into the candidate list before at least one first motion vector inheritance candidate, insert one or more adjacent motion vector construction candidates into the candidate list after the at least one first motion vector inheritance candidate and before the at least one first motion vector construction candidate, and insert at least one second motion vector inheritance candidate into the candidate list before the at least one second motion vector construction candidate.

[0171] In some embodiments, the first and second thresholds may be the same as or different from one another, and each of the first and second thresholds may comprise at least one of a positive integer, a signal value determined by an encoder, or a value configurable by an encoder or decoder.

[0172] In some embodiments, an apparatus for video encoding is provided, the apparatus including a processor 1820 and a memory 1840 configured to store instructions executable by the processor, the processor, upon execution of the instructions, being configured to perform a method such as that shown in FIG.

[0173] In some other embodiments, a non-transitory computer-readable storage medium is provided that stores instructions that, when executed by the processor 1820, cause the processor to perform a method such as that shown in FIG.

[0174] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure in accordance with its general principles, including departures from the present disclosure that come within known or customary practice in the art. It is intended that the specification and examples be considered as illustrative only.

[0175] It will be understood that the present disclosure is not strictly limited to the embodiments described above and illustrated by the accompanying drawings, and that various modifications and changes can be made thereto without departing from the scope of the present disclosure.

Claims

1. obtaining one or more candidate motion vectors from a plurality of non-adjacent neighboring blocks of a current block based on at least one scanning distance, wherein one of the at least one scanning distance indicates a number of blocks away from one side of the current block, and the number is a positive integer; obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more motion vector candidates, wherein the one or more motion vector candidates include one or more motion vector inheritance candidates and one or more motion vector construction candidates; inserting the one or more candidate motion vectors into a candidate list according to a predetermined order; the candidate list further includes one or more neighboring motion vector candidates, including one or more neighboring motion vector inheritance candidates and one or more neighboring motion vector construction candidates; the predetermined order comprises a first order; In the first order, inserting the one or more motion vector candidates into the candidate list comprises: inserting the one or more neighboring motion vector inheritance candidates before the one or more neighboring motion vector construction candidates; inserting the one or more neighboring motion vector construction candidates before the one or more motion vector inheritance candidates; and inserting the one or more motion vector inheritance candidates before the one or more motion vector construction candidates.

2. 2. The method of claim 1, wherein the candidate list includes one or more sub-block-based temporal motion vector prediction (SbTMVP) candidates and one or more zero motion vectors (MVs), and the one or more motion vector candidates are inserted after the one or more SbTMVP candidates and before the one or more zero MVs.

3. The method of claim 2 , wherein the one or more neighboring motion vector candidates are from a plurality of adjacent neighboring blocks that are adjacent to the current block.

4. the predetermined order comprises a second order replacing the first order; In the second order, inserting the one or more motion vector candidates into the candidate list comprises: inserting the one or more adjacent motion vector inheritance candidates before the one or more motion vector inheritance candidates; inserting the one or more motion vector inheritance candidates before the one or more neighboring motion vector construction candidates; The method of claim 1 , further comprising inserting the one or more neighboring motion vector construction candidates before the one or more motion vector construction candidates.

5. The method of claim 1 , further comprising inserting one or more zero MVs in the candidate list between the one or more neighboring motion vector construction candidates and the one or more motion vector inheritance candidates.

6. the one or more motion vector succession candidates include at least one first motion vector succession candidate having a scanning distance less than a first threshold and at least one second motion vector succession candidate having a scanning distance greater than the first threshold; the one or more motion vector construction candidates include at least one first motion vector construction candidate having a scanning distance less than a second threshold and at least one second motion vector construction candidate having a scanning distance greater than the second threshold; the predetermined order comprises a third order replacing the first order; In the third order, the step of inserting the one or more motion vector candidates into the candidate list comprises: inserting the at least one first motion vector succession candidate before the at least one second motion vector succession candidate in the candidate list; 2. The method of claim 1, further comprising inserting the at least one first motion vector construction candidate before the at least one second motion vector construction candidate in the candidate list.

7. In the third order, the step of inserting the one or more motion vector candidates into the candidate list comprises: inserting the one or more adjacent motion vector succession candidates in the candidate list before the at least one first motion vector succession candidate; inserting the one or more neighboring motion vector construction candidates after the at least one first motion vector succession candidate and before the at least one first motion vector construction candidate in the candidate list; 7. The method of claim 6, further comprising inserting the at least one second motion vector inheritance candidate before the at least one second motion vector construction candidate in the candidate list.

8. the first threshold and the second threshold are the same as or different from each other; The first threshold and the second threshold are respectively: a positive integer, signal value, or The method of claim 7 , including at least one of the configurable values.

9. one or more processors; a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors; Apparatus for video decoding, wherein the one or more processors are configured to perform the method of any one of claims 1 to 8 upon execution of the instructions.

10. A non-transitory computer readable storage medium storing computer executable instructions and bitstreams which, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1 to 8.

11. performing an encoding method to generate a bitstream; storing the bitstream; The encoding method comprises: determining one or more candidate motion vectors from a plurality of non-adjacent neighboring blocks of a current block based on at least one scanning distance, wherein one of the at least one scanning distance indicates a number of blocks away from one side of the current block, and the number is a positive integer; obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more motion vector candidates, wherein the one or more motion vector candidates include one or more motion vector inheritance candidates and one or more motion vector construction candidates; inserting the one or more candidate motion vectors into a candidate list according to a predetermined order; the candidate list further includes one or more neighboring motion vector candidates, including one or more neighboring motion vector inheritance candidates and one or more neighboring motion vector construction candidates; the predetermined order comprises a first order; In the first order, inserting the one or more motion vector candidates into the candidate list comprises: inserting the one or more neighboring motion vector inheritance candidates before the one or more neighboring motion vector construction candidates; inserting the one or more neighboring motion vector construction candidates before the one or more motion vector inheritance candidates; inserting the one or more motion vector inheritance candidates before the one or more motion vector construction candidates.

12. A computer program for storing instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Motion Vector Prediction

    JP2020523853A