Candidate derivation for affine merge modes in video coding
Patent Information
- Application Number
- KR1020247012203
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-24
- Filing Date
- 2022-09-21
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2042-09-21
Smart Images

Figure R1020247012203_ABST
Abstract
Description
Technology Field
[0001] Cross-reference regarding related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 248,401, titled “Derivation of Candidates for Affine Merging Modes in Video Coding,” filed September 24, 2021, the entire contents of which are incorporated for all purposes.
[0003] The present disclosure relates to video coding and compression, and in particular to a method and apparatus for improving the derivation of affine merge candidates for an affine motion prediction mode in a video encoding or decoding process, but is not limited thereto. Background Technology
[0004] Various video coding technologies can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some currently well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC; also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC; also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to the previous standard VP9. Audio Video Coding (AVS), representing a digital audio and digital video compression standard, is another series of video compression standards developed by the Audio and Video Coding Standards Working Group in China. Most existing video coding standards are built upon well-known hybrid video coding frameworks, which use block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce redundancy in video images or sequences and employ transform coding to compress the energy of prediction errors. A key goal of video coding technology is to compress video data into a form that uses lower bitrates while avoiding or minimizing video quality degradation.
[0005] The first generation AVS standards include the Chinese national standards "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding, Part 16: Radio Television Video" (known as AVS+). These can provide a bitrate reduction of approximately 50% at the same perceptual quality compared to the MPEG-2 standard. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation AVS standards include the Chinese national standard "Information Technology, Efficient Multimedia Coding" (known as AVS2) series, which aims for the transmission of additional HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. AVS2 was published as a Chinese national standard in May 2016. Meanwhile, the video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as one of the international standard applications. The AVS3 standard is one of the next-generation video coding standards for UHD video applications that aims to surpass the coding efficiency of HEVC, the latest international standard. The AVS3-P2 standard was finalized at the 68th AVS meeting held in March 2019, providing a bitrate reduction of approximately 30% compared to the HEVC standard. Currently, the AVS Group maintains reference software called the High Performance Model (HPM) to demonstrate a reference implementation of the AVS3 standard.
[0006] The present disclosure provides an example of a technique related to improving the derivation of affine merge candidates for an affine motion prediction mode in a video encoding or decoding process.
[0007] According to a first aspect of the present disclosure, a video coding method is provided. The method may include the step of obtaining one or more affine candidates from a plurality of non-adjacent neighbor blocks that are not adjacent to the current block. Additionally, the method may include the step of obtaining one or more control point motion vectors (CPMV) for the current block based on the one or more affine candidates.
[0008] According to a second aspect of the present disclosure, a method for pruning affine candidates is provided. The method may include the step of calculating a first set of affine model parameters associated with one or more CPMVs of a first affine candidate. Additionally, the method may include the step of calculating a second set of affine model parameters associated with one or more CPMVs of a second affine candidate. Additionally, the method may include the step of performing a similarity check between a first affine candidate and a second affine candidate based on the first set of affine model parameters and the second set of affine model parameters.
[0009] According to a third aspect of the present disclosure, a video coding device is provided. The device includes one or more processors and a memory configured to store instructions executable by one or more processors. Additionally, one or more processors are configured to perform a method according to a first aspect or a second aspect upon execution of instructions.
[0010] According to a fourth aspect of the present disclosure, a non-transient computer-readable storage medium is provided for storing a computer-executable instruction that, when executed by one or more computer processors, causes one or more computer processors to perform a method according to a first aspect or a second aspect. Brief explanation of the drawing
[0011] A more specific description of the embodiments of the present disclosure will be provided with reference to specific embodiments illustrated in the accompanying drawings. Considering that these drawings illustrate only some embodiments and are therefore not to be construed as limiting the scope, embodiments will be described and explained with additional specificity and detail using the accompanying drawings. FIG. 1 is a block diagram of an encoder according to some embodiment of the present disclosure. FIG. 2 is a block diagram of a decoder according to some embodiment of the present disclosure. FIG. 3A is a drawing illustrating block partitioning of a multi-type tree structure according to some embodiments of the present disclosure. FIG. 3B is a drawing illustrating block partitioning of a multi-type tree structure according to some embodiments of the present disclosure. FIG. 3C is a drawing illustrating block partitioning of a multi-type tree structure according to some embodiments of the present disclosure. FIG. 3D is a drawing illustrating block partitioning of a multi-type tree structure according to some embodiments of the present disclosure. FIG. 3E is a drawing illustrating block partitioning of a multi-type tree structure according to some embodiments of the present disclosure. FIG. 4A illustrates a 4-parameter affine model according to some embodiments of the present disclosure. FIG. 4B illustrates a 4-parameter affine model according to some embodiments of the present disclosure. FIG. 5 illustrates a 6-parameter affine model according to some embodiments of the present disclosure. FIG. 6 illustrates adjacent neighbor blocks for an inherited affine merge candidate according to some embodiments of the present disclosure. FIG. 7 illustrates adjacent neighbor blocks for a constructed affine merge candidate according to some embodiments of the present disclosure. FIG. 8 illustrates non-adjacent neighbor blocks for an inherited affine merge candidate according to some embodiments of the present disclosure. FIG. 9 illustrates the derivation of configured affine merge candidates using non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 10 illustrates vertical scanning of non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 11 illustrates parallel scanning of non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 12 illustrates combined vertical and parallel scanning of non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 13A illustrates a neighboring block of the same size as the current block according to some embodiments of the present disclosure. FIG. 13B illustrates a neighboring block having a different size from the current block according to some embodiments of the present disclosure. FIG. 14A illustrates an example in which, according to some embodiments of the present disclosure, the lower left or upper right block of the lower or rightmost block of the previous street is used as the lower or rightmost block of the current street. FIG. 14A illustrates an example in which, according to some embodiments of the present disclosure, the left or upper block of the lowest or rightmost block of the previous street is used as the lowest or rightmost block of the current street. FIG. 15A illustrates scanning positions of the lower left and upper right positions used for upper and left non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 15B illustrates a scanning position of the bottom right location used for both the upper and left non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 15C illustrates a scanning position of the lower left position used for both the upper and left non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 15D illustrates a scanning position of the upper right position used for both the upper and left non-adjacent neighbor blocks according to some embodiments of the present disclosure. FIG. 16 illustrates a simplified scanning process for deriving configured merge candidates according to some embodiments of the present disclosure. FIG. 17 is a drawing illustrating a computing environment combined with a user interface according to some embodiments of the present disclosure. FIG. 18 is a flowchart illustrating a video coding method according to some embodiments of the present disclosure. FIG. 19 is a flowchart illustrating a method for pruning affine candidates according to some embodiments of the present disclosure. FIG. 20 is a block diagram illustrating a system for encoding and decoding video blocks according to some embodiments of the present disclosure. Specific details for implementing the invention
[0012] Specific embodiments will now be described in detail, examples of which are illustrated in the attached drawings. In the following detailed description, numerous non-limiting specific details are provided to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented in various types of electronic devices having digital video capabilities.
[0013] Throughout this specification, references to “one embodiment,” “an embodiment,” “an example,” “some embodiment,” “some example,” or similar language mean that a specific feature, structure, or characteristic described is included in at least one embodiment or example. A feature, structure, element, or characteristic described in connection with one or some embodiments may also apply to other embodiments unless explicitly specified otherwise.
[0014] Throughout this disclosure, terms such as “first,” “second,” “third,” etc., are used as nomenclature to refer to related elements, e.g., devices, components, configurations, steps, etc., without implying a spatial or temporal order unless explicitly otherwise specified. For example, “first device” and “second device” may mean two separately formed devices, or two parts, components, operating states, etc. of the same device, and may be named arbitrarily.
[0015] The terms “module,” “submodule,” “circuit,” “subcircuit,” “network,” “subnetwork,” “unit,” or “subunit” may include memory (shared, private, or group) that stores code or instructions that can be executed on one or more processors. A module may include one or more circuits where code or instructions are stored or are not stored. A module or circuit may include one or more components connected directly or indirectly. These components may be physically attached to each other, adjacent to each other, or not.
[0016] The terms “if” or “when” as used herein may be understood to mean “upon” or “in response to” depending on the context. These terms may not indicate that the relevant limitations or features are conditional or optional when appearing in the claims. For example, the method may include i) a step in which a function or operation X’ is performed if condition X exists, and ii) a step in which a function or operation Y’ is performed if condition Y exists. The method may be embodied in both the ability to perform a function or operation X’ and the ability to perform a function or operation Y’. Thus, functions X’ and Y’ may be performed at different times when the method is executed multiple times.
[0017] A unit or module may be implemented purely in software, purely in hardware, or as a combination of hardware and software. For example, in a purely software implementation, a unit or module may include functionally related code blocks or software components connected together, either directly or indirectly, to perform a specific function.
[0018] FIG. 20 is a block diagram illustrating an exemplary system (10) for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. As illustrated in FIG. 20, the system (10) includes a source device (12) that generates and encodes video data to be later decoded by a destination device (14). The source device (12) and the destination device (14) may include any of various electronic devices, such as a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some embodiments, the source device (12) and the destination device (14) are equipped with wireless communication capabilities.
[0019] In some implementations, the destination device (14) may receive encoded video data to be decoded via a link (16). The link (16) may include any type of communication medium or device capable of moving the encoded video data from the source device (12) to the destination device (14). As one example, the link (16) may include a communication medium that enables the source device (12) to transmit the encoded video data directly to the destination device (14) in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device (14). The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, switch, base station, or any other equipment that may be useful for facilitating communication from a source device (12) to a destination device (14).
[0020] In some other implementations, encoded video data may be transmitted from the output interface (22) to the storage device (32). Subsequently, the encoded video data within the storage device (32) may be accessed by the destination device (14) via the input interface (28). The storage device (32) may include any of various distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a Digital Versatile Disc (DVD), a Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or other suitable digital storage media for storing encoded video data. As another example, the storage device (32) may correspond to a file server or other intermediate storage device capable of holding the encoded video data generated by the source device (12). The destination device (14) may access the stored video data from the storage device (32) via streaming or download. A file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to a destination device (14). Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive.The destination device (14) can access the encoded video data through any standard data connection including a wireless channel (e.g., Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on a file server. The transmission of the encoded video data from the storage device (32) may be a streaming transmission, a download transmission, or a combination of both.
[0021] As illustrated in FIG. 20, the source device (12) includes a video source (18), a video encoder (20), and an output interface (22). The video source (18) may include a source such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video supply interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data that is the source video, or a combination of such sources. As one example, if the video source (18) is a video camera of a security surveillance system, the source device (12) and the destination device (14) may form a camera phone or a video phone. However, the embodiments described in this application may be applied to general video coding and may be applied to wireless and / or wired applications.
[0022] Captured, pre-captured, or computer-generated video may be encoded by a video encoder (20). The encoded video data may be transmitted directly to a destination device (14) through an output interface (22) of a source device (12). The encoded video data may also (or alternatively) be stored in a storage device (32) for subsequent access by the destination device (14) or another device for decoding and / or playback. The output interface (22) may further include a modem and / or transmitter.
[0023] The destination device (14) includes an input interface (28), a video decoder (30), and a display device (34). The input interface (28) may include a receiver and / or a modem and may receive encoded video data via a link (16). The encoded video data transmitted via the link (16) or provided to a storage device (32) may include various syntax elements generated by the video encoder (20) for use by the video decoder (30) when decoding the video data. These syntax elements may be included within the encoded video data that is transmitted via a communication medium, stored on a storage medium, or stored on a file server.
[0024] In some embodiments, the destination device (14) may include a display device (34) which may be an integrated display device and an external display device configured to communicate with the destination device (14). The display device (34) displays decoded video data to a user and may include any various display devices such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0025] The video encoder (20) and video decoder (30) may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of these standards. It should be understood that this application is not limited to a specific video encoding / decoding standard and may apply to other video encoding / decoding standards. Generally, the video encoder (20) of the source device (12) is considered to be configured to encode video data according to any standard of current or future standards. Likewise, the video decoder (30) of the destination device (14) is considered to be configured to decode video data according to any standard of current or future standards.
[0026] The video encoder (20) and the video decoder (30) may each be implemented with any of the various suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. In the case of partial software implementation, the electronic device may store instructions for said software on a suitable non-transient computer-readable medium and execute instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. The video encoder (20) and the video decoder (30) may each be included in one or more encoders or decoders, and any one of them may be integrated into each device as part of a combined encoder / decoder (CODE).
[0027] Like HEVC, VVC is constructed based on a block-based hybrid video coding framework. FIG. 1 is a block diagram illustrating a block-based video encoder according to some implementation of the present disclosure. In the encoder (100), the input video signal is processed block by block, which is called a coding unit (CU). The encoder (100) may be a video encoder (20) as illustrated in FIG. 20. In VTM-1.0, the CU can be up to 128x128 pixels. However, unlike HEVC, which divides blocks based only on a quad-tree, in VVC, a single coding tree unit (CTU) is divided into CUs based on quad / binary / ternary trees to adapt to various local characteristics. Additionally, the concept of multiple partition unit types in HEVC is eliminated; that is, in VVC, the separation of coding units (CU), prediction units (PU), and transform units (TU) no longer exists. Instead, each CU is always used as the base unit for both prediction and transformation without further splitting. In a multi-type tree structure, one CTU is first split into a quad-tree structure. Then, each quad-tree leaf node can be further split into binary and ternary tree structures.
[0028] FIGS. 3A through 3E are schematic diagrams illustrating multi-type tree splitting modes according to some embodiments of the present disclosure. FIGS. 3A through 3E each represent five splitting types including photo-segmentation (Fig. 3A), vertical binary splitting (Fig. 3B), horizontal binary splitting (Fig. 3C), vertically extended ternary splitting (Fig. 3D), and horizontally extended ternary splitting (Fig. 3E).
[0029] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") predicts the current video block using pixels from already coded neighboring block samples (called reference samples) within the same video picture / slice. Spatial prediction reduces spatial redundancy inherent in the video signal. Temporal prediction (also called "inter prediction" or "motion compensated prediction") predicts the current video block using reconstructed pixels from already coded video pictures. Temporal prediction reduces temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and the temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is additionally transmitted, which is used to identify which reference picture in the reference picture store the temporal prediction signal originates from.
[0030] After spatial and / or temporal prediction, the intra / inter mode determination circuit (121) of the encoder (100) selects the best prediction mode based, for example, on a rate-distortion optimization method. Then, the block predictor (120) is subtracted from the current video block; the generated prediction residual is de-correlated using a transform circuit (102) and a quantization circuit (104). The generated quantized residual coefficient is de-quantized by an inverse quantization circuitry (116) and inverse-transformed by an inverse transform circuitry (118) to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Additionally, in-loop filtering (115), such as a deblocking filter, a sample adaptive offset (SAO), and / or an adaptive in-loop filter (ALF), may be applied to the reconstructed CU before being stored in the reference picture storage of the picture buffer (117) and used to code future video blocks. To form the output video bitstream (114), the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all transmitted to an entropy coding unit (106) to be further compressed and packed to form the bitstream.
[0031] For example, deblocking filters are available in AVC, HEVC, and the latest version of VVC. In HEVC, an additional in-loop filter called SAO is defined to further improve coding efficiency. In the current version of the VVC standard, another in-loop filter called ALF is actively under research and is likely to be included in the final standard. These in-loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. These operations can also be disabled at the discretion of the encoder (100) to reduce computational complexity.
[0032] It should be noted that intra-prediction is generally based on unreconstructed pixels that are not filtered, whereas inter-prediction is based on reconstructed pixels that are filtered when these filter options are enabled by the encoder (100).
[0033] FIG. 2 is a block diagram illustrating a block-based video decoder (200) that can be used with many video coding standards. This decoder (200) is similar to the reconstruction-related section present in the encoder (100) of FIG. 1. The block-based video decoder (200) may be a video decoder (30) as illustrated in FIG. 1. In the decoder (200), an input video bitstream (201) is first decoded through entropy decoding (202) to derive quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed through inverse quantization (204) and inverse transformation (206) to obtain reconstructed prediction residuals. A block predictor mechanism implemented within an intra / inter mode selector (212) is configured to perform intra-prediction (208) or motion compensation (210) based on the decoded prediction information. An unfiltered reconstructed pixel set is obtained by summing the prediction residuals reconstructed from the inverse transformation (206) and the prediction output generated by the block predictor mechanism using a summer (summer, 214).
[0034] The reconstructed block may pass through an in-loop filter (209) further before being stored in a picture buffer (213) that functions as a reference picture storage. The reconstructed video in the picture buffer (213) can be transmitted to drive a display device, as well as used to predict future video blocks. When the in-loop filter (209) is enabled, a filtering operation is performed on the reconstructed pixels to yield a final reconstructed video output (222).
[0035] In current VVC and AVS3 standards, motion information of the current coding block is copied from spatial or temporal neighbor blocks specified by a merge candidate index or obtained through explicit signaling of motion estimation. The focus of the present invention is to improve the method for deriving affine merge candidates to enhance the accuracy of motion vectors for affine merge modes. To facilitate the description of this disclosure, the existing affine merge mode design of the VVC standard is used as an example to illustrate the proposed idea. It should be noted that while the existing affine mode design of the VVC standard is used as an example throughout this disclosure, the proposed technique can be applied to different designs of affine motion prediction modes or to other coding tools having the same or similar design principles to those skilled in modern video coding technology.
[0036] Affine Model
[0037] In HEVC, only translational motion models are applied for motion compensation prediction. However, the real world has various types of motion, such as zoom in / out, rotation, perspective motion, and other irregular motions. VVC and AVS3 apply affine motion compensation prediction by signaling a flag for each intercoding block to indicate whether a translational motion model or an affine motion model is applied for interpretation. In the current VVC and AVS3 design, two affine modes are supported for a single affine coding block, including a 4-parameter affine mode and a 6-parameter affine mode.
[0038] The 4-parameter affine model has the following parameters: two parameters for translational movement in the horizontal and vertical directions, one parameter for zoom movement, and one parameter for bidirectional rotation movement. In this model, the horizontal zoom parameter is identical to the vertical zoom parameter, and the horizontal rotation parameter is identical to the vertical rotation parameter. To achieve better accommodation of motion vectors and affine parameters, these affine parameters must be derived from two MVs (also called Control Point Motion Vectors (CPMVs)) located at the top-left and top-right corners of the current block. As illustrated in Figures 4A and 4B, the affine motion field of the block is described by two CPMVs (V0, V1). Based on the control point motion, the motion field (v) of a single affine-coded block x , v y ) is explained as follows:
[0039] (1)
[0040] The 6-parameter affine mode has the following parameters: two parameters for translational movement in the horizontal and vertical directions, respectively, and two other parameters for vertical zoom movement and rotational movement, respectively. The 6-parameter affine movement model is coded with three CPMVs. As shown in Fig. 5, the three control points of a single 6-parameter affine block are located at the top-left, top-right, and bottom-left corners of the block. The movement of the top-left control point is related to translational movement, the movement of the top-right control point is related to horizontal rotation and zoom movement, and the movement of the bottom-left control point is related to vertical rotation and zoom movement. Compared to the 4-parameter affine movement model, the horizontal rotation and zoom movement of the 6-parameter model may not be the same as the vertical movement. Assuming (V0, V1, V2) are the MVs of the top-left, top-right, and bottom-left corners of the current block in Fig. 5, each sub-block (v x , v y The motion vector of ) is derived as follows using three MVs at the control point:
[0041] (2)
[0042] Affine Merge Mode
[0043] In Affine Merge mode, the CPMV for the current block is not explicitly signaled but is derived from neighboring blocks. Specifically, in this mode, the CPMV for the current block is generated using motion information from spatial neighboring blocks. The list of Affine Merge mode candidates has a limited size. For example, there may be up to five candidates in the current VVC design. The encoder can evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder. Affine merge candidates can be determined in three ways. In the first way, the Affine merge candidate can be inherited from neighboring Affine-coded blocks. In the second way, the Affine merge candidate can be constructed from the translational MV of neighboring blocks. In the third way, the zero MV is used as the Affine merge candidate.
[0044] In the case of the method of inheritance, there may be up to two candidates. The candidates are obtained, if available, from neighbor blocks located at the bottom left of the current block (e.g., scanning order from A0 to A1 as shown in FIG. 6) and neighbor blocks located at the top right of the current block (e.g., scanning order from B0 to B2 as shown in FIG. 6).
[0045] In the case of the method being constructed, the candidates are combinations of neighboring translational MVs, and these can be generated in two steps.
[0046] Step 1: Obtain four translational MVs including MV1, MV2, MV3, and MV4 from available neighbors.
[0047] MV1: This is the MV from one of the three neighboring blocks close to the top-left corner of the current block. As shown in FIG. 7, the scanning order is B2, B3, and A2.
[0048] MV2: This is the MV from one of the two neighboring blocks closest to the top right of the current block. As shown in FIG. 7, the scanning order is B1 and B0.
[0049] MV3: This is the MV from one of the two neighboring blocks close to the bottom-left corner of the current block. As shown in FIG. 7, the scanning order is A1 and A0.
[0050] MV4: This is the MV of the temporally juxtaposed block of the neighbor block closest to the bottom right of the current block. As shown in FIG. 7, the neighbor block is T.
[0051] Step 2: Derive combinations based on the four translational MVs from Step 1.
[0052] Combination 1: MV1, MV2, MV3;
[0053] Combination 2: MV1, MV2, MV4;
[0054] Combination 3: MV1, MV3, MV4;
[0055] Combination 4: MV2, MV3, MV4;
[0056] Combination 5: MV1, MV2;
[0057] Combination 6: MV1, MV3.
[0058] If the merge candidate list is not full after being filled with inherited candidates and configured candidates, zero MV is inserted at the end of the list.
[0059] In the case of current video standards VVC and AVS, only adjacent neighbor blocks are used to derive affine merge candidates for the current block, as shown in FIGS. 6 and 7 for inherited candidates and configured candidates, respectively. To increase the diversity of merge candidates and explore spatial correlations in more detail, it is simple to extend the range of neighbor blocks from adjacent regions to non-adjacent regions.
[0060] In the present disclosure, the candidate derivation process for an affine merge mode is extended to use not only adjacent neighbor blocks but also non-adjacent neighbor blocks. The specific method can be summarized in three aspects, including affine merge candidate pruning, a non-adjacent neighbor-based derivation process for affine inherited merge candidates, and a non-adjacent neighbor-based derivation process for affine configured merge candidates.
[0061] Affine merger candidate Pruning
[0062] In typical video coding standards, since the list of affine merge candidates generally has a limited size, candidate pruning is an essential process for removing duplicate candidates. This pruning process is required for both inherited and constructed candidates in the affine merge. As explained in the introduction, the CPMV of the current block is not used directly for affine motion compensation. Instead, the CPMV must be converted into translational MV at the location of each sub-block within the current block. The conversion process follows a general affine model as shown below:
[0063] (3)
[0064] Here ( a,b ) is the delta translation parameter, and ( CD ) are the delta zoom and rotation parameters for the horizontal direction, and ( e,f ) are the delta zoom and rotation parameters for the vertical direction, and ( x,y ) is the current block (e.g., the coordinates shown in FIG. 5 ( x,y It is the horizontal and vertical distance of the pivot position (e.g., center or top-left corner) of the sub-block relative to the top-left corner of )), and ( v x ,v y ) is the target translation MV of the sub-block.
[0065] For the 6-parameter affine model, three CPMVs named V0, V1, and V2 are available. Then, the six model parameters a, b, c, d, e and f can be calculated as follows:
[0066] (4)
[0067] For a 4-parameter affine model, if the top-left corner CPMV and the top-right corner CPMV named V0 and V1 are available, a, b, c, d, e and f The 6 parameters of can be calculated as follows:
[0068] (5)
[0069] For a 4-parameter affine model, if the top-left corner CPMV and bottom-left corner CPMV named V0 and V2 are available, a, b, c, d, e and f The 6 parameters of can be calculated as follows:
[0070] (6)
[0071] In the above equations (4), (5) and (6), w and h represent the width and height of the current block, respectively.
[0072] When two sets of CPMV merge candidates are compared for redundancy checks, we propose checking the similarity of six affine model parameters. Thus, the candidate pruning process can be performed in two stages.
[0073] In Step 1, given two sets of CPMV candidates, corresponding affine model parameters for each set of candidates are derived. More specifically, two sets of CPMV candidates can be represented by two sets of affine model parameters, e.g., (a1, b1, c1, d1, e1, f1) and (a2, b2, c2, d2, e2, f2).
[0074] In step 2, a similarity check is performed between two sets of affine model parameters based on one or more predefined thresholds. In one embodiment, if the absolute values of (a1-a2), (b1-b2), (c1-c2), (d1-d2), (e1-e2) and (f1-f2) are all lower than a positive threshold value, such as 1, the two candidates are considered similar and one of them is pruned or removed and is not included in the merge candidate list.
[0075] In some embodiments, the calculation of the CPMV pruning process can be simplified by removing the division or right shift operation of step 1.
[0076] Specifically, c, d, e and f The model parameters can be calculated without being divided by the width (w) and height (h) of the current block. For example, taking the above equation (4) as an example, c', d', e' and f' The approximate model parameters of can be calculated as shown in the equation (7) below:
[0077] (7)
[0078] If only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters that depend on the width or height of the current block. In this case, the model parameters can be transformed to account for the influence of the width and height. For example, in the case of Equation (5), c', d', e' and f' The approximate model parameters of can be calculated based on the following equation (8). In the case of equation (6), c', d', e' and f' The approximate model parameters of can be calculated based on the following equation (9):
[0079] (8)
[0080] (9)
[0081] In Step 1 above c', d', e' and f' Once the approximate model parameters are calculated, the calculation of the absolute values required for the similarity check in Step 2 above can be changed accordingly as follows: (a1- a2), (b1- b2), (c1'- c2'), (d1'- d2'), (e1' - e2'), and (f1'- f2').
[0082] In Step 2 above, a threshold is required to evaluate the similarity between two sets of CPMV candidate CPMVs. There may be various ways to define the threshold. In one embodiment, the threshold may be defined for each comparable parameter. Table 1 is one example of this embodiment showing thresholds defined for each comparable model parameter. In another embodiment, the threshold may be defined considering the size of the current coding block. Table 2 is one example of this embodiment showing thresholds defined by the size of the current coding block.
[0083] Comparable parameters threshold a 1 b 1 c 2 d 2 e 2 f 2
[0084] Current block size threshold Size <= 64 pixels 1 64 pixels < size <= 256 pixels 2 256 pixels < size <= 1024 pixels 4 1024 pixels < size 8
[0085] In another embodiment, the threshold may be defined by taking into account the weight or height of the current block. Tables 3 and 4 are examples of the present embodiment. Table 3 shows the threshold defined according to the width of the current coding block, and Table 4 shows the threshold defined according to the height of the current coding block.
[0086] Current block width threshold Width <= 8 pixels 1 8 pixels < width <= 32 pixels 2 32 pixels < width <= 64 pixels 4 64 pixels < width 8
[0087] Current block height threshold Height <= 8 pixels 1 8 pixels < height <= 32 pixels 2 32 pixels < height <= 64 pixels 4 64 pixels < height 8
[0088] In another embodiment, the threshold value may be defined as a series of fixed values. In another embodiment, the threshold value may be defined by any combination of the above embodiments. As one example, the threshold value may be defined by taking into account different parameters and the weight and height of the current block. Table 5 is an example of this embodiment showing a threshold value defined by the height of the current coding block. It should be noted that in any of the above proposed embodiments, if necessary, the comparable parameter may represent any parameter defined in any of the equations from equation (4) to equation (9).
[0089] Comparable parameters threshold a 1 b 1 c Width <= 8 pixels: 18 pixels < Width <= 32 pixels: 232 pixels < Width <= 64 pixels: 464 pixels < Width: 8 e d Height <= 8 pixels: 18 pixels < Height <= 32 pixels: 232 pixels < Height <= 64 pixels: 464 pixels < Height: 8 f
[0090] The advantages of using transformed affine model parameters for candidate redundancy checks are as follows: it creates an integrated similarity check process for candidates with different affine model types, for example, one merge candidate may use a 6-parameter affine model with three CPMVs, while another candidate may use a 4-parameter affine model with two CPMVs; it considers the different influences of each CPMV in the merge candidate when deriving the target MV in each sub-block; and it provides the similarity significance of two affine merge candidates regarding the width and height of the current block.
[0091] Non-neighbor-based derivation process for affine inheritance merger candidates
[0092] For inherited merge candidates, the non-neighbor-based derivation process can be performed in three steps. Step 1 is the candidate scan step. Step 2 is the CPMV projection step. Step 3 is the candidate pruning step.
[0093] In Step 1, non-adjacent neighbor blocks are scanned and selected in the following way.
[0094] Scanning area and distance
[0095] In some embodiments, non-adjacent neighbor blocks may be scanned from the left and top regions of the current coding block. The scanning distance may be defined as the number of coding blocks from the scanning position to the left or top side of the current coding block.
[0096] As illustrated in FIG. 8, multiple lines of non-adjacent neighbor blocks may be scanned to the left or above the current coding block. The distances illustrated in FIG. 8 represent the number of coding blocks from each candidate location to the left or top side of the current block. For example, the "Distance 2" area to the left of the current block indicates that the candidate neighbor block located in that area is two blocks away from the current block. Similar indications may be applied to other scanning areas with different distances.
[0097] In one or more embodiments, as illustrated in FIG. 13A, the non-adjacent neighbor blocks of each distance may have the same block size as the current coding block. As illustrated in FIG. 13A, the non-adjacent neighbor block on the left (1301) and the non-adjacent neighbor block above (1302) have the same size as the current block (1303). In some embodiments, as illustrated in FIG. 13B, the non-adjacent neighbor blocks of each distance may have a different block size from the current coding block. Neighbor block (1304) is a neighbor block adjacent to the current block (1303). As illustrated in FIG. 13B, the non-adjacent neighbor block on the left (1305) and the non-adjacent neighbor block above (1306) have the same size as the current block (1307). Neighbor block (1308) is a neighbor block adjacent to the current block (1307).
[0098] It is worth noting that when the non-neighboring neighbor blocks of each distance have the same block size as the current coding block, the value of the block size changes adaptively according to the partition granularity in each different region of the image. It is worth noting that when the non-neighboring neighbor blocks of each distance have a different block size than the current coding block, the value of the block size can be predefined as a constant value, such as 4x4, 8x8, or 16x16.
[0099] Based on the defined scanning distance, the total size of the scanning area to the left or above the current coding clock can be determined by a configurable distance value. In one or more embodiments, the maximum scanning distances to the left and above may use the same value or different values. FIG. 13 shows an example where the maximum distances to both the left and above share the same value of 2. The maximum scanning distance value(s) can be determined by the encoder side and signaled as a bitstream. Alternatively, the maximum scanning distance value(s) can be predefined as fixed value(s), such as a value of 2 or 4. If the maximum scanning distance is predefined as a value of 4, it indicates that the scanning process has ended first, either when the candidate list is full or when all non-adjacent neighbor blocks with a maximum distance of 4 have been scanned.
[0100] In one or more embodiments, within each scanning area at a specific distance, the neighbor blocks that start and end may vary depending on the location.
[0101] In some embodiments, for the left scanning area, the starting neighbor block may be the adjacent bottom-left block of the starting neighbor block of the adjacent scanning area at a shorter distance. For example, as illustrated in FIG. 8, the starting neighbor block of the "distance 2" scanning area to the left of the current block is the adjacent bottom-left neighbor block of the starting neighbor block of the "distance 1" scanning area. The ending neighbor block may be the adjacent left block of the ending neighbor block of the upper scanning area at a shorter distance. For example, as illustrated in FIG. 8, the ending neighbor block of the "distance 2" scanning area to the left of the current block is the adjacent left neighbor block of the ending neighbor block of the "distance 1" scanning area above the current block.
[0102] Similarly, for the upper scanning area, the starting neighbor block may be the adjacent block to the top right of the starting neighbor block of the shorter adjacent scanning area. The ending neighbor block may be the adjacent block to the top left of the ending neighbor block of the shorter adjacent scanning area.
[0103] Scanning order
[0104] When neighbor blocks are scanned in a non-adjacent area, a specific order and / or rule may be followed to determine the selection of the scanned neighbor blocks.
[0105] In some embodiments, the left area may be scanned first, followed by the upper area. As shown in FIG. 8, three lines of a non-adjacent area on the left (e.g., distances 1 through 3) may be scanned first, followed by three lines of a non-adjacent area on the current block.
[0106] In some embodiments, the left area and the upper area may be scanned alternately. For example, as shown in FIG. 8, the left scanning area of "distance 1" is scanned first, and then the upper area of "distance 1" is scanned.
[0107] For scanning areas located on the same side (e.g., the left or top area), the scanning sequence proceeds from the shortest distance area to the farthest distance area. This sequence can be flexibly combined with other embodiments of the scanning sequence. For example, the left area and the top area may be scanned alternately, and the sequence for the same side area is scheduled from short distance to far distance.
[0108] A scanning order can be defined within each scanning area at a specific distance. In one embodiment, for the left scanning area, scanning may start from the bottom neighbor block to the top neighbor block. For the top scanning area, scanning may start from the right block to the left block.
[0109] Scanning complete
[0110] In the case of inherited merge candidates, neighboring blocks coded in affine mode are defined as eligible candidates. In some embodiments, the scanning process may be performed interactively. For example, scanning performed in a specific area at a specific distance may be stopped when the first X eligible candidates are identified, where X is a predefined positive value. For example, as illustrated in FIG. 8, scanning in the left scanning area at distance 1 may be stopped when the first one or more eligible candidates are identified. Then, the next iteration of the scanning process targeting another scanning area is started according to a predefined scanning order / rule.
[0111] In some embodiments, the scanning process may be performed continuously. For example, scanning performed in a specific area of a specific distance may be stopped when all included neighboring blocks are scanned and no more eligible candidates are identified or when the maximum number of allowed candidates is reached.
[0112] During the candidate scanning process, each candidate non-neighboring neighbor block is determined and scanned according to the proposed scanning method. For easier implementation, each candidate non-neighboring neighbor block may be indicated or located by a specific scanning location. Once the distance from a specific scanning area is determined according to the proposed method, the scanning location can be appropriately determined based on the following method.
[0113] In one method, as shown in FIG. 15A, the bottom-left and top-right positions are used for the top and left non-adjacent neighbor blocks, respectively.
[0114] In another method, as shown in Fig. 15B, the bottom right position is used for both the upper and left non-adjacent neighbor blocks.
[0115] In another method, as shown in Fig. 15C, the bottom-left position is used for both the upper and left non-adjacent neighbor blocks.
[0116] In another method, as shown in Fig. 15D, the top right position is used for both the upper and left non-adjacent neighbor blocks.
[0117] For easier explanation, in FIGS. 15A through 15D, each non-adjacent neighbor block is assumed to have the same block size as the current block. Without loss of generality, this example can be easily extended to non-adjacent neighbor blocks having different block sizes.
[0118] Additionally, in step 2, the same process of CPMV projection used in current AVS and VVC standards may be used. In this CPMV projection process, it is assumed that the current block shares the same affine model as a selected neighbor block, and two or three corner pixel coordinates (e.g., if the current block uses a 4-parameter model, two coordinates (top-left pixel / sample position) are used; if the current block uses a 6-parameter model, three coordinates (top-left pixel / sample position, top-right pixel / sample position and bottom-left pixel / sample position) are used) are substituted into equation (1) or (2), and depending on whether the neighbor block is coded with a 4-parameter or 6-parameter affine model, the equation is selected to generate two or three CPMVs.
[0119] In Step 3, all eligible candidates identified in Step 1 and transformed in Step 2 may undergo a similarity check against all existing candidates already in the merge candidate list. Details regarding the similarity check have already been described in the section on pruning affine merge candidates above. If a new eligible candidate is found to be similar to an existing candidate in the candidate list, this new eligible candidate is removed / pruned.
[0120] Non-neighbor-based derivation process for affine-configured merge candidates
[0121] When deriving an inherited merge candidate, one neighbor block is identified at a time, which must be coded in affine mode and may contain two or three CPMVs. When deriving a configured merge candidate, two or three neighbor blocks may be identified at a time, each identified neighbor block does not need to be coded in affine mode, and only one translational MV is retrieved from this block.
[0122] FIG. 9 illustrates an example in which a constructed affine merge candidate can be derived using non-adjacent neighbor blocks. In FIG. 9, A, B, and C represent the geographical locations of three non-adjacent neighbor blocks. A virtual coding block is formed using the location of A as the top-left corner, the location of B as the top-right corner, and the location of C as the bottom-left corner. If the virtual CU is considered as an affine-coded block, the MVs at locations A', B', and C' can be derived according to Equation (3), where the model parameters (a, b, c, d, e, f) can be calculated by the translational MVs at locations A, B, and C. Once derived, the MVs at locations A', B', and C' can be used as three CPMVs for the current block, and an existing process for generating a constructed affine merge candidate (a process used in AVS and VVC standards) can be used.
[0123] For a configured merge candidate, the non-neighbor-based derivation process can be performed in five steps. The non-neighbor-based derivation process can be performed in five steps on a device such as an encoder or decoder. Step 1 is the candidate scanning step. Step 2 is the affine model determination step. Step 3 is the CPMV projection step. Step 4 is the candidate generation step. And Step 5 is the candidate pruning step. In Step 1, non-neighbor blocks can be scanned and selected in the following way.
[0124] Scanning area and distance
[0125] In some embodiments, to maintain a rectangular coding block, the scanning process is performed for only two non-adjacent neighbor blocks. The third non-adjacent neighbor block may depend on the horizontal and vertical positions of the first and second non-adjacent neighbor blocks.
[0126] In some embodiments, as illustrated in FIG. 9, the scanning process is performed only for the positions of B and C. The position of A can be uniquely determined by the horizontal position of C and the vertical position of B. In this case, the scanning area and distance can be defined according to a specific scanning direction.
[0127] In some embodiments, the scanning direction may be perpendicular to the side of the current block. One example is illustrated in FIG. 10, where the scanning area is defined as a line of continuous motion fields located to the left or above the current block. The scanning distance is defined as the number of motion fields from the scanning position to the side of the current block. It should be noted that the size of the motion field may depend on the maximum granularity of the applicable video coding standard. In the example illustrated in FIG. 10, the size of the motion field was set to 4x4, assuming it matches the current VVC standard.
[0128] In some embodiments, the scanning direction may be parallel to the side of the current block. One example is illustrated in FIG. 11, where the scanning area is defined as a line of consecutive coding blocks to the left or above the current block.
[0129] In some embodiments, the scanning direction may be a combination of vertical and parallel scanning with respect to the side of the current block. One example is illustrated in FIG. 12. As illustrated in FIG. 12, the scanning direction may be a combination of parallel and diagonal. Scanning at position B starts from left to right and then proceeds diagonally to the left and upper blocks. Scanning at position B will be repeated as illustrated in FIG. 12. Similarly, scanning at position C starts from top to bottom and then proceeds diagonally to the left and upper blocks. Scanning at position C will be repeated as illustrated in FIG. 12.
[0130] Scanning order
[0131] In some embodiments, the scanning order may be defined from a position with a short distance to a position with a long distance to the current coding block. This order may also apply to vertical scanning.
[0132] In some embodiments, the scanning order may be defined as a fixed pattern. This fixed pattern scanning order may be used for candidate locations of similar distance. One example is the case of parallel scanning. As one example, the scanning order may be defined from top to bottom for the left scanning area and from left to right for the upper scanning area, as shown in the example in FIG. 11.
[0133] In the case of a combined scanning method, the scanning sequence may be a combination of a fixed pattern and a distance-dependent one, as shown in the example in FIG. 12.
[0134] Scanning complete
[0135] For configured merge candidates, there is no need to affine code eligible candidates because only translational MV is required.
[0136] Depending on the number of candidates required, the scanning process can be terminated when the first X eligible candidates are identified, where X is a positive value.
[0137] As illustrated in FIG. 9, three corners A, B, and C are required to form a virtual coding block. For easier implementation, the scanning process of step 1 may be performed only to identify non-adjacent neighbor blocks located at corners B and C, whereby the coordinates of A can be accurately determined by considering the horizontal coordinates of C and the vertical coordinates of C. In this way, the formed virtual coding block is limited to a rectangle. If either point B or C is unavailable, for example, if it is out of bounds, or if movement information of the non-adjacent neighbor block corresponding to B or C is unavailable, the horizontal coordinates or vertical coordinates of C can be defined as the horizontal coordinates or vertical coordinates of the top-left point of the current block, respectively.
[0138] To maintain consistency, the scanning area and distance definitions, scanning order, and scanning termination methods proposed for deriving inherited merge candidates may be reused wholly or partially to deriving configured merge candidates. In one or more embodiments, the same methods defined for the inherited merge candidate scanning, including but not limited to the scanning area and distance, scanning order, and scanning termination, may be reused wholly for the configured merge candidate scanning.
[0139] In some embodiments, the same method defined for inherited merge candidate scanning may be partially reused for configured merge candidate scanning. FIG. 16 illustrates an example of such a case. In FIG. 16, the block size of each non-adjacent neighbor block is the same as the current block, which is defined similarly to inherited candidate scanning, but the scanning at each distance is limited to only one block, thus simplifying the entire process.
[0140] In Step 2, the translational MV at the positions of the candidates selected after Step 1 is evaluated, and an appropriate affine model can be determined. To explain more easily and without loss of generality, Figure 9 is used again as an example.
[0141] Due to factors such as hardware constraints, implementation complexity, and various reference indices, the scanning process may be terminated before a sufficient number of candidates are identified. For example, motion information in the motion field of one or more of the candidates selected after Step 1 may be unavailable.
[0142] If motion information for all three candidates is available, the corresponding virtual coding block represents a 6-parameter affine model. If motion information for one of the three candidates is unavailable, the corresponding virtual coding block represents a 4-parameter affine model. If motion information for one or more of the three candidates is unavailable, the corresponding virtual coding block will not represent a valid affine model.
[0143] In some embodiments, if the motion information of the upper left corner of the virtual coding block, e.g., corner A of FIG. 9, is unavailable, or if the motion information of the upper right corner, e.g., corner B of FIG. 9, and the lower left corner, e.g., corner C of FIG. 9, is unavailable, the virtual block is set to invalid so that it cannot represent a valid model, and steps 3 and 4 can be skipped in the current iteration.
[0144] In some embodiments, only one of the top right corner, e.g., corner B in FIG. 9, or the bottom left corner, e.g., corner C in FIG. 9, is unavailable, but if both are unavailable, the virtual block may represent a valid 4-parameter affine model.
[0145] In step 3, if the virtual coding block can represent a valid affine model, the same projection process used for the inherited merge candidate can be used.
[0146] In one or more embodiments, the same projection process used for the inherited merge candidate may be used. In this case, the 4-parameter model represented by the virtual coding block in step 2 may be projected as a 4-parameter model for the current block, and the 6-parameter model represented by the virtual coding block in step 2 may be projected as a 6-parameter model for the current block.
[0147] In some embodiments, the affine model represented by the virtual coding block in step 2 is always projected as a 4-parameter model or a 6-parameter model for the current block.
[0148] It is worth noting that according to equations (5) and (6), there may be two types of 4-parameter affine models, where type A is the case where the top-left corner CPMV and the top-right corner CPMV named V0 and V1 are available, and type B is the case where the top-left corner CPMV and the bottom-left corner CPMV named V0 and V2 are available.
[0149] In one or more embodiments, the type of the projected 4-parameter affine model is the same type as the 4-parameter affine model represented by the virtual coding block. For example, the affine model represented by the virtual coding block in step 2 is a 4-parameter affine model of type A or type B, and the affine model projected onto the current block is also of type A or type B, respectively.
[0150] In some embodiments, the 4-parameter affine model represented by the virtual coding block in step 2 is always projected as the same type of 4-parameter model for the current block. For example, Type A or Type B of the 4-parameter affine model represented by the virtual coding block is always projected as the 4-parameter affine model of Type A.
[0151] In Step 4, based on the projected CPMV from Step 3 onwards, the same candidate generation process used in the current VVC or AVS standard may be used, as an example. In another embodiment, the temporal motion vector used in the candidate generation process for the current VVC or AVS standard may not be used in the non-neighbor block-based derivation method. If the temporal motion vector is not used, this indicates that the generated combination does not include the temporal motion vector.
[0152] In Step 5, newly generated candidates after Step 4 may undergo a similarity check against all existing candidates already in the merge candidate list. Details regarding the similarity check have already been explained in the section on pruning affine merge candidates. If a newly generated candidate is found to be similar to an existing candidate in the candidate list, the newly generated candidate is removed or pruned.
[0153] FIG. 17 illustrates a computing environment (or computing device) (1710) combined with a user interface (1760). The computing environment (1710) may be part of a data processing server. In some embodiments, the computing device (1710) may perform any of the various methods or processes (such as encoding / decoding methods or processes) as previously described according to various embodiments of the present disclosure. The computing environment (1710) may include a processor (1720), memory (1740), and an I / O interface (1750).
[0154] The processor (1720) generally controls the overall operation of the computing environment (1710), such as operations related to display, data acquisition, data communication, and image processing. The processor (1720) may include one or more processors that execute instructions to perform all or some steps of the method described above. Additionally, the processor (1720) may include one or more modules that facilitate interaction between the processor (1720) and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0155] Memory (1740) is configured to store various types of data to support the operation of the computing environment (1710). Memory (1740) may include certain software (1742). Examples of such data may include instructions for any application or method operating in the computing environment (1710), video datasets, image data, etc. Memory (1740) may be implemented using any type of volatile or non-volatile memory device, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk, or a combination thereof.
[0156] The I / O interface (1750) provides an interface between the processor (1720) and peripheral interface modules such as a keyboard, click wheel, buttons, etc. Buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. The I / O interface (1750) may be combined with an encoder and a decoder.
[0157] In some embodiments, a non-transient computer-readable storage medium is also provided that includes a plurality of programs, such as those contained in memory (1740), which are executable by a processor (1720) of a computing environment (1710) to perform the method described above. For example, the non-transient computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0158] A non-transient computer-readable storage medium stores a plurality of programs to be executed by a computing device having one or more processors, wherein the plurality of programs cause the computing device to perform the motion prediction method described above when executed by one or more processors.
[0159] In some embodiments, the computing environment (1710) may be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing units (DSPDs), programmable logic devices (PLDs), field programmable gates (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components.
[0160] FIG. 18 is a flowchart illustrating a video coding method according to an embodiment of the present disclosure.
[0161] In step 1801, the processor (1720) may obtain one or more affine candidates from a plurality of non-adjacent neighbor blocks that are not adjacent to the current block or CU.
[0162] In some embodiments, a plurality of non-adjacent neighbor blocks may include non-adjacent coding blocks as shown in FIG. 11, FIG. 12, FIG. 13A, FIG. 13B, FIG. 14A, FIG. 14B, FIG. 15A to FIG. 15D and FIG. 16.
[0163] In some embodiments, the processor (1720) may acquire one or more affine candidates according to a scanning rule.
[0164] In some embodiments, the scanning rule may be determined based on at least one scanning area, at least one scanning distance, and a scanning order.
[0165] In some embodiments, at least one scanning distance represents the number of blocks away from the side of the current block.
[0166] In some embodiments, one of the plurality of non-adjacent neighbor blocks at at least one scanning distance may have the same size as the current block as shown in FIG. 13A, or may have a different size from the current block as shown in FIG. 13B.
[0167] In some embodiments, at least one scanning area may include a first scanning area and a second scanning area, wherein the first scanning area is determined by a first maximum scanning distance representing the maximum number of blocks away from a first side of the current block, and the second scanning area is determined by a second maximum scanning distance representing the maximum number of blocks away from a second side of the current block, and the first maximum scanning distance is the same as or different from the second maximum scanning distance. In some embodiments, the first maximum scanning distance or the second maximum scanning distance may be set to a fixed value such as 3, 4, etc.
[0168] For example, the first scanning area may be the left area of the current block (1303), and the first maximum scanning distance is three blocks away from the left side of the current block (1303). That is, the block (1301) is three blocks away from the left side of the current block (1303) at the first maximum scanning distance. Additionally, the second scanning area may be the upper surface area of the current block (1303), and the second maximum scanning distance is three blocks away from the top or upper surface of the current block (1303). That is, the block (1302) is three blocks away from the upper / upper surface of the current block (1303) at the second maximum scanning distance.
[0169] In some embodiments, the encoder may signal a first maximum scanning distance and a second maximum scanning distance to be transmitted to the decoder.
[0170] In some embodiments, the processor (1720) may stop scanning in at least one scanning area as a scan termination in response to determining that the first or second maximum scanning distance is equal to a fixed value, and in response to determining that the candidate list is full or that all non-adjacent neighbor blocks within the first or second maximum scanning distance have been scanned.
[0171] In some embodiments, the processor (1720) can scan a plurality of non-adjacent neighbor blocks within a first scanning area to obtain one or more non-adjacent neighbor blocks coded in an affine mode, and can determine one or more non-adjacent neighbor blocks coded in an affine mode as one or more affine candidates.
[0172] In some embodiments, the processor (1720) may scan from a first starting non-adjacent neighbor block along a scanning line parallel to the left of the current block, wherein the first starting non-adjacent block is a bottom block within a first scanning area, and the blocks within the first scanning area are at a first scanning distance from the left of the current block, e.g., D2 in FIG. 14A.
[0173] In some embodiments, the first starting non-adjacent block may be located at the bottom and left of the second starting non-adjacent neighbor block within the second scanning area, and the blocks within the second scanning area may be located at a second scanning distance from the left side of the current block, e.g., D1 in FIG. 14A, as shown in FIG. 14A. In some other examples, the first starting non-adjacent block may be located to the left of the second starting non-adjacent neighbor block within the second scanning area, and the blocks within the second scanning area may be located at a second scanning distance from the left side of the current block, as shown in FIG. 14B.
[0174] In some embodiments, the processor (1720) may scan from a third starting non-adjacent neighbor block along a scanning line parallel to the top face of the current block, wherein the third starting non-adjacent block may be a right block within the first scanning area, and the blocks within the first scanning area may be at a first scanning distance from the top face of the current block, e.g., D2 in FIG. 14A.
[0175] In some embodiments, the third starting non-adjacent block may be located above and to the right of the fourth starting non-adjacent neighbor block within the second scanning area, and the blocks within the second scanning area may be located at a second scanning distance from the upper face of the current block, e.g., D1 in FIG. 14A, as shown in FIG. 14A. In some other embodiments, the third starting non-adjacent block may be located to the right of the fourth starting non-adjacent neighbor block within the second scanning area, and the blocks within the second scanning area may be located at a second scanning distance from the upper face of the current block, as shown in FIG. 14B.
[0176] In some embodiments, the processor (1720) may position non-neighboring neighbor blocks at scanning locations. For example, for easier implementation, each candidate non-neighboring neighbor block may be indicated or positioned by a specific scanning location.
[0177] In some embodiments, the scanning positions may include a lower-left position of a non-adjacent neighbor block within a second scanning area above the current block as shown in FIG. 15A, a upper-right position of a non-adjacent neighbor block within a first scanning area to the left of the current block as shown in FIG. 15A, a lower-right position of a non-adjacent neighbor block within a first scanning area or a second scanning area as shown in FIG. 15B, a lower-left position of a non-adjacent neighbor block within a first scanning area or a second scanning area as shown in FIG. 15C, and a upper-right position of a non-adjacent neighbor block within a first scanning area or a second scanning area as shown in FIG. 15D.
[0178] In some embodiments, the processor (1720) obtains a first candidate location for a first affine candidate and a second candidate location for a second affine candidate based on a scanning rule; determines a third candidate location for a third affine candidate based on the first and second candidate locations; obtains a virtual block based on the first candidate location, the second candidate location, and the third candidate location; obtains three CPMVs for the virtual block based on translational MVs at the first candidate location, the second candidate location, and the third candidate location; and can obtain two or three CPMVs for the current block based on the three CPMVs of the virtual block using the same projection process used for inherited candidate derivation.
[0179] In some embodiments, the virtual block may be a rectangular coding block, and the third candidate position may be determined based on the vertical position of the first candidate position and the horizontal position of the second candidate position. For example, the virtual block may be a virtual block including positions A, B, and C as shown in FIG. 9.
[0180] In some embodiments, the processor (1720) may determine the vertical position of the third candidate position as the vertical position of the top-left point of the current block, and in response to determining that the first candidate position or the second candidate position is unavailable, or in response to determining that movement information at the first candidate position or the second candidate position is unavailable, the horizontal position of the third candidate position may be determined as the horizontal position of the top-left point of the current block.
[0181] In some embodiments, the processor (1720) may determine that the virtual block cannot represent a valid affine model in response to determining that motion information at the first, second, or third candidate location is unavailable.
[0182] In some embodiments, the processor (1720) may determine that the virtual block can represent a valid affine model in response to determining that at least one motion information at the first candidate location or the second candidate location is unavailable.
[0183] In some embodiments, one or more affine candidates may include one or more affine inherited candidates and one or more affine configured candidates, and the processor (1720) may further acquire one or more affine inherited candidates according to a first scanning rule and acquire one or more affine configured candidates according to a second scanning rule, wherein the second scanning rule is wholly or partially identical to the first scanning rule.
[0184] In some embodiments, the processor (1720) may further determine a second scanning rule based on at least one second scanning area, at least one second scanning distance, and a second scanning order, and may scan at least one second scanning area at each distance corresponding to the same block size as the current block.
[0185] In some embodiments, the processor (1720) may obtain two or three CPMVs for the current block based on three CPMVs of the virtual block using the same projection process used for inherited candidate derivation, which includes the following cases: the processor (1720) may obtain two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the first type of affine model represented by the virtual block onto the first type of affine model for the current block in response to determining that the virtual block represents a first type of affine model; and the processor (1720) may obtain two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the second type of affine model represented by the virtual block onto the second type of affine model for the current block in response to determining that the virtual block represents a second type of affine model; Alternatively, the processor (1720) may obtain two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the affine model represented by the virtual block onto an affine model of the type for the current block, where the type of the current block is a first type or a second type.
[0186] In step 1802, the processor (1720) may obtain one or more CPMVs for the current block based on one or more affine candidates.
[0187] FIG. 19 is a flowchart illustrating a method for pruning affine candidates according to an embodiment of the present disclosure.
[0188] In step 1901, the processor (1720) can calculate a first set of affine model parameters associated with one or more CPMVs of the first affine candidate.
[0189] In step 1902, the processor (1720) can calculate a second set of affine model parameters associated with one or more CPMVs of the second affine candidate.
[0190] In step 1903, the processor (1720) can perform a similarity check between a first affine candidate and a second affine candidate based on a first set of affine model parameters and a second set of affine model parameters.
[0191] In some embodiments, the processor (1720) may determine that a first affine candidate is similar to a second affine candidate, and in response to determining that a first set of affine model parameters is similar to a second set of affine model parameters, may prun one of the first affine candidate and the second affine candidate.
[0192] In some embodiments, the processor (1720) may determine that a first affine candidate is similar to a second affine candidate in response to determining that a plurality of differences are each smaller than a plurality of thresholds, wherein the plurality of differences include the difference between one parameter of a first set of affine model parameters and a corresponding parameter of a second set of affine model parameters.
[0193] In some embodiments, as shown in Table 1, a plurality of threshold values may be determined according to a first set of affine model parameters comparable to a second set of affine model parameters.
[0194] In some embodiments, a plurality of threshold values may be determined according to the size of the current block. For example, a plurality of threshold values may be determined according to the width or height of the current block as shown in Table 2, Table 3, or Table 4. As another example, a plurality of threshold values may be determined as a group of fixed values as shown in Table 5.
[0195] In some examples, the processor (1720) can calculate one or more of the first set of affine model parameters associated with one or more CPMVs of the first affine candidate according to the width and height of the current block, and can calculate one or more of the second set of affine model parameters associated with one or more CPMVs of the second affine candidate according to the width and height of the current block.
[0196] In some embodiments, a video coding device is provided. The device includes a processor (1720) and a memory (1740) configured to store instructions executable by the processor; wherein the processor is configured to perform the method illustrated in FIG. 18 when executing instructions.
[0197] In some other embodiments, a non-transient computer-readable storage medium in which a command is stored is provided. When the command is executed by a processor (1720), the command causes the processor to perform the method illustrated in FIG. 18.
[0198] In some embodiments, a video coding device is provided. The device includes a processor (1720) and a memory (1740) configured to store instructions executable by the processor; wherein the processor is configured to perform the method illustrated in FIG. 19 when executing instructions.
[0199] In some other embodiments, a non-transient computer-readable storage medium in which a command is stored is provided. When the command is executed by a processor (1720), the command causes the processor to perform the method illustrated in FIG. 19.
[0200] Other embodiments of the present disclosure will be apparent to those skilled in the art by considering the specification and practice of the disclosure disclosed herein. This application is intended to include any modification, use, or adaptation of the present disclosure, including deviations from the present disclosure that follow the general principles of the present disclosure and are within the known or customary practice of the art. The specification and embodiments are to be regarded merely as illustrative.
[0201] It will be understood that the present disclosure is not limited to the exact examples described above and illustrated in the attached drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
Claim 1 A video decoding method, the method comprising: obtaining one or more affine candidates from a plurality of non-adjacent neighbor blocks not adjacent to a current block; and obtaining one or more control point motion vectors (CPMV) for the current block based on one or more affine candidates, wherein the one or more affine candidates include one or more affine inherited candidates and one or more affine constructed candidates, and the method further comprising: obtaining one or more affine inherited candidates according to a first scanning rule; and obtaining one or more affine constructed candidates according to a second scanning rule, wherein the second scanning rule is wholly or partially identical to the first scanning rule. Claim 2 The method of claim 1, wherein the step of acquiring one or more affine candidates comprises: the step of acquiring one or more affine candidates according to a scanning rule. Claim 3 A method according to claim 2, further comprising the step of determining a scanning rule based on at least one scanning area, at least one scanning distance, and a scanning order. Claim 4 A method according to claim 3, further comprising the step of determining at least one scanning area according to at least one scanning distance. Claim 5 In claim 4, at least one scanning area includes a first scanning area and a second scanning area, wherein the first scanning area is determined according to a first maximum scanning distance representing the maximum number of blocks separated from a first side of the current block, and the second scanning area is determined according to a second maximum scanning distance representing the maximum number of blocks separated from a second side of the current block, and the first maximum scanning distance is the same as or different from the second maximum scanning distance. Claim 6 A method according to claim 5, further comprising the step of receiving a first maximum scanning distance and a second maximum scanning distance from a bitstream. Claim 7 A method according to claim 5, further comprising the step of predetermining a first maximum scanning distance or a second maximum scanning distance as a fixed value. Claim 8 A method according to claim 7, further comprising the step of stopping scanning in at least one scanning area in response to determining that the first maximum scanning distance or the second maximum scanning distance is equal to 4, in response to determining that a candidate list containing one or more affine candidates is full, or in response to determining that all non-adjacent neighbor blocks within the first maximum scanning distance and the second maximum scanning distance are scanned. Claim 9 The method of claim 3 further comprises the step of scanning from a first starting non-adjacent neighbor block along a scanning line parallel to the left of the current block, wherein the first starting non-adjacent block is a bottom block within a first scanning area, and the blocks within the first scanning area are at a first scanning distance from the left of the current block. Claim 10 The method of claim 3 further comprises the step of scanning from a third starting non-adjacent neighbor block along a scanning line parallel to the upper face of the current block, wherein the third starting non-adjacent block is a right block within a first scanning area, and the blocks within the first scanning area are at a first scanning distance from the upper face of the current block. Claim 11 A method according to claim 3, further comprising the step of positioning a non-adjacent neighbor block at a scanning location. Claim 12 A method according to claim 1, further comprising: a step of obtaining a first candidate location for a first affine candidate and a second candidate location for a second affine candidate based on a scanning rule; a step of determining a third candidate location for a third affine candidate based on the first and second candidate locations; a step of obtaining a virtual block based on the first candidate location, the second candidate location, and the third candidate location; a step of obtaining three CPMVs for the virtual block based on translational MVs at the first candidate location, the second candidate location, and the third candidate location; and a step of obtaining two or three CPMVs for the current block based on the three CPMVs of the virtual block using the same projection process used for deriving inherited candidates. Claim 13 In claim 12, the virtual block is a rectangular coding block, and the third candidate position is determined based on the vertical position of the first candidate position and the horizontal position of the second candidate position, or the method further comprises the step of: determining the vertical position of the third candidate position as the vertical position of the top-left point of the current block and determining the horizontal position of the third candidate position as the horizontal position of the top-left point of the current block in response to a determination that the first candidate position or the second candidate position is unavailable or in response to a determination that movement information at the first candidate position or the second candidate position is unavailable. Claim 14 The method of claim 12 further comprises the step of determining that the virtual block cannot represent a valid affine model in response to determining that motion information at a first, second, or third candidate location is unavailable, or further comprises the step of determining that the virtual block can represent a valid affine model in response to determining that at least one motion information at a first or second candidate location is unavailable. Claim 15 The method of claim 1 further comprises the steps of: determining a second scanning rule based on at least one second scanning area, at least one second scanning distance, and a second scanning order; and scanning at least one second scanning area at each distance corresponding to the same block size as the current block. Claim 16 In claim 12, the step of obtaining two or three CPMVs for the current block based on three CPMVs of the virtual block using the same projection process used for deriving inherited candidates comprises: a step of obtaining two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the affine model of the first type represented by the virtual block onto the affine model of the first type for the current block in response to determining that the virtual block represents an affine model of the first type; a step of obtaining two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the affine model of the second type represented by the virtual block onto the affine model of the second type for the current block in response to determining that the virtual block represents an affine model of the second type; A method comprising at least one of the steps of obtaining two or three CPMVs for the current block based on three CPMVs of the virtual block by projecting the affine model represented by the virtual block onto an affine model of the type for the current block, wherein the type of the current block is a first type or a second type. Claim 17 A video encoding method comprising: a step of obtaining one or more affine candidates from a plurality of non-adjacent neighbor blocks that are not adjacent to the current block; a step of obtaining one or more control point motion vectors (CPMV) for the current block based on the one or more affine candidates; wherein the one or more affine candidates include one or more affine inherited candidates and one or more affine configured candidates, and the method further comprises: a step of obtaining the one or more affine inherited candidates according to a first scanning rule; and a step of obtaining the one or more affine configured candidates according to a second scanning rule, wherein the second scanning rule is wholly or partially identical to the first scanning rule. Claim 18 A video coding device comprising: one or more processors; and a memory coupled to one or more processors and configured to store instructions executable by one or more processors, wherein the one or more processors are configured to perform a method according to any one of claims 1 to 17 upon execution of instructions. Claim 19 A non-transient computer-readable storage medium that stores a bitstream formed by an instruction that causes one or more computer processors to perform a video encoding method according to claim 17 when executed by one or more computer processors. Claim 20 A computer program stored on a computer-readable storage medium, wherein the computer program comprises a plurality of programs to be executed by a computing device having one or more processors, and the plurality of programs, when executed by one or more processors, cause the computing device to perform a method according to any one of claims 1 to 17. Claim 21 A method for storing a bitstream, comprising the step of performing a video encoding method according to claim 17 to generate a bitstream; and the step of storing the bitstream. Claim 22 delete Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete Claim 28 delete Claim 29 delete Claim 30 delete Claim 31 delete Claim 32 delete Claim 33 delete Claim 34 delete Claim 35 delete Claim 36 delete Claim 37 delete Claim 38 delete
Citation Information
Patent Citations
Motion vector prediction
US20200221116A1
Usage for history-based affine parameters
US20210266584A1