Candidate Derivation for Affine Merge Mode in Video Coding
By deriving affine merge candidates from non-adjacent neighboring blocks and performing similarity checks, the method enhances the accuracy and efficiency of video encoding and decoding processes in video coding standards.
Patent Information
- Application Number
- JP2024518556
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-24
- Filing Date
- 2022-09-21
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-09-21
AI Technical Summary
Existing video coding standards face challenges in accurately deriving affine merge candidates for the affine motion prediction mode, which affects the efficiency of video encoding and decoding processes.
The method involves obtaining affine candidates from non-adjacent neighboring blocks and calculating control point motion vectors (CPMVs) for the current block. Additionally, a similarity check is performed between different affine candidates to remove redundant ones, enhancing the derivation process.
This approach improves the accuracy and diversity of affine merge candidates, leading to more efficient video encoding and decoding processes while maintaining video quality.
Smart Images

Figure 0007697145000014 
Figure 0007697145000015 
Figure 0007697145000016
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the priority of U.S. Provisional Application No. 63 / 248,401, entitled "Candidate Derivation for Affine Merge Mode in Video Coding", filed on September 24, 2021, the entire content of which is incorporated by reference for all purposes.
[0002] This disclosure relates to video encoding and compression, and more particularly, but not limited to, methods and apparatuses for improving affine - merge candidate derivation for affine motion prediction mode in a video - encoding or decoding process.
Background Art
[0003] To compress video data, various video coding techniques may be used. Video coding is performed according to one or more video coding standards. For example, today, some well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which have been developed together by ISO / IEC MPEG and ITU-T VECG. AOMedia Video1 (AV1) has been developed by the Alliance for Open Media (AOM) as a successor to its predecessor VP9. Audio Video Coding (AVS) is another series of video compression standards, called digital audio and digital video compression standards, developed by the Audio and Video Coding Standards Workgroup of China. Most of the existing video coding standards are built on top of well-known hybrid video coding frameworks, that is, block-based prediction methods (e.g., inter prediction, intra prediction) are used to reduce the redundancy present in video images or sequences, and transform coding is used to condense the energy of the prediction error. An important goal of video coding techniques is to compress video data in a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality.
[0004] The first-generation AVS standard includes the domestic Chinese standard "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1), and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+). The first-generation AVS standard can provide about 50% bit-rate savings at the same perceived quality compared to the MPEG-2 standard. The video part of the AVS1 standard was published as a domestic Chinese standard in February 2006. The second-generation AVS standard includes a series of domestic Chinese standards "Information Technology, Efficient Multimedia Coding" (known as AVS2), mainly targeting the transmission of special HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. AVS2 was published as a domestic Chinese standard in May 2016. On the other hand, the video part of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for applications. The AVS3 standard is a new-generation video coding standard for UHD video applications aiming to exceed the coding efficiency of the latest international standard HEVC. In March 2019, at the 68th AVS meeting, the baseline of AVS3-P2 was completed, and the AVS3-P2 baseline brings about approximately 30% bit-rate savings compared to the HEVC standard. Currently, there is a reference software called High Performance Model (HPM), which is maintained and managed by the AVS group to show the reference implementation form of the AVS3 standard.
Summary of the Invention
[0005] The present disclosure provides an example of a technique related to improving the derivation of affine merge candidates for the affine motion prediction mode in a video encoding or decoding process.
[0006] According to a first aspect of the present disclosure, a method of video encoding is provided. The method may include obtaining one or more affine candidates from a plurality of non-adjacent neighboring blocks that are not adjacent to a current block. Further, the method may include obtaining one or more control point motion vectors (CPMVs) of the current block based on the one or more affine candidates.
[0007] According to a second aspect of the present disclosure, a method for removing affine candidates is provided. The method may include calculating a first set of affine model parameters associated with one or more CPMVs of a first affine candidate. Further, the method may include calculating a second set of affine model parameters associated with one or more CPMVs of a second affine candidate. Moreover, the method may include performing a similarity check between the first affine candidate and the second affine candidate based on the first set of affine model parameters and the second set of affine model parameters.
[0008] According to a third aspect of the present disclosure, an apparatus for video encoding is provided. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. Further, when the one or more processors execute the instructions, they are configured to implement the method according to the first aspect or the second aspect.
[0009] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to implement the method according to the first aspect or the second aspect is provided.
[0010] A more detailed description of the examples of the present disclosure will be given by referring to the specific examples illustrated in the accompanying drawings. On the premise that these drawings only depict some examples and are therefore not considered to limit the scope, the examples will be described and explained more specifically and in detail through the use of the accompanying drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 14A
Figure 14B
Figure 15A
Figure 15B
Figure 15C
Figure 15D
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
DETAILED DESCRIPTION OF THE INVENTION
[0012] Particular implementations are referred to in detail herein, and examples thereof are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to facilitate understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative forms may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented in many types of electronic devices having digital video capabilities.
[0013] References throughout this specification to "one embodiment", "an embodiment", "an example", "some embodiments", "some examples", or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are applicable to other embodiments as well, unless expressly specified otherwise.
[0014] Throughout the present disclosure, the terms "first", "second", "third", etc. are all used as a terminology system only for referring to related elements, such as devices, components, structures, steps, etc., and do not imply any spatial or chronological order unless otherwise specifically specified. For example, "the first device" and "the second device" may refer to two separately formed devices, or two parts, components, or operable states of the same device, and may be arbitrarily named.
[0015] The terms "module", "sub-module", "circuit", "sub-circuit", "circuit device", "sub-circuit device", "unit", or "sub-unit" may include a memory (shared, dedicated, or group) that stores code or instructions executable by one or more processors. A module may include one or more circuits regardless of the presence of the stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached, or may or may not be placed adjacent to each other.
[0016] As used herein, the terms "if" or "when" may be understood to mean "simultaneously with" or "in response to" depending on the context. These terms may not indicate that the related limitations or features are conditional or optional when they appear in the claims. For example, a method may include the steps of i) when or if condition X exists, function or action X' is performed, and ii) when or if condition Y exists, function or action Y' is performed. The method may be performed with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at different times during multiple executions of the method.
[0017] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a particular function.
[0018] FIG. 20 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 20, system 10 includes a source device 12 that generates and encodes video data that will be decoded later by a destination device 14. The source device 12 and the destination device 14 may include any of a variety of electronic devices, including a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, and the like. In some implementations, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.
[0019] In some embodiments, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium to enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be beneficial in facilitating communication from the source device 12 to the destination device 14.
[0020] In some other implementations, the encoded video data may be transmitted from the output interface 22 to the storage device 32. Thereafter, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disk, a digital versatile disk (DVD), a compact disk read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a wireless fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0021] As shown in FIG. 20, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include, for example, a video capture device such as a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a source such as a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As one example, if the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the implementations described in this application may generally be applicable to video encoding and may also be applied to wireless and / or wired applications.
[0022] Captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may further (or alternatively) be stored in the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.
[0023] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and may receive encoded video data via link 16. The encoded video data communicated via link 16 or provided to the storage device 32 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 during decoding of the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0024] In some implementations, the destination device 14 may include a display device 34, and the display device 34 may be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 may display the decoded video data to a user and may include any of various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0025] Video encoder 20 and video decoder 30 may operate according to dedicated or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extended versions of such standards. It should be understood that this application is not limited to a specific video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally assumed that the video encoder 20 of source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is generally assumed that the video decoder 30 of destination device 14 may be configured to decode video data according to any of these current or future standards.
[0026] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device may store software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and any of these may be integrated as part of a combined encoder / decoder (codec) in their respective devices.
[0027] Like HEVC, VVC is built on a block - based hybrid video coding framework. FIG. 1 is a block diagram illustrating a block - based video encoder according to some implementations of the present disclosure. In encoder 100, the input video signal is processed for each block called a coding unit (CU). Encoder 100 may be a video encoder 20 as shown in FIG. 20. In VTM - 1.0, the CU can be up to 128×128 pixels. However, unlike HEVC which divides blocks based only on the quadtree, in VVC, one coding tree unit (CTU) is divided into CUs to adapt to various local characteristics based on the quadtree / bi - tree / tri - tree. Additionally, the concept of multiple split unit types in HEVC is abolished, that is, the distinction between the CU, prediction unit (PU), and transform unit (TU) no longer exists in VVC. Instead, each CU is always used as a basic unit for both prediction and transform without further division. In multiple types of tree structures, one CTU is first divided by the quadtree structure. Then, the leaf nodes of each quadtree can be further divided by the bi - tree and tri - tree structures.
[0028] FIGS. 3A - 3E are schematic diagrams illustrating multiple types of tree partition modes according to some implementations of the present disclosure. FIGS. 3A - 3E show five partition types including 4 - way split (FIG. 3A), vertical 2 - way split (FIG. 3B), horizontal 2 - way split (FIG. 3C), vertical extended 3 - way split (FIG. 3D), and horizontal extended 3 - way split (FIG. 3E), respectively.
[0029] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") uses pixels from samples of already encoded neighboring blocks (referred to as reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter prediction" or "motion compensation prediction") uses pixels reconstructed from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference pictures are supported, one reference picture index is sent additionally, and the reference picture index is used to identify from which reference picture in the reference picture store the temporal prediction signal came.
[0030] After spatial and / or temporal prediction, the intra / inter mode decision circuitry 121 of the encoder 100 selects the best prediction mode, for example, based on a rate distortion optimization method. The block predictor 120 then subtracts from the current video block, and the resulting prediction residual is decorrelated using the transform circuitry 102 and quantization circuitry 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuitry 116 and inverse transformed by the inverse transform circuitry 118 to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, the reconstructed CU is placed in the reference picture store of the picture buffer 117, and loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), may be applied to the reconstructed CU before it is used to encode future video blocks. All of the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy encoding unit 106 so that they are further compressed and packed to form the bitstream to form the output video bitstream 114.
[0031] For example, deblocking filters are available in AVC, HEVC, and the current version of VVC. In HEVC, an additional loop filter called SAO is defined to further improve encoding efficiency. In the current version of the VVC standard, yet another loop filter called ALF is being actively investigated and is likely to be included in the final standard.
[0032] These loop filter operations are optional. Performing these operations helps to improve encoding efficiency and visual quality. They may further be turned off as a decision made by the encoder 100 to save computational complexity.
[0033] Intra prediction is typically based on non-filtered reconstructed pixels, while inter prediction, it should be noted, is based on filtered reconstructed pixels when these filter options are turned on by the encoder 100.
[0034] FIG. 2 is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section resident in the encoder 100 of FIG. 1. The block-based video decoder 200 may be a video decoder 30 as shown in FIG. 20. In decoder 200, the incoming video bitstream 201 is first decoded through entropy decoding 202 to derive quantization coefficient levels and prediction-related information. The quantization coefficient levels are then processed through inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residuals. The block predictor mechanism executed by the intra / inter mode selector 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. An adder 214 is used to sum the reconstructed prediction residuals from the inverse transform 206 and the predictive output generated by the block predictor mechanism, thereby obtaining a set of non-filtered reconstructed pixels.
[0035] The reconstructed block may further pass through the in-loop filter 209 and is then stored in the picture buffer 213 that functions as a reference picture store. The reconstructed video in the picture buffer 213 is sent to drive a display device and may also be used to predict future video blocks. In a situation where the in-loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.
[0036] In the current VVC and AVS3 standards, the motion information of the current coding block is copied from spatially or temporally adjacent blocks specified by the merge candidate index, or obtained by explicit signaling of motion estimation. The main objective of this disclosure is to improve the accuracy of the motion vectors for the affine merge mode by improving the method for deriving affine merge candidates. To facilitate the description of this disclosure, the existing affine merge mode design in the VVC standard is used as an example to illustrate the proposed concept. Although the existing affine mode design in the VVC standard is used as an example throughout this disclosure, it should be noted that the proposed techniques are also applicable to different designs of affine motion prediction modes or other coding tools having the same or similar design spirit for those skilled in the art of modern video coding technology.
[0037] Affine model In HEVC, only the translation motion model is applied for motion compensation prediction. On the other hand, in the real world, there are many types of motions such as, for example, zoom-in / zoom-out, rotation, viewpoint motion, and other irregular motions. In VVC and AVS3, affine motion compensation prediction is applied by signaling one flag for each inter-coded block to indicate whether the translation motion model or the affine motion model is applied for inter prediction. In the current VVC and AVS3 designs, two affine modes are supported for one affine coding block, including the four-parameter affine mode and the six-parameter affine mode.
[0038] The 4-parameter affine model has parameters including two parameters for translational movement in the horizontal and vertical directions respectively, one parameter for zoom movement, and one parameter for rotational movement in both directions. In this model, the horizontal zoom parameter is equal to the vertical zoom parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. To achieve better adaptation of the motion vector and the affine parameters, these affine parameters will be derived from two MVs (also called control point motion vectors (CPMVs)) located at the upper left and upper right corners of the current block. As shown in FIGS. 4A to 4B, the affine motion field of the block is represented by two CPMVs (V0, V1). Based on the motion of the control points, the motion field (v x ,v y ) of one affine-coded block is
Equation
[0039] The 6-parameter affine mode has parameters including two parameters for translational movement in the horizontal and vertical directions respectively, two parameters for zoom movement and rotational movement in the horizontal direction respectively, and another two parameters for zoom movement and rotational movement in the vertical direction respectively. The 6-parameter affine motion model is coded with three CPMVs. As shown in FIG. 5, the three control points of one 6-parameter affine block are located at the upper left, upper right, and lower left corners of the block. The motion at the upper left control point is related to translational motion, the motion at the upper right control point is related to rotational and zoom motion in the horizontal direction, and the motion at the lower left control point is related to rotational and zoom motion in the vertical direction. Compared with the 4-parameter affine motion model, the 6-parameter rotational and zoom motion in the horizontal direction may not be the same as the rotational and zoom motion in the vertical direction. Assuming that (V0, V1, V2) are the MVs at the upper left, upper right, and lower left corners of the current block in FIG. 5, the sub-block (v x ,vy ) Each motion vector is derived using three MVs at the control points, [Number] as derived below.
[0040] Affine Merge Mode In the Affine Merge Mode, the CPMV of the current block is not explicitly signaled and is derived from adjacent blocks. Specifically, in this mode, the motion information of spatially adjacent blocks is used to generate the CPMV of the current block. The Affine Merge Mode candidate list has a limited size. For example, in the current VVC design, there may be up to five candidates. The encoder may evaluate and select the best candidate index based on the rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. The Affine Merge candidates can be determined in three ways. In the first way, the Affine Merge candidates may be inherited from adjacent affine-coded blocks. In the second way, the Affine Merge candidates may be constructed from translational MVs from adjacent blocks. In the third way, zero MVs are used as Affine Merge candidates.
[0041] In the case of the inheritance method, there may be up to two candidates. The candidates, if available, are obtained from the adjacent block to the lower left of the current block (e.g., as shown in Figure 6, the scanning order is from A0 to A1), and from the adjacent block to the upper right of the current block (e.g., as shown in Figure 6, the scanning order is from B0 to B2).
[0042] In the case of the construction method, the candidates are a combination of translational MVs of adjacent ones that can be generated in two steps.
[0043] Step 1: Obtain four translational MVs including MV1, MV2, MV3, and MV4 from the available neighboring ones. MV1: The MV from one of the three adjacent blocks near the upper left corner of the current block. As shown in Figure 7, the scan order is B2, B3, and A2. MV2: The MV from one of the ones from two adjacent blocks near the upper right corner of the current block. As shown in Figure 7, the scan order is B1 and B0. MV3: The MV from one of the ones from two adjacent blocks near the lower left corner of the current block. As shown in Figure 7, the scan order is A1 and A0. MV4: The MV from the temporally arranged blocks of the adjacent block near the lower right corner of the current block. As shown in this figure, the adjacent block is T.
[0044] Step 2: Derive combinations based on the four translational MVs from Step 1. Combination 1: MV1, MV2, MV3, Combination 2: MV1, MV2, MV4, Combination 3: MV1, MV3, MV4, Combination 4: MV2, MV3, MV4, Combination 5: MV1, MV2, Combination 6: MV1, MV3.
[0045] After satisfying the inheritance and construction candidates, when the merge candidate list is not full, a zero MV is inserted at the end of the list.
[0046] For the current video standards VVC and AVS, for each of the inheritance candidates and construction candidates, as shown in Figures 6 and 7, only the adjacent neighboring blocks are used to derive the affine merge candidates of the current block. To increase the diversity of the merge candidates and further investigate the spatial correlation relationship, it is easy to expand the coverage of the neighboring blocks from the adjacent areas to the non - adjacent areas.
[0047] In the present disclosure, the candidate derivation process for affine merge mode is extended by using not only adjacent neighboring blocks but also non - adjacent neighboring blocks. A detailed method may be outlined in three aspects, including affine merge candidate removal, a non - adjacent neighbor - based derivation process for affine inheritance merge candidates, and a non - adjacent neighbor - based derivation process for affine construction merge candidates.
[0048] Affine Merge Candidate Removal The affine merge candidate list of a typical video coding standard usually has a limited size, so candidate removal is an essential process to remove redundant candidates. This removal process is necessary for both affine merge inheritance candidates and construction candidates. As explained in the introduction section, the CPMV of the current block is not directly used for affine motion compensation. Instead, the CPMV needs to be converted to translational MVs at the locations of each sub - block within the current block. The conversion process is carried out by following a general affine model as shown below,
Equation
[0049] For the six - parameter affine model, three CPMVs called V0, V1, and V2 are available. Then, the six model parameters a, b, c, d, e, and f are,
Number
[0050] In the case of a 4-parameter affine model, when the CPMVs at the upper left corner and upper right corner, called V0 and V1, are available, the six parameters a, b, c, d, e, and f are
Number
[0051] In the case of a 4-parameter affine model, when the CPMVs at the upper left corner and lower left corner, called V0 and V2, are available, the six parameters a, b, c, d, e, and f are
Number
[0052] In the above equations (4), (5), and (6), w and h represent the width and height of the current block, respectively.
[0053] When two merge candidate sets of CPMVs are compared for redundancy checking, it is proposed to check the similarity of the six affine model parameters. Therefore, the candidate removal process can be implemented in two steps.
[0054] In step 1, for two candidate sets of CPMVs, the corresponding affine model parameters for each candidate set are derived. More specifically, the two candidate sets of CPMVs may be represented by two sets of affine model parameters, for example, (a1, b1, c1, d1, e1, f1) and (a2, b2, c2, d2, e2, f2).
[0055] In step 2, a similarity check is performed between two sets of affine model parameters based on one or more predefined thresholds. In one embodiment, when the absolute values of (a1 - a2), (b1 - b2), (c1 - c2), (d1 - d2), (e1 - e2), and (f1 - f2) are all below a positive threshold, such as the value 1, the two candidates are considered similar, and one of these can be removed / deleted and not added to the merge candidate list.
[0056] In some embodiments, the division or right shift operation in step 1 may be removed to simplify the calculations in the CPMV removal process.
[0057] Specifically, the model parameters c, d, e, and f may be calculated without being divided by the width w and height h of the current block. For example, taking the above equation (4) as an example, the approximate model parameters c’, d’, e’, and f’ may be calculated as in the following equation (7).
Number
[0058] In the case where only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters, and the model parameters are dependent on the width or height of the current block. In this case, the model parameters may be transformed taking into account the effects of the width and height. For example, in the case of equation (5), the approximate model parameters c’, d’, e’, and f’ may be calculated based on the following equation (8). In the case of equation (6), the approximate model parameters c’, d’, e’, and f’ may be calculated based on the following equation (9).
Number
[0059] In step 2 above, a threshold value is required to evaluate the similarity between two candidate sets of CPMV. There may be multiple ways to define the threshold value. In one embodiment, the threshold value may be defined for each comparable parameter. Table 1 is an example in this embodiment, showing the threshold values defined for each comparable model parameter. In another embodiment, the threshold value may be defined considering the size of the current encoded block. Table 2 is an example in this embodiment, showing the threshold values defined by the size of the current encoded block.
[0060] [Table 1]
[0061] [Table 2]
[0062] In another embodiment, the threshold value may be defined considering the weight or height of the current block. Tables 3 and 4 are examples in this embodiment. Table 3 shows the threshold values defined by the width of the current encoded block, and Table 4 shows the threshold values defined by the height of the current encoded block.
[0063] [Table 3]
[0064] [Table 4]
[0065] In another embodiment, the threshold value may be defined as a group of fixed values. In another embodiment, the threshold value may be defined in any combination of the above embodiments. In one example, the threshold value may be defined considering different parameters and the weights and heights of the current block. Table 5 is an example in this embodiment and shows the threshold value defined by the height of the current coding block. It should be noted that in any of the above proposed embodiments, the comparable parameters may represent any parameter defined by any of the equations from Equation (4) to Equation (9) if necessary.
[0066]
Table 5
[0067] The benefits of using the transformed affine model parameters for candidate redundancy checks are to create an integrated similarity check process for candidates with different affine model types. For example, one merge candidate may use a 6-parameter affine model with 3 CPMVs, while another candidate may use a 4-parameter affine model with 2 CPMVs, and to consider the different effects of each CPMV in the merge candidate when deriving the target MV for each sub-block, and to provide the importance of the similarity of two affine merge candidates with respect to the width and height of the current block.
[0068] Adjacent neighbor-based derivation process for affine inheritance merge candidates In the case of inheritance merge candidates, the adjacent neighbor-based derivation process may be performed in three steps. Step 1 is related to candidate scanning. Step 2 is related to CPMV projection. Step 3 is related to candidate removal.
[0069] In step 1, non - adjacent neighboring blocks are scanned and selected in the following manner.
[0070] Scanning area and distance In some examples, non - adjacent neighboring blocks may be scanned from the area to the left and the area above the current encoded block. The scanning distance may be defined as the number of encoded blocks from the scanning position to the left side or the top - most side of the current encoded block.
[0071] As shown in FIG. 8, multiple lines of non - adjacent neighboring blocks may be scanned on the left or above the current encoded block. The distances shown in FIG. 8 represent the number of encoded blocks from each candidate position to the left side or the top - most side of the current block. For example, the area with "distance 2" to the left of the current block indicates that the candidate neighboring blocks within this area are 2 blocks away from the current block. Similar indications apply to other scanning areas with different distances.
[0072] In one or more embodiments, the non - adjacent neighboring blocks at each distance may have the same block size as the current encoded block, as shown in FIG. 13A. As shown in FIG. 13A, the left non - adjacent neighboring block 1301 and the upper non - adjacent neighboring block 1302 have the same size as the current block 1303. In some embodiments, the non - adjacent neighboring blocks at each distance may have a different block size from the current encoded block, as shown in FIG. 13B. The neighboring block 1304 is a neighboring block adjacent to the current block 1303. As shown in FIG. 13B, the left non - adjacent neighboring block 1305 and the upper non - adjacent neighboring block 1306 have the same size as the current block 1307. The neighboring block 1308 is a neighboring block adjacent to the current block 1307.
[0073] Note that when non - adjacent neighboring blocks at each distance have the same block size as the current encoded block, the value of the block size is adaptively changed according to the segmentation granularity for each different area in the image. Note that when non - adjacent neighboring blocks at each distance have a different block size from the current encoded block, the value of the block size may be predefined as a constant value, such as 4×4, 8×8, or 16×16.
[0074] Based on the defined scan distance, the overall size of the scan area to the left or above the current encoded block may be determined by configurable distance values. In one or more embodiments, the maximum scan distances on the left and above may use the same value or different values. FIG. 13 shows an example where the maximum distances on both the left and above share the same value of 2. The maximum scan distance value may be determined on the encoder side and signaled in the bitstream. Alternatively, the maximum scan distance value may be predefined as a fixed value, such as a value of 2 or 4. When the maximum scan distance is predefined as a value of 4, this indicates that the scan process has ended, either when the candidate list is full or when all non - adjacent neighboring blocks having a distance of at most 4 have been scanned, whichever occurs first.
[0075] In one or more embodiments, the starting and ending neighboring blocks within each scan area at a particular distance may be in any position.
[0076] In some embodiments, for the left scan area, the adjacent block at the start may be the adjacent lower left block of the adjacent block at the start of the adjacent scan area having a smaller distance. For example, as shown in FIG. 8, the adjacent block at the start of the scan area of "distance 2" to the left of the current block is the adjacent lower left adjacent block of the adjacent block at the start of the scan area of "distance 1". The adjacent block at the end may be the adjacent left block of the adjacent block at the end of the upper scan area having a smaller distance. For example, as shown in FIG. 8, the adjacent block at the end of the scan area of "distance 2" to the left of the current block is the adjacent left adjacent block of the adjacent block at the end of the scan area of "distance 1" above the current block.
[0077] Similarly, for the upper scan area, the adjacent block at the start may be the adjacent upper right block of the adjacent block at the start of the adjacent scan area having a smaller distance. The adjacent block at the end may be the adjacent upper left block of the adjacent block at the end of the adjacent scan area having a smaller distance.
[0078] Scanning order When adjacent blocks are scanned in non-adjacent areas, the selection of the adjacent blocks to be scanned may be determined according to a specific order or / and rules.
[0079] In some embodiments, the left area may be scanned first, and then the scanning of the upper area may follow. As shown in FIG. 8, three lines (e.g., from distance 1 to distance 3) of the non-adjacent area on the left may be scanned first, and then the scanning of three lines of the non-adjacent area above the current block may follow.
[0080] In some embodiments, alternatively, the left area and the upper area may be scanned. For example, as shown in FIG. 8, the left scan area having "Distance 1" is scanned first, and then the scan of the upper area having "Distance 1" follows.
[0081] For scan areas on the same side (e.g., the left or upper area), the scan order is from the area with a smaller distance to the area with a larger distance. This order may be flexibly combined with other embodiments of the scan order. For example, alternatively, the left and upper areas may be scanned, and the order of the areas on the same side is scheduled to be from a smaller distance to a larger distance.
[0082] Within each scan area at a specific distance, the scan order may be defined. In one embodiment, for the left scan area, the scan may start from the bottom adjacent block to the top adjacent block. For the upper scan area, the scan may start from the right block to the left block.
[0083] Scan completed In the case of inheritance merge candidates, adjacent blocks encoded in affine mode are defined as eligible candidates. In some embodiments, the scan processes may be performed interactively. For example, the scan performed in a specific area at a specific distance may be stopped the moment the first X eligible candidates are identified, where X is a predefined positive value. For example, as shown in FIG. 8, the scan in the left scan area having Distance 1 may be stopped when the first one or more eligible candidates are identified. Then, the next iteration of the scan process is started by targeting another scan area as defined by the predefined scan order / rules.
[0084] In some embodiments, the scan process may be performed continuously. For example, a scan performed in a particular area at a particular distance may be stopped when all covered adjacent blocks have been scanned and no more eligible candidates are identified, or when the maximum allowable number of candidates has been reached.
[0085] During the candidate scan process, by following the proposed scan method above, non-adjacent adjacent blocks of each candidate are determined and scanned. As a simpler implementation, the non-adjacent adjacent blocks of each candidate may be indicated or located at a particular scan position. By following the proposed method above, when a particular scan area and distance are determined, the scan position may be appropriately determined based on the following method.
[0086] In one method, as shown in FIG. 15A, lower left and upper right positions are used for the non-adjacent adjacent blocks above and to the left, respectively.
[0087] In another method, as shown in FIG. 15B, a lower right position is used for the non-adjacent adjacent blocks both above and to the left.
[0088] In another method, as shown in FIG. 15C, a lower left position is used for the non-adjacent adjacent blocks both above and to the left.
[0089] In another method, as shown in FIG. 15D, an upper right position is used for the non-adjacent adjacent blocks both above and to the left.
[0090] As a simpler illustration, in FIGS. 15A-15D, each non-adjacent adjacent block is assumed to have the same block size as the current block. Without loss of generality, this illustration may be easily extended to non-adjacent adjacent blocks having different block sizes.
[0091] Furthermore, in step 2, the same process of CPMV projection as that used in the current AVS and VVC standards may be utilized. In this CPMV projection process, the current block is assumed to share the same affine model with the selected adjacent blocks, and then the coordinates of two or three corner pixels (for example, if the current block uses a 4-parameter model, two coordinates (the top-left pixel / sample location and the top-right pixel / sample location) are used, and if the current block uses a 6-parameter model, three coordinates (the top-left pixel / sample location, the top-right pixel / sample location, and the bottom-left pixel / sample location) are used) are plugged into equation (1) or (2), which depends on whether the adjacent blocks are encoded with a 4-parameter affine model or a 6-parameter affine model to generate two or three CPMVs.
[0092] In step 3, any eligible candidates identified in step 1 and transformed in step 2 may pass a similarity check against all existing candidates already in the merge candidate list. Details of the similarity check have already been described in the section on affine merge candidate removal above. If it is found that the new eligible candidate is similar to any existing candidate in the candidate list, this new eligible candidate is deleted / removed.
[0093] Adjacent non-adjacent neighbor-based derivation process for affine construction merge candidates In the case of deriving inheritance merge candidates, one adjacent block is identified at a time, and this single adjacent block needs to be encoded in affine mode and may contain two or three CPMVs. In the case of deriving construction merge candidates, two or three adjacent blocks may be identified at a time, and each identified adjacent block does not need to be encoded in affine mode, and only one translational MV is taken out from this block.
[0094] FIG. 9 presents an example in which construction affine merge candidates can be derived by using non - adjacent neighboring blocks. In FIG. 9, A, B, and C are the geographical positions of three non - adjacent neighboring blocks. The virtual coded block is formed by using the position of A as the upper - left corner, the position of B as the upper - right corner, and the position of C as the lower - left corner. When considering the virtual CU as an affine - coded block, the MVs at the positions of A', B', and C may be derived by following Equation (3), and the model parameters (a, b, c, d, e, f) may be calculated with the translational MVs at the positions of A, B, and C. Once derived, the MVs at the positions of A', B', and C may be used as the three CPMVs of the current block, and the existing process (used in the AVS and VVC standards) for generating construction affine merge candidates may be used.
[0095] For construction merge candidates, the non - adjacent neighbor - based derivation process may be performed in five steps. The non - adjacent neighbor - based derivation process may be performed in five steps in a device such as an encoder or a decoder. Step 1 is related to candidate scanning. Step 2 is related to affine model determination. Step 3 is related to CPMV projection. Step 4 is related to candidate generation. Also, Step 5 is related to candidate removal. In Step 1, non - adjacent neighboring blocks may be scanned and selected in the following way.
[0096] Scan Area and Distance In some embodiments, to maintain rectangular coded blocks, the scan process is performed only for two non - adjacent neighboring blocks. The third non - adjacent neighboring block may depend on the horizontal and vertical positions of the first and second non - adjacent neighboring blocks.
[0097] In some embodiments, as shown in FIG. 9, the scan process is performed only for positions B and C. The position of A may be uniquely determined by the horizontal position of C and the vertical position of B. In this case, the scan area and distance may be defined according to a specific scan direction.
[0098] In some embodiments, the scan direction may be perpendicular to the side of the current block. One example is shown in FIG. 10, where the scan area is defined as one line of a continuous motion field to the left or above the current block. The scan distance is defined as the number of motion fields from the scan position to the side of the current block. Note that the size of the motion field may depend on the maximum granularity of the applicable video coding standard. In the example shown in FIG. 10, it is assumed that the size of the motion field is aligned with the current VVC standard and set to 4x4.
[0099] In some embodiments, the scan direction may be parallel to the side of the current block. One example is shown in FIG. 11, where the scan area is defined as one line of a continuous coded block to the left or above the current block.
[0100] In some embodiments, the scan direction may be a combination of perpendicular and parallel scans to the side of the current block. One example is shown in FIG. 12. As shown in FIG. 12, the scan direction may also be a combination of parallel and diagonal. The scan at position B starts from left to right and then starts moving diagonally to the blocks on the left and above. The scan at position B will repeat as shown in FIG. 12. Similarly, the scan at position C starts from top to bottom and then starts moving diagonally to the blocks on the left and above. The scan at position C will repeat as shown in FIG. 12.
[0101] Scan Order In some embodiments, the scan order may be defined as going from positions with smaller distances to positions with larger distances for the current encoding block. This order may be applied to the case of a right-angle scan.
[0102] In some embodiments, the scan order may be defined as a fixed pattern. This fixed-pattern scan order may be used for candidate positions having similar distances. One example is the case of a parallel scan. In one example, as in the example shown in FIG. 11, the scan order may be defined as the direction from top to bottom for the left scan area and as the direction from left to right for the above-mentioned scan area.
[0103] In the case of a combined scan method, the scan order may be a combination of a fixed pattern and distance dependence, as in the example shown in FIG. 12.
[0104] Scan end In the case of constructing merge candidates, since only translational MVs are required for eligible candidates, they do not need to be affine-encoded.
[0105] Depending on the number of candidates required, the scan process may end when the first X eligible candidates are identified, where X is a positive value.
[0106] As shown in FIG. 9, three corners named A, B, and C are required to form a virtual coding block. As a simpler implementation form, the scan process in step 1 may be performed only to identify non-adjacent neighboring blocks at corners B and C, while the coordinates of A may be accurately determined by taking the horizontal coordinate of C and the vertical coordinate of B. In this way, the formed virtual coding block is restricted to be rectangular. When the B or C point is unavailable, such as being outside the boundary for example, or when the motion information in the non-adjacent neighboring block corresponding to B or C is unavailable, the horizontal or vertical coordinate of C may be defined as the horizontal or vertical coordinate of the upper left point of the current block respectively.
[0107] For unification, the method proposed for deriving inheritance merge candidates to define the scan area and distance, scan order, and scan end may be fully or partially reused to derive construction merge candidates. In one or more embodiments, the same method defined for inheritance merge candidate scan, including but not limited to the scan area and distance, scan order, and scan end, may be fully reused for construction merge candidate scan.
[0108] In some embodiments, the same method defined for inheritance merge candidate scan may be partially reused for construction merge candidate scan. FIG. 16 shows an example in this case. In FIG. 16, the block size for each non-adjacent neighboring block is the same as the current block and is defined in the same way as the inheritance candidate scan, but since the scan at each distance is limited to only one block, the overall process is a simplified version.
[0109] In step 2, the translational MV at the position of the selected candidate after step 1 may be evaluated and an appropriate affine model may be determined. As a simpler illustration, without loss of generality, FIG. 9 is reused as an example again.
[0110] Due to factors such as hardware constraints, implementation complexity, and different reference indices, the scan process may end before a sufficient number of candidates are identified. For example, the motion information of the motion field in one or more of the selected candidates after step 1 may be unavailable.
[0111] If the motion information of all three candidates is available, the corresponding virtual coding block represents a 6-parameter affine model. If the motion information of one of the three candidates is unavailable, the corresponding virtual coding block represents a 4-parameter affine model. If the motion information of two or more of the three candidates is unavailable, the corresponding virtual coding block may not be able to represent a valid affine model.
[0112] In some embodiments, if the motion information at the upper left corner of the virtual coding block, such as corner A in FIG. 9, is unavailable, or if the motion information at both the upper right corner, such as corner B in FIG. 9, and the lower left corner, such as corner C in FIG. 9, is unavailable, the virtual block is set to be invalid and may not be able to represent a valid model. Then, steps 3 and 4 may be skipped for the current iteration.
[0113] In some embodiments, if the upper right corner, such as corner B in FIG. 9, or the lower left corner, such as corner C in FIG. 9, is unavailable, but not both, the virtual block may represent a valid 4-parameter affine model.
[0114] In step 3, if the virtual coding block can represent a valid affine model, the same projection process used for the inheritance merge candidate may be used.
[0115] In one or more embodiments, the same projection process used for inheritance merge candidates may be used. In this case, the four-parameter model represented by the virtual coding block from step 2 is projected onto the four-parameter model for the current block, and the six-parameter model represented by the virtual coding block from step 2 is projected onto the six-parameter model for the current block.
[0116] In some embodiments, the affine model represented by the virtual coding block from step 2 is always projected onto the four-parameter model or the six-parameter model for the current block.
[0117] According to equations (5) and (6), note that there are two types of four-parameter affine models. Type A is the one where the upper left corner CPMV and the upper right corner CPMV, called V0 and V1, are available, and type B is the one where the upper left corner CPMV and the lower left corner CPMV, called V0 and V2, are available.
[0118] In one or more embodiments, the type of the projected four-parameter affine model is the same as the type of the four-parameter affine model represented by the virtual coding block. For example, the affine model represented by the virtual coding block from step 2 is a four-parameter affine model of type A or B, and then the projected affine model for the current block is also of type A or B, respectively.
[0119] In some embodiments, the four-parameter affine model represented by the virtual coding block from step 2 is always projected onto the same type of four-parameter model for the current block. For example, type A or B of the four-parameter affine model represented by the virtual coding block is always projected onto the four-parameter affine model of type A.
[0120] In step 4, based on the projected CPMV after step 3, in one example, the same candidate generation process used in the current VVC or AVS standard may be used. In another embodiment, the temporal motion vectors used in the candidate generation process for the current VVC or AVS standard may not be used for the non-adjacent neighboring block-based derivation method. When the temporal motion vectors are not used, this indicates that the generated combinations do not include any temporal motion vectors.
[0121] In step 5, any newly generated candidates after step 4 may pass a similarity check against all existing candidates already in the merge candidate list. Details of the similarity check have already been described in the section on affine merge candidate removal. If it is found that a newly generated candidate is similar to any existing candidate in the candidate list, this newly generated candidate is deleted or removed.
[0122] FIG. 17 shows a computing environment (or computing device) 1710 coupled to a user interface 1760. The computing environment 1710 can be part of a data processing server. In some embodiments, the computing device 1710 can implement any of the various methods or processes (such as encoding / decoding methods or processes) as described above according to the various examples of the present disclosure. The computing environment 1710 may include a processor 1720, a memory 1740, and an I / O interface 1750.
[0123] Processor 1720 typically controls the overall operation of computing environment 1710, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1720 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Additionally, processor 1720 may include one or more modules to facilitate the interaction between processor 1720 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, or the like.
[0124] Memory 1740 is configured to store various types of data to support the operation of computing environment 1710. Memory 1740 may include pre-determined software 1742. Examples of such data include instructions for any application or method operating in computing environment 1710, video data sets, image data, and the like. Memory 1740 may be implemented by using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disks, or a combination thereof.
[0125] I / O interface 1750 provides an interface between processor 1720 and peripheral interface modules such as a keyboard, click wheel, buttons, and the like. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. I / O interface 1750 may be coupled to an encoder and a decoder.
[0126] In some embodiments, to implement the above method, a non-transitory computer-readable storage medium having a plurality of programs, such as those included in memory 1740, executable by a processor 1720 in computing environment 1710 is also provided. For example, the non-transitory computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0127] The non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, and when the plurality of programs are executed by one or more processors, the computing device is caused to implement the above method for motion prediction.
[0128] In some embodiments, computing environment 1710 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to implement the above method.
[0129] FIG. 18 is a flowchart illustrating a method for video encoding according to an example of the present disclosure.
[0130] In step 1801, processor 1720 may obtain one or more affine candidates from a plurality of non-adjacent neighboring blocks that are not adjacent to the current block or CU.
[0131] In some examples, the plurality of non-adjacent neighboring blocks may include non-adjacent encoded blocks as shown in FIGS. 11-12, FIGS. 13A-13B, FIGS. 14A-14B, FIGS. 15A-15D, and FIG. 16.
[0132] In some examples, the processor 1720 may obtain one or more affinity candidates according to a scanning rule.
[0133] In some examples, the scanning rule may be determined based on at least one scan area, at least one scan distance, and a scan order.
[0134] In some examples, at least one scan distance indicates the number of blocks away from the side of the current block.
[0135] In some examples, one of the plurality of non - adjacent neighboring blocks at one of the at least one scan distances may have the same size as the current block, as shown in FIG. 13A, or may have a different size from the current block, as shown in FIG. 13B.
[0136] In some examples, the at least one scan area may include a first scan area and a second scan area. The first scan area is determined according to a first maximum scan distance indicating the maximum number of blocks away from the first side of the current block, and the second scan area is determined according to a second maximum scan distance indicating the maximum number of blocks away from the second side of the current block. The first maximum scan distance is the same as or different from the second maximum scan distance. In some examples, the first maximum scan distance or the second maximum scan distance may be set as a fixed value such as 3, 4, etc.
[0137] For example, the first scan area may be the left area of the current block 1303, and the first maximum scan distance is 3 blocks away from the left side of the current block 1303. That is, block 1301 is at the first maximum scan distance, that is, 3 blocks away from the left side of the current block 1303. Further, the second scan area may be the upper area of the current block 1303, and the second maximum scan distance is 3 blocks away from above or the upper side of the current block 1303. That is, block 1302 is at the second maximum scan distance, that is, 3 blocks away from the upper / upper side of the current block 1303.
[0138] In some examples, the encoder may signal the first maximum scan distance and the second maximum scan distance in the bitstream that is to be sent to the decoder.
[0139] In some examples, in response to determining that the first or second maximum scan distance is equal to a fixed value, and in response to determining that the candidate list is full or that all non-adjacent neighboring blocks within the first or second maximum scan distance have been scanned, the processor 1720 may stop scanning at least one scan area as the scan is complete.
[0140] In some examples, the processor 1720 may scan a plurality of non-adjacent neighboring blocks within the first scan area to obtain one or more non-adjacent neighboring blocks encoded in affine mode, and determine one or more non-adjacent neighboring blocks encoded in affine mode as one or more affine candidates.
[0141] In some examples, the processor 1720 may scan from a non-adjacent neighboring block of the first start along a scanning line parallel to the left side of the current block. The non-adjacent neighboring block of the first start is the bottommost block in the first scan area, and the blocks in the first scan area are at a first scan distance away from the left side of the current block, such as D2 in FIG. 14A.
[0142] In some examples, the non-adjacent neighboring block of the first start may be at the bottom and left of the non-adjacent neighboring block of the second start in the second scan area. The blocks in the second scan area may be at a second scan distance away from the left side of the current block, such as D1 in FIG. 14A as shown. In some other examples, the non-adjacent neighboring block of the first start may be on the left of the non-adjacent neighboring block of the second start in the second scan area, and the blocks in the second scan area may be at a second scan distance away from the left side of the current block as shown in FIG. 14B.
[0143] In some examples, the processor 1720 may scan from a non-adjacent neighboring block of the third start along a scanning line parallel to the upper side of the current block. The non-adjacent neighboring block of the third start may be the right block in the first scan area, and the blocks in the first scan area may be at a first scan distance away from the upper side of the current block, such as D2 in FIG. 14A.
[0144] In some examples, the non-adjacent block adjacent to the third start may be above and to the right of the non-adjacent block adjacent to the fourth start within the second scan area, and the blocks within the second scan area may be at a second scan distance away from the upper side of the current block, such as D1 in FIG. 14A, as shown in FIG. 14A. In some other examples, the non-adjacent block adjacent to the third start may be to the right of the non-adjacent block adjacent to the fourth start within the second scan area, and the blocks within the second scan area may be at a second scan distance away from the upper side of the current block, as shown in FIG. 14B.
[0145] In some examples, the processor 1720 can place non-adjacent neighboring blocks at scan positions. For example, as a simpler implementation form, each non-adjacent neighboring block of each candidate may be indicated or located at a specific scan position.
[0146] In some examples, the scan positions may include the lower left position of the non-adjacent block within the second scan area above the current block as shown in FIG. 15A, the upper right position of the non-adjacent block within the first scan area to the left of the current block as shown in FIG. 15A, the lower right position of the non-adjacent block within the first scan area or the second scan area as shown in FIG. 15B, the lower left position of the non-adjacent block within the first scan area or the second scan area as shown in FIG. 15C, and the upper right position of the non-adjacent block within the first scan area or the second scan area as shown in FIG. 15D.
[0147] In some examples, the processor 1720 obtains a first candidate position for a first affine candidate and a second candidate position for a second affine candidate based on scan rules, determines a third candidate position for a third affine candidate based on the first and second candidate positions, obtains a virtual block based on the first candidate position, the second candidate position, and the third candidate position, obtains three CPMVs of the virtual block based on the translational MV at the first candidate position, the second candidate position, and the third candidate position, and may obtain two or three CPMVs of the current block based on the three CPMVs of the virtual block by using the same projection process used for inheritance candidate derivation.
[0148] In some examples, the virtual block may be a rectangular coded block, and the third candidate position may be determined based on the vertical position of the first candidate position and the horizontal position of the second candidate position. For example, the virtual block may be a virtual block including positions A, B, and C as shown in FIG. 9.
[0149] In some examples, in response to determining that the first candidate position or the second candidate position is unavailable, or in response to determining that the motion information at the first candidate position or the second candidate position is unavailable, the processor 1720 may determine the vertical position of the third candidate position as the vertical position of the upper left point of the current block and determine the horizontal position of the third candidate position as the horizontal position of the upper left point of the current block.
[0150] In some examples, in response to determining that the motion information at the first, second, or third candidate position is unavailable, the processor 1720 may determine that the virtual block cannot represent a valid affine model.
[0151] In some examples, in response to determining that at least one motion information at the first or second candidate position is available, the processor 1720 may determine that the virtual block can represent a valid affine model.
[0152] In some examples, one or more affinity candidates may include one or more affinity inheritance candidates and one or more affinity construction candidates, and the processor 1720 may further obtain one or more affinity inheritance candidates according to a first scan rule, and obtain one or more affinity construction candidates according to a second scan rule, and the second scan rule may be completely or partially the same as the first scan rule.
[0153] In some examples, the processor 1720 may further determine a second scan rule based on at least one second scan area, at least one second scan distance, and a second scan order, and scan at least one second scan area at each distance equal to the same block size as the current block.
[0154] In one example, in response to determining that the virtual block represents a first type of affinity model, the processor 1720 projects the first type of affinity model represented by the virtual block onto the first type of affinity model for the current block, by which the processor 1720 may obtain two or three CPMVs of the current block based on three CPMVs of the virtual block, including using the same projection process used for inheritance candidate derivation, or the processor 1720, in response to determining that the virtual block represents a second type of affinity model, projects the second type of affinity model represented by the virtual block onto the second type of affinity model for the current block, by which the processor 1720 may obtain two or three CPMVs of the current block based on three CPMVs of the virtual block, or the processor 1720 projects the affinity model represented by the virtual block onto the type of affinity model for the current block, by which the processor 1720 may obtain two or three CPMVs of the current block based on three CPMVs of the virtual block, and the type of the current block is the first type or the second type.
[0155] In step 1802, the processor 1720 may obtain one or more CPMVs of the current block based on one or more affinity candidates.
[0156] FIG. 19 is a flowchart illustrating a method for removing affinity candidates according to an example of the present disclosure.
[0157] In step 1901, the processor 1720 may calculate a first set of affinity model parameters associated with one or more CPMVs of the first affinity candidate.
[0158] In step 1902, the processor 1720 may calculate a second set of affinity model parameters associated with one or more CPMVs of the second affinity candidate.
[0159] In step 1903, the processor 1720 may perform a similarity check between the first affinity candidate and the second affinity candidate based on the first set of affinity model parameters and the second set of affinity model parameters.
[0160] In some examples, in response to determining that the first set of affinity model parameters is similar to the second set of affinity model parameters, the processor 1720 may determine that the first affinity candidate is similar to the second affinity candidate and remove one of the first affinity candidate and the second affinity candidate.
[0161] In some examples, in response to determining that a plurality of differences are each smaller than a plurality of thresholds, the processor 1720 may determine that the first affinity candidate is similar to the second affinity candidate, where the plurality of differences includes the difference between one parameter of the first set of affinity model parameters and one corresponding parameter of the second set of affinity model parameters.
[0162] In some examples, the plurality of thresholds may be determined according to the first set of affinity model parameters comparable to the second set of affinity model parameters, as shown in Table 1.
[0163] In some examples, the plurality of thresholds may be determined according to the size of the current block. For example, the plurality of thresholds may be determined according to the width or height of the current block, as shown in Tables 2, 3, or 4. As another example, the plurality of thresholds may be determined as a group of fixed values, as shown in Table 5.
[0164] In some examples, the processor 1720 may calculate one or more affine model parameters from a first set of affine model parameters associated with one or more CPMVs of a first affine candidate according to the width and height of the current block, and calculate one or more affine model parameters from a second set of affine model parameters associated with one or more CPMVs of a second affine candidate according to the width and height of the current block.
[0165] In some examples, an apparatus for video encoding is provided. The apparatus includes a processor 1720 and a memory 1740 configured to store instructions executable by the processor, and the processor is configured to perform a method as illustrated in FIG. 18 when the instructions are executed.
[0166] In some other examples, a non-transitory computer-readable storage medium storing instructions is provided. When the instructions are executed by the processor 1720, the instructions cause the processor to perform a method as illustrated in FIG. 18.
[0167] In some examples, an apparatus for video encoding is provided. The apparatus includes a processor 1720 and a memory 1740 configured to store instructions executable by the processor, and the processor is configured to perform a method as illustrated in FIG. 19 when the instructions are executed.
[0168] In some other examples, a non-transitory computer-readable storage medium storing instructions is provided. When the instructions are executed by the processor 1720, the instructions cause the processor to perform a method as illustrated in FIG. 19.
[0169] Other examples of the present disclosure will be apparent to those skilled in the art from a consideration of the specification and practice of the disclosure as disclosed herein. This application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles thereof and including such departures from the present disclosure as come within known or customary practice in the art. The specification and examples are intended to be considered as exemplary only.
[0170] It is to be understood that the present disclosure is not limited to the exact examples described above and illustrated in the accompanying drawings, and that various modifications and changes can be made without departing from the scope of the present disclosure.
Claims
1. A method for video encoding, comprising: obtaining one or more affine candidates from a plurality of non - adjacent neighboring blocks that are not adjacent to the current block; and obtaining one or more control point motion vectors (CPMVs) of the current block based on the one or more affine candidates; The step of obtaining one or more affine candidates includes: obtaining the one or more affine candidates according to a scanning rule; The scanning rule is obtained based on at least one scanning area, at least one scanning distance, and a scanning order. A method for video encoding.
2. further comprising determining the at least one scanning area according to the at least one scanning distance. The method according to claim 1.
3. The at least one scanning area includes a first scanning area and a second scanning area. The first scanning area is determined according to a first maximum scanning distance indicating the maximum number of blocks away from a first side of the current block, and the second scanning area is determined according to a second maximum scanning distance indicating the maximum number of blocks away from a second side of the current block. The first maximum scanning distance is the same as or different from the second maximum scanning distance. The method according to claim 2.
4. further comprising receiving the first maximum scanning distance and the second maximum scanning distance from a bitstream. The method according to claim 3.
5. Determining in advance the first maximum scan distance or the second maximum scan distance as a fixed value The method according to claim 3, further comprising this step.
6. Responsive to determining that the first maximum scan distance or the second maximum scan distance is equal to 4, responsive to determining that a candidate list including the one or more affinity candidates is full, or responsive to determining that all non-adjacent neighboring blocks within the first maximum scan distance and the second maximum scan distance have been scanned, stopping the scan of the at least one scan area The method according to claim 5, further comprising this step.
7. Scanning from a first starting non-adjacent neighboring block along a scanning line parallel to the left side of the current block, wherein the first starting non-adjacent block is the bottommost block within a first scan area, and the blocks within the first scan area are at a first scan distance away from the left side of the current block The method according to claim 1, further comprising this step.
8. Scanning from a third starting non-adjacent neighboring block along a scanning line parallel to the upper side of the current block, wherein the third starting non-adjacent block is the rightmost block within a first scan area, and the blocks within the first scan area are at a first scan distance away from the upper side of the current block The method according to claim 1, further comprising this step.
9. Placing a non-adjacent neighboring block at a scan position The method according to claim 1, further comprising this step.
10. A step of obtaining a first candidate position for a first affine candidate and a second candidate position for a second affine candidate based on the scan rule; A step of determining a third candidate position for a third affine candidate based on the first and second candidate positions; A step of obtaining a virtual block based on the first candidate position, the second candidate position, and the third candidate position; A step of obtaining three CPMVs of the virtual block based on translational MV at the first candidate position, the second candidate position, and the third candidate position; A step of obtaining two or three CPMVs of the current block based on the three CPMVs of the virtual block by using the same projection process used for inherited candidate derivation The method according to claim 1, further comprising.
11. The virtual block is a rectangular coded block, and the third candidate position is determined based on the vertical position of the first candidate position and the horizontal position of the second candidate position, or, The method is In response to determining that the first candidate position or the second candidate position is unavailable, or in response to determining that the motion information at the first candidate position or the second candidate position is unavailable, determining the vertical position of the third candidate position as the vertical position of the upper left point of the current block, and determining the horizontal position of the third candidate position as the horizontal position of the upper left point of the current block. The method according to claim 10, further comprising.
12. In response to determining that the motion information at the first, second, or third candidate position is unavailable, determining that the virtual block cannot represent a valid affine model, or, In response to determining that at least one motion information at the first or second candidate position is available, determining that the virtual block can represent a valid affinity model The method according to claim 10, further comprising **Claim 13** wherein the one or more affinity candidates include one or more affinity inheritance candidates and one or more affinity construction candidates, and the method comprises obtaining the one or more affinity inheritance candidates according to a first scan rule; and obtaining the one or more affinity construction candidates according to a second scan rule, wherein the second scan rule is the same as the first scan rule, either completely or partially The method according to claim 1, further comprising **Claim 14** determining the second scan rule based on at least one second scan area, at least one second scan distance, and a second scan order; and scanning the at least one second scan area at distances equal to the same block size as the current block The method according to claim 13, further comprising **Claim 15** obtaining two or three CPMVs of the current block based on the three CPMVs of the virtual block by using the same projection process used for inheritance candidate derivation, obtaining the two or three CPMVs of the current block based on the three CPMVs of the virtual block by projecting the first type of affinity model represented by the virtual block onto the first type of affinity model for the current block in response to determining that the virtual block represents the first type of affinity model In response to determining that the virtual block represents a second type of affinity model, obtaining the two or three CPMVs of the current block based on the three CPMVs of the virtual block by projecting the second type of affinity model represented by the virtual block onto the second type of affinity model of the current block, or Obtaining the two or three CPMVs of the current block based on the three CPMVs of the virtual block by projecting the affinity model represented by the virtual block onto the affinity model of the type of the current block, wherein the type of the current block is the first type or the second type The method according to claim 10, further comprising at least one of the above.
16. An apparatus for video encoding, comprising: One or more processors; A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors and a decoded bitstream And comprising The one or more processors are configured to perform the method according to any one of claims 1 to 15 using the bitstream when executing the instructions. An apparatus for video encoding.
17. A non-transitory computer-readable storage medium storing computer-executable instructions and a decoded bitstream, wherein when the computer-executable instructions are executed by one or more computer processors, the method according to any one of claims 1 to 15 is caused to be performed by the one or more computer processors. A non-transitory computer-readable storage medium.
18. A computer program for execution by a computing device comprising one or more processors, wherein when the computer program is executed by the one or more processors, the computing device is caused to perform the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Motion Vector Prediction
JP2020523853A