Method, apparatus, and computer program for video coding
The local warp motion delta mode in video coding addresses inefficiencies in handling complex local motions, improving compression performance and reducing complexity in video coding technologies like AV1 and VVC.
Patent Information
- Application Number
- JP2024545802
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-03
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-25
AI Technical Summary
Existing video coding technologies, such as AV1 and VVC, face challenges in efficiently handling complex local motion patterns, leading to suboptimal compression performance and increased computational complexity.
The proposed method involves using a local warp motion delta mode for video coding, where a warp model is generated based on motion vectors of adjacent blocks, and a base model is selected to decode blocks, enhancing coding efficiency and reducing complexity.
This approach improves coding efficiency by better handling complex local motions, reducing computational overhead, and enhancing compression performance.
Smart Images

Figure 2025523734000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 388,886, filed Jul. 13, 2022, and U.S. Patent Application No. 17 / 980,125, filed Nov. 3, 2022, and incorporates the disclosures of each of them herein in their entireties by reference.
[0002] The present disclosure generally relates to advanced image and video coding techniques, and more specifically to coding and / or decoding with local warp motion modes.
Background Art
[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video - on - demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project are based on the previous research efforts of alliance members. Individual contributors had started experimental technical platforms years earlier. Xiph / Mozilla's Daala had already made its code public in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor was announced on August 11, 2015. Based on the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The alliance released the AV1 bitstream specification on March 28, 2018, along with a reference software - based encoder and decoder. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, a verified version 1.0.0 with the Errata 1 specification was released. The AV1 bitstream specification includes a reference video codec.
[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) issued the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). They also investigated the potential need for standardization of future video coding technologies that could be significantly more performant than HEVC in terms of compression capabilities. In October 2017, they jointly solicited proposals for video compression with capabilities beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team or Joint Video Expert Team) meeting. As a result of this meeting, JVET officially started the standardization of next-generation video coding beyond HEVC. This new standard was named Versatile Video Coding (VVC).
Summary of the Invention
[0005] Embodiments of the present disclosure relate to a video coding method, apparatus, and computer-readable medium for improving the local warp motion delta mode.
[0006] According to one aspect of one or more embodiments, a video coding method executed by at least one processor includes obtaining a coded video bitstream, obtaining a plurality of blocks from the video bitstream, the plurality of blocks including a current block and one or more adjacent blocks, determining that a warp delta mode is used to predict the current block based on syntax elements in the video bitstream, determining positions and motion vectors of the one or more adjacent blocks, generating a first warp model of the current block based on the motion vectors of the one or more adjacent blocks, selecting a warp model from among the first warp model of the current block and a second warp model associated with one of the one or more adjacent blocks as a base model for a coding mode, and decoding the plurality of blocks in the coding mode based on the base model.
[0007] According to another aspect of one or more embodiments, an apparatus / device and a non-transitory computer-readable medium consistent with the video coding method are also provided.
[0008] Further embodiments are described in the following description, and will become apparent, in part, from the description and / or can be realized by the practice of the presented embodiments of the disclosure.
Brief Description of the Drawings
[0009] To more clearly explain the technical solutions of the embodiments of this disclosure, the following briefly introduces the accompanying drawings for explaining the embodiments. The accompanying drawings here, which are incorporated into the specification and form a part of this specification, show embodiments that conform to this disclosure and are used to explain the principles of this disclosure together with this specification. Obviously, the accompanying drawings in the following description only show some embodiments. Also, as will be understood by those skilled in the art, the aspects of the embodiments may be combined together or implemented alone.
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
DETAILED DESCRIPTION OF THE INVENTION
[0010] The following detailed description of example embodiments refers to the accompanying drawings. The same reference numerals in different figures may identify the same or similar elements.
[0011] The above disclosure provides examples and explanations and is not intended to be exhaustive or to limit implementation to exactly the disclosed forms. Changes and modifications are possible in light of the above disclosure or may be obtained from practice of the implementation. Also, one or more mechanisms or components of one embodiment may be incorporated into or combined with another embodiment (or one or more mechanisms of another embodiment). Further, in the flowcharts and descriptions of operations provided below, it is to be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be (at least partially) executed simultaneously, and the order of one or more operations may be interchanged.
[0012] It will become apparent that the systems and / or methods described herein can be implemented in various forms of hardware, software, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0013] Even if certain combinations of features are recited in the claims and / or disclosed in the specification, those combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not specifically disclosed in the specification. Each of the dependent claims listed below may directly depend on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.
[0014] The features of the proposals described below can be used separately or combined in any order. Also, embodiments can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0015] Any element, act, or instruction used herein is so expressly described Unless otherwise indicated, it should not be construed as important or essential. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar words are used. Also, as used herein, the terms "have," "possess," "possessing," "include," "including," or the like are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least in part based on" unless expressly stated otherwise. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0016] Here, aspects will be described with reference to flowchart diagrams and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It is to be understood that each block of the flowchart diagrams and / or block diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, can be implemented by computer-readable program instructions.
[0017] Next, refer to FIG. 1, which is a functional block diagram of a networked computer environment showing a video coding system 100 (hereinafter, "the system") for encoding and decoding video data according to an exemplary embodiment such as those described herein. It should be understood that FIG. 1 merely provides an example of one implementation and does not imply any limitation with respect to an environment in which multiple different embodiments may be implemented. Numerous changes to the illustrated environment can be made based on design and implementation requirements.
[0018] System 100 may include computer 102 and server computer 114. Computer 102 may communicate with server computer 114 via communication network 110 (hereinafter, “network”). Computer 102 may include processor 104 and software program 108 stored in data storage device 106, and is enabled to interface with a user and communicate with server computer 114. As will be described later with reference to FIG. 10, computer 102 may include internal component 800A and external component 900A respectively, and server computer 114 may include internal component 800B and external component 900B respectively. Computer 102 may be, for example, a mobile device, a telephone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running a program, accessing a network, and accessing a database.
[0019] Server computer 114 may also operate in a cloud computing service model, such as, for example, software - as - a - service (SaaS), platform - as - a - service (PaaS), or infrastructure - as - a - service (IaaS), as will be described later with reference to FIGS. 11 and 12. Server computer 114 may also be placed within a cloud computing deployment model, such as, for example, a private cloud, a community cloud, a public cloud, or a hybrid cloud.
[0020] Server computer 114, which can be used to encode video data, is enabled to execute a video encoding program 116 (hereinafter referred to as "the program") that can interact with database 112. The video encoding program method will be described in more detail later with respect to FIG. 4. In one embodiment, computer 102 can operate as an input device including a user interface, while program 116 can run mainly on server computer 114. In an alternative embodiment, program 116 may run mainly on one or more computers 102, and server computer 114 may be used for processing and storing data used by program 116. Note that program 116 may be a stand-alone program or may be integrated into a larger video encoding program.
[0021] However, it should be noted that the processing for program 116 may be shared between computer 102 and server computer 114 in any ratio in some examples. In another embodiment, program 116 may operate on two or more computers, server computers, or some combination of computers and server computers, for example, multiple computers 102 that communicate with a single server computer 114 across network 110. In another embodiment, for example, program 116 may operate on multiple server computers 114 that communicate with multiple client computers across network 110. Alternatively, the program may operate on a network server that communicates with servers and multiple client computers across the network.
[0022] Network 110 may include a wired connection, a wireless connection, a fiber optic connection, or some combination thereof. Generally, Network 110 can be any combination of connections and protocols that support communication between Computer 102 and Server Computer 114. Network 110 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as the public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, a fiber optic-based network, or the like, and / or a combination of these or other types of networks.
[0023] The number and configuration of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks configured differently than those shown in FIG. 1. Further, two or more of the devices shown in FIG. 1 may be implemented within a single device, or alternatively, a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) of System 100 may perform one or more functions described as being performed by another set of devices of System 100.
[0024] As described above, AV1 is an open video coding format developed as a successor to VP9 and designed for video transmission over the Internet. As shown in Figure 2A, VP9 starts from the 64×64 level and goes down to the 4×4 level, using a 4-way partition tree, with some additional constraints for blocks 8×8 and below (as shown in the upper half of Figure 2A). Note that the partition designated as R can be referred to as recursive in that the same partition tree can be repeated at a lower scale until the partition reaches the lowest 4×4 level. As shown in Figure 2B, AV1 not only extends the partition tree to a 10-way structure but also increases the maximum size (referred to as a superblock in VP9 / AV1 terms) to start from 128×128. Note that this can include 4:1 / 1:4 rectangular partitions that did not exist in VP9 (shown in Figure 2A). None of these rectangular partitions can be further subdivided. Additionally, AV1 adds more flexibility in the use of partitions below the 8×8 level in that 2×2 chroma inter prediction is possible in certain cases.
[0025] In HEVC, in order to adapt to various local characteristics, a coding tree unit (CTU) can be divided into coding units (CUs) by using a quadtree (QT) structure called a coding tree. A decision on whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU partitioning type. Inside one PU, the same prediction process is applied, and related information can be transmitted to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU partitioning type, the CU can be partitioned into transform units (TUs) according to another QT structure such as the coding tree for that CU. One of the important features of the HEVC structure is that it has multiple partitioning concepts including CUs, PUs, and TUs. In HEVC, a CU or TU is only square in shape, but a PU can be square or rectangular for an inter-prediction block. One coding block can be further divided into four square sub-blocks, and a transform is performed on each sub-block (i.e., TU). Each TU can be further recursively divided into smaller TUs (using quadtree partitioning), which is called a residual quadtree (RQT). At the picture boundary, HEVC adopts implicit quadtree partitioning so that the block maintains quadtree partitioning until its size conforms to the picture boundary.
[0026] As shown in FIG. 3A, the coding tree unit (CTU) is first partitioned by the QT structure, and the QT leaf nodes are further partitioned by the binary tree (BT) structure to form a quadtree + binary tree (QTBT) block structure. The QTBT block structure eliminates the concept of multi-partition types. That is, the QTBT structure eliminates the separation of the concepts of CU, PU, and TU and supports more flexibility for the CU partition shape. In the QTBT block structure, the CU can have a square or rectangular shape. There are two types of BT splits: symmetric horizontal split and symmetric vertical split. The BT leaf node is a CU, and its segmentation is used for prediction and conversion processing without any further partitioning. What this means is that in the QTBT coding block structure, the CU, PU, and TU have the same block size. In the joint exploration model (JEM) used by JVET, the CU can be composed of coding blocks (CBs) of different color components. For example, in the case of a prediction (P) slice and a binary (B) slice in the 4:2:0 chroma format, one CU can include one luma CB and two chroma CBs. The CU can also be composed of a single-component CB. For example, in the case of an I slice, one CU includes only one luma CB or only two chroma CBs.
[0027] In the QTBT partitioning method, the parameters include, but are not limited to, the CTU size (i.e., the QT root node size similar to the concept in HEVC), the minimum allowable QT leaf node size (i.e., MinQTSize), the maximum allowable BT root node size (i.e., MaxBTSize), the maximum allowable BT depth (i.e., MaxBTDepth), and the minimum allowable BT leaf node size (i.e., MinBTSize).
[0028] For example, the QTBT partitioning method can be as follows. The CTU size can be set to 128×128 luma samples having two corresponding chroma sample 64×64 blocks, MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4×4, and MaxBTDepth is set to 4. QT partitioning can be first applied to the CTU to generate QT leaf nodes. The QT leaf nodes can have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf QT node is 128×128, it will not be further split by the BT because its size exceeds MaxBTSize (i.e., 64×64). In some embodiments, the leaf QT node can be further partitioned by the BT. Thus, the QT leaf node is also the root node of the BT and has a BT depth of zero. When the BT depth reaches MaxBTDepth (i.e., 4), no further splitting is considered. When a BT node has a width equal to MinBTSize (i.e., 4), no further horizontal splitting is considered. Similarly, when a BT node has a height equal to MinBTSize, no further vertical splitting is considered. The leaf nodes of the BT are further processed by prediction and transformation processing without further partitioning. In JEM, for example, the maximum CTU size is 256×256 luma samples.
[0029] FIG. 3A shows an example of block partitioning of the QTBT structure, and FIG. 3B shows the corresponding tree representation. Solid lines indicate QT splits, and dotted lines indicate BT splits. At each split (i.e., non-leaf) node of the BT, one flag can be signaled to indicate which split type (i.e., horizontal or vertical) is used. For example, as shown in FIG. 3B, 0 indicates a horizontal split and 1 indicates a vertical split. In the case of QT splitting, since the QT split always splits the block both horizontally and vertically to generate four sub-blocks of equal size, there is no need to indicate the split type.
[0030] The QTBT partitioning method supports the flexibility for luma and chroma to have separate QTBT structures. Currently, in P slices and B slices, the luma CTB and chroma CTB within one CTU share the same QTBT structure. However, in I slices, the luma CTB is partitioned into CUs by the QTBT structure, and the chroma CTB is partitioned into chroma CUs by a different QTBT structure. What this means is that the CUs within an I slice are composed of coding blocks of the luma component or coding blocks of two chroma components, while the CUs within a P slice or B slice are composed of coding blocks of all three color components.
[0031] In HEVC, in order to reduce memory access for motion compensation, the inter prediction of small blocks is restricted. As a result, bi-prediction is not supported for 4×8 and 8×4 blocks, and inter prediction is not supported for 4×4 blocks. In QTBT implemented in JEM-7.0, these restrictions are removed.
[0032] Figure 4 shows the multi-type tree (MTT) structure in VCC, which, in addition to QTBT, further adds (a) vertical center-side triple-tree partitioning and (b) horizontal center-side triple-tree partitioning. The advantages of the triple-tree partitioning shown in Figure 4 include, but are not limited to, complementing quad-tree partitioning and binary-tree partitioning. While quad-trees and binary-trees always split along the block center, triple-tree partitioning can capture objects located at the block center. Also, the width and height of the proposed triple-tree partitions are always powers of two, and therefore no additional transformation is required. The two-level tree design is mainly motivated by complexity reduction. Theoretically, the complexity of traversing a tree is T D where T represents the number of split types and D represents the depth of the tree.
[0033] In merge mode, the motion information derived implicitly is directly used for generating the prediction samples of the current CU. Merge mode with motion vector differences (MMVD) has been introduced in VVC. To specify whether the MMVD mode is used for that CU, the MMVD flag is signaled immediately after the skip flag and the merge flag are sent. In MMVD, after a merge candidate is selected, it can be further refined by the signaled MVD information. The MVD information may include a merge candidate flag, an index defining the magnitude of the motion, and an index for indicating the direction of the motion. In merge mode, one of the first two merge candidate flags in the merge list is selected and used as the MV basis. The merge candidate flag is signaled to specify which flag is used.
[0034] The distance index defines the magnitude information of the motion and indicates a predetermined offset from the starting point. FIG. 5 shows the MMDV search points of two reference frames according to some embodiments. As shown in FIG. 5, an offset can be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predetermined offset is defined in Table 1 below.
Table 1
[0035] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent one of four directions as shown in Table 2 below. Note that the meaning of the sign of the MVD can vary according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where both lists point to the same side of the current picture (i.e., both POCs of the two references are greater than the POC of the current picture, or both POCs of the two references are less than the POC of the current picture), the signs in Table 2 define the signs of the MV offsets added to the starting MV. When the starting MV is a bidirectional prediction MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the POC difference in List 0 (L0) is greater than the POC difference in List 1 (L1), the signs in Table 1 define the signs of the MV offsets added to the L0 MV component of the starting MV, and the sign for the L1 MV has the opposite value. If the POC difference in L1 is greater than that in L0, the signs in Table 2 define the signs of the MV offsets added to the L1 MV component of the starting MV, and the sign for the L0 MV has the opposite value.
[0036] The MVD is scaled according to the POC differences in each direction. If the POC differences in both lists are the same, no scaling is required. If the POC difference in L0 is greater than that in L1, the MVD of L1 is scaled. If the POC difference in L1 is greater than that in L0, the MVD of L0 is similarly scaled. When the starting MV is unidirectionally predicted, the MVD is added to the available MV
Table 2
[0037] In VVC, in addition to the normal unidirectional and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional MVD signaling can be applied. In the symmetric MVD mode, motion information including both reference picture indices of L0 and L1 and the MVD of L1 can be derived (not signaled).
[0038] The decoding process of the symmetric MVD mode is as follows. First, at the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. For example, if mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. If the closest reference picture in L0 and the closest reference picture in L1 form a forward and backward pair of reference pictures, or a backward and forward pair of reference pictures, BiDirPredFlag is set to 1, and both the L0 reference picture and the L1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. Second, at the CU level, when the CU is bi-predicted coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true (e.g., equal to 1), only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of L0 and L1 are set equal to the pair of reference pictures, respectively. Finally, MVD1 is set equal to (-MVD0).
[0039] In AV1, for each block coded in an inter-frame, when the mode of the current block is an inter-coding mode rather than a skip mode, another flag is signaled to indicate whether a single reference mode or a composite reference mode is used for the current block. In the single reference mode, a predicted block can be generated by one motion vector. In the composite reference mode, a predicted block is generated by the weighted average of two predicted blocks derived from two motion vectors. The modes that can be signaled in the single reference case are detailed in Table 3 below.
Table 3
[0040] The modes that can be signaled in the case of composite references are detailed in Table 4 below. [Table 4]
[0041] AV1 enables 1 / 8 pixel motion vector accuracy (or precision). Using the syntax, the motion vector difference in reference frame L0 or L1 can be signaled as follows. The syntax mv_joint specifies which components of the motion vector difference are non-zero. A syntax mv_joint value of 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, a value of 1 indicates that there is a non-zero MVD only along the horizontal direction, a value of 2 indicates that there is a non-zero MVD only along the vertical direction, and a value of 3 indicates that there are non-zero MVDs along both the horizontal and vertical directions. The syntax mv_sign specifies whether the motion vector difference is positive or negative. The syntax mv_class specifies the class of the motion vector difference. As shown in Table 5 below, the higher the class, the larger the magnitude of the motion vector difference. The syntax Hhmv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude of each MV class. The syntax mv_fr specifies the first two fractional bits of the motion vector difference. The syntax mv_hp specifies the third fractional bit of the motion vector difference. [Table 5]
[0042] For the NEW_NEARMV mode and the NEAR_NEWMV mode (shown in Table 4), the accuracy of the MVD depends on the associated class and the size of the MVD. First, fractional MVDs are allowed only when the size of the MVD is 1 pixel or less. Second, when the value of the associated MV class is MV_CLASS_1 or greater, only one MVD value is allowed, and the MVD values for each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). Table 6 shows the allowable MVD values for each MV class.
Table 6
[0043] In some embodiments, when the current block is coded using the NEW_NEARMV mode or the NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class. When the current block is not coded using the NEW_NEARMV mode or the NEAR_NEWMV mode, a different context is used to signal mv_joint or mv_class.
[0044] To indicate whether the MVDs for two reference lists are signaled together, a new inter-coding mode (i.e., JOINT_NEWMV) can be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs for reference L0 and reference L1 are signaled together. Thus, only one MVD named joint_mvd can be signaled and sent to the decoder, and the delta MVs for reference L0 and reference L1 can be derived from joint_mvd. The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, and GLOBAL_GLOBALMV mode. It can be assumed that no additional context is added.
[0045] When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for reference L0 or reference L1 based on the POC distance. Specifically, let the distance between reference frame L0 and the current frame be represented as td0, and the distance between reference frame L1 and the current frame be represented as td1. If td0 is greater than or equal to td1, joint_mvd is directly used for reference L0, and the mvd for reference L1 is derived from joint_mvd by the following formula (1):
Number
[0046] If td1 is greater than or equal to td0, joint_mvd is directly used for reference L1, and the mvd for reference L0 is derived from joint_mvd by the following formula (2):
Number
[0047] A new inter-coding mode (i.e., AMVDMV) can be added for the single-reference case. When the AMVDMV mode is selected, it indicates that AMVD is applied to the signal MVD. Under the JOINT_NEWMV mode, a flag such as amvd_flag can be added to indicate whether AMVD is applied to the joint MVD coding mode. When the adaptive MVD resolution is applied to the joint MVD coding mode, it is called joint AMVD coding, where the MVDs for two reference frames are signaled together, and the accuracy of the MVD is implicitly determined by the magnitude of the MVD. The MVDs for two (or more) reference frames are signaled together and MVD coding is applied.
[0048] In the adaptive motion vector resolution (AMVR) first proposed in CWG-C012, a total of seven MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, the AVM encoder searches all the supported accuracy values and signals the best accuracy to the decoder. To reduce the encoder execution time, two accuracy sets are supported. Each accuracy set contains four predetermined accuracies. At the frame level, the accuracy set is adaptively selected based on the value of the maximum accuracy of that frame. Similar to AV1, the maximum accuracy is signaled in the frame header. Table 7 summarizes the accuracy values supported based on the frame-level maximum accuracy.
Table 7
[0049] In AVM software (similar to AV1), there is a frame-level flag indicating whether the MV of a frame includes sub-pel accuracy. AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the accuracy of a block is lower than the maximum accuracy, the motion model and interpolation filter are not signaled. When the accuracy of a block is lower than the maximum accuracy, the motion mode is estimated as translational motion, and the interpolation filter is estimated as the REGULAR interpolation filter. Similarly, when the accuracy of a block is either 4-pel or 8-pel, the inter-intra mode is not signaled and is assumed to be 0.
[0050] Motion compensation typically assumes a translational motion model between a reference block and a target block. However, warp motion (distorted motion) utilizes an affine model. The affine motion model is given by Equation (3):
Equation
[0051] Here, [x, y] are the coordinates of the original pixel, and [x’, y’] are the distorted coordinates of the reference block. According to Equation (3), up to six parameters are required to define warp motion, where a3 and b3 define the translational MV, a1 and b2 define the scaling along the MV, and a2 and b1 define the rotation.
[0052] In global warp motion compensation, for each inter-reference frame, global motion information including the global motion type and several motion parameters is signaled. Table 9 lists the global motion type and the number of related parameters.
Table 8
[0053] After signaling the reference frame index, if global motion is selected, the global motion type and parameters associated with a given reference frame are used for the current coding block.
[0054] In local warp motion compensation, local warp motion is allowed for an inter-coding block when the following conditions are met. First, the current block must use single-reference prediction. The width or height of the coding block must be 8 or more. Finally, at least one of the adjacent neighboring blocks must use the same reference frame as the current block.
[0055] When local warp motion is used for the current block, based on the MVs of the current block and its adjacent neighboring blocks, the affine model parameters are estimated by minimizing the mean squared difference between the reference projection and the modeled projection. To estimate the parameters of local warp motion, when an adjacent block uses the same reference frame as the current block, a projection sample pair of the central sample in the adjacent block and its corresponding sample in the reference frame is obtained. Subsequently, three additional samples are created by shifting the central position by a quarter sample in one or both dimensions. These additional samples can also be considered as projection sample pairs to ensure the stability of the model parameter estimation process.
[0056] The MVs of the adjacent blocks used to derive the motion parameters are referred to as motion samples. The motion samples are selected from adjacent blocks that use the same reference frame as the current block. Note that the warp motion prediction mode is only enabled for blocks using a single reference frame.
[0057] FIG. 6 shows exemplary motion samples used to derive block model parameters using local warp motion prediction. As shown in FIG. 6, the MVs of adjacent blocks B0, B1, and B2 are referred to as MV0, MV1, and MV2, respectively. The current block is predicted using one-sided prediction with reference frame Ref0. Adjacent block B0 is predicted using composite prediction with reference frames Ref0 and Ref1. Adjacent block B1 is predicted using one-sided prediction with reference frame Ref0. Adjacent block B2 is predicted using composite prediction with reference frames Ref0 and Ref2. The motion vector MV0 of B0 Ref0 , the motion vector MV1 of B1 Ref0 , and the motion vector MV2 of B2 Ref0 can be used as motion samples for deriving the affine motion parameters of the current block.
[0058] In addition to translational motion, the AVM also supports warp motion compensation. Two types of warp motion models are supported: a global warp model and a local warp model. The global warp model is associated with each reference frame, each of the four non-translational parameters has 12-bit accuracy, and the translational motion vector is coded with 15-bit accuracy. The coding block may choose to use it directly (a reference frame index is provided). The global warp model captures frame-level scaling and rotation. Thus, the global warp model mainly focuses on rigid body motion across the entire frame. A local warp model at the coding block level is also supported. In the local warp mode, also known as WARPED_CAUSAL, the warp parameters of the current block are derived by fitting a model to neighboring motion vectors using least squares.
[0059] A new warp motion mode is called WARP_EXTEND. In the WARP_EXTEND mode, the motion of adjacent blocks is smoothly extended into the current block and has the ability to improve warp parameters. This enables representing complex distorted motions while minimizing blocking artifacts and spreading them across multiple blocks. To achieve this, the WARP_EXTEND mode applied to the NEWMV block constructs a new warp model based on two constraints. The two constraints are that the per-pixel motion vector generated by the new warp model should be continuous with the per-pixel motion vector in the adjacent blocks, and that the pixel at the center of the current block should have a per-pixel motion vector that matches the motion vector signaled for the block as a whole. FIG. 7 shows the motion vectors in a block using the warp extension mode according to some embodiments. As shown in FIG. 7, for example, when an adjacent block 710 to the left of the current block 720 is distorted, a model that fits the motion vectors shown in FIG. 7 is used as the warp model.
[0060] The two constraints for constructing the new warp model mean specific equations that include the warp parameters of the adjacent and current blocks. By solving these equations, the warp model for the current block can be calculated. For example, if (A, …, F) represents the warp model of an adjacent (neighbor) and (A’, …, F’) represents the new warp model, the first constraint is that at each point along the common edge, it is as follows:
Equation
[0061] Note that the points along the edge have different values of y, but they all have the same value of x. This means that the coefficients of y must be the same on both sides (i.e., B’ = B and D’ = D). On the other hand, the x coefficient provides two equations regarding the other coefficients and is defined by the following equations (5)-(8): [Number]
[0062] Here, in equations (7)-(8), x is the horizontal position of the pixels in the vertical column and is thus, in effect, a constant.
[0063] The second constraint stipulates that the motion vector at the center of the block must be equal to that signaled using the NEWMV mechanism. This provides two more equations, resulting in a system of six equations in six variables with a unique solution. These equations can be efficiently solved both in software and hardware. The solution can be obtained using basic addition, subtraction, multiplication, and division by powers of two. Thus, this mode is significantly less complex than the least-squares based local warp mode.
[0064] Note that there may be multiple adjacent blocks that can be extended from. Therefore, a method for selecting which block to extend from is required. This problem is also encountered similarly in motion vector prediction. Specifically, there may be several possible motion vectors from nearby blocks, and one of them must be selected as the basis for NEWMV coding. The solution for this can be extended to handle the needs of WARP_EXTEND. This is done by tracking the source of each motion vector prediction. And WARP_EXTEND is enabled only when the selected motion vector prediction is taken from a directly adjacent block. Then, that block is used as a single "adjacent block" in the rest of the algorithm.
[0065] Note that Naver's warp model is sometimes very good as it is without requiring further changes. To code this case more inexpensively, WARP_EXTEND can be used for the NEARMV block. The Naver selection is the same as for NEWMV, except that it requires that in the selection in NEWMV, the Naver is distorted (not simply translated by translational motion). However, if this is true and WARP_EXTEND is selected, the Naver's warp model parameters are copied to the current block.
[0066] In some embodiments, a motion mode called WARP_DELTA can be used. In this mode, the block's warp model is coded as a delta from the predicted warp model, similar to how the motion vector is coded as a delta from the predicted motion vector. The prediction can be sourced from either the global motion model (if any) or an adjacent block.
[0067] Constraints can be applied to avoid multiple ways encoding the same predicted warp model. For example, when the mode is NEARMV or NEWMV, the same neighbor selection logic as described for WARP_EXTEND is used. If this results in a distorted adjacent block, the model of that adjacent block is used as the prediction (without applying the rest of the WARP_EXTEND logic). Otherwise, the global warp model is used as the basis. Other constraints may be applied. This example is not intended to limit the scope of the embodiments. Then, deltas for each of the non-translational parameters can be coded. Finally, the translational part of the model is adjusted so that the motion vector per pixel at the center of the block matches the overall motion vector of the block.
[0068] This tool (i.e., WARP_DELTA) involves explicitly coding deltas for each warp parameter, so it uses more bits for encoding than other warp modes. Therefore, WARP_DELTA can be disabled for blocks smaller than 16×16. However, the decode logic is extremely simple, so it can represent more complex motions that cannot be represented in other warp modes.
[0069] Merge using motion vector difference (MMVD) can be used in either skip mode or merge mode using a motion vector representation method. MMVD reuses merge candidates in VVC. A candidate is selected from among the merge candidates and can be further extended by the proposed motion vector representation method. MMVD provides a new motion vector representation using simplified signaling. The representation method includes a starting point, the magnitude of the motion, and the direction of the motion. The MMVD technique uses the merge candidate list in VVC. However, only candidates with the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of MMVD. Figure 8 shows the MMDV search process for the current frame using, for example, the two reference frames shown in Figure 5. The base candidate index defines the starting point of the motion vector representation method. The base candidate index indicates the best candidate among the candidates in the list according to Table 10 below.
Table 9
[0070] When the number of base candidates is equal to 1, the base candidate IDX is not signaled. The distance index represents the magnitude information of the motion. The distance index indicates a predetermined distance from the starting point information. The predetermined distance can be as shown in Table 11.
Table 10
[0071] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent four directions as shown in Table 12.
Table 11
[0072] The MMVD flag can be signaled immediately after sending the skip flag and the merge flag. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is in AFFINE mode. If the AFFINE flag is not equal to 1, the skip / merge index is parsed for the skip / merge mode in the reference software (e.g., VTM).
[0073] In the design of the local warp extension mode and the local warp delta mode (e.g., according to CWG-C050), for the warp delta mode, the base warp model is only allowed to be merged from one spatially adjacent block indicated by the MV prediction index (i.e., the MVP index). If this adjacent block does not use the local warp mode, the global warp model is used as the base mode for the warp delta mode. However, this is a second-best option because a better base warp model selection will result in a smaller warp delta that reduces the signaling cost and improves the quality of the coded block.
[0074] In some embodiments, the warp model of the current block can be derived using a regression method based on the positions and MVs of adjacent blocks. This warp model can be used as the basis for the warp delta mode. In some embodiments, when a warp model is used, the current block warp model is always used as the base model instead of the adjacent block warp model. When the warp delta mode is used, a flag can be signaled to indicate whether the current block warp model or the adjacent block warp model is used as the base model of the warp delta mode. When the warp delta model is used and the adjacent block indicated by the MVP index does not use local warp, the current block warp model can be used instead of the global warp model. If the current block does not have a warp model, a two-parameter translational warp model constructed using the MV of the current block can be used as the base model for the warp delta mode. In some embodiments, when using the warp model of the current block as the basis for the warp delta mode, the range of delta values can be different from the range when using the warp model of the adjacent block as the basis. For example, when using the warp model of the current block as the base warp model, the range of delta values is smaller than the range when using the warp model of the adjacent block as the base model.
[0075] In some embodiments, one or more warp models from other MVP adjacent blocks (including spatial neighbors and temporal neighbors) can be used as the basis for the warp delta mode.
[0076] In some embodiments, the warp model of adjacent blocks at fixed positions can be used as the basis. For example, the warp model of adjacent blocks from above, from the left, from the upper left corner, etc. can be used as the basis.
[0077] In some embodiments, multiple adjacent block warp models can be checked in a predetermined order by a pruning process. For example, the warp model of the adjacent block that is first in the scanning order using the local warp model can be used as the base model for the warp delta mode. As another example, when the warp delta mode is used and the adjacent block indicated by the MVP index does not use the local warp, another (predetermined) adjacent block warp model can be used (instead of the global warp model). In another example, when the MVP index indicates an upper adjacent block that does not use the local warp, the warp model of the left adjacent block can be used as the base. In another example, when the MVP index indicates a left adjacent block that does not use the local warp, the warp model of the upper adjacent block can be used as the base. In another example, when both the upper and left adjacent blocks do not use the local warp mode, the global warp model or the current block warp model can be used.
[0078] In some embodiments, additional flags / indexes can be signaled to indicate which warp model from which adjacent block is used. When there are multiple warp models from an adjacent block, the average (or median, or most frequent value) of all or selected warp models from the adjacent block can be used as the base for the warp delta mode. In one example, when there are multiple warp models from an adjacent block, the warp model parameters extrapolated from the warp model parameters used in the adjacent block can be used as the base for the warp delta mode.
[0079] In some embodiments, different base warp models may be used depending on whether the MV is coded using the NEAR mode or the NEW mode. For example, when the current MV is coded using NEARMV, the adjacent block as the SOTA is used as the base model. When the current MV is coded using the NEW mode, the projection model from the adjacent block indicated by the MVP may be used as the base model. The projection method may be the same as the method described with reference to the warp extension mode.
[0080] In some embodiments, the warp model of the already-coded block may be stored in a bank, a storage component, or the like. One of the bank (or stored) entries may be used as the basis for the warp delta mode. In one example, the bank warp model candidate may replace the adjacent block model using aspects of some embodiments. In some embodiments, the bank warp model may be indicated by the signaled index. In some embodiments, to reduce the memory bandwidth, the model accuracy stored in the bank can be made lower than that of the normal warp model, and when it is used as the base model, it is shifted back to the normal accuracy (bit depth). There may be multiple banks depending on the number of model parameters of the warp motion. For example, separate banks may be used to store two, four, and / or six warp motion model parameters. In some embodiments, there is a single bank for storing the warp motion model parameters, and this bank is updated each time the coding block is inter-coded. In some embodiments, the bank may be updated conditionally only according to the prediction mode of the current coding block. For example, the bank is updated only when the current block is coded by the warp motion mode.
[0081] In some embodiments, the warp delta mode can be used as an extended mode of the warp mode (hereinafter, the "normal warp mode") and the warp extended mode (i.e., instead of a separate mode parallel to the normal warp mode and the warp extended mode), and can be used as such. That is, when normal warp is used, a flag can be signaled to determine whether a warp delta is additionally signaled in addition to the current block warp model. When this warp delta flag indicates true, the warp delta is signaled. When this warp delta flag indicates false, the current block warp model is used. In some embodiments, when the warp extended mode is used, a flag can be signaled to determine whether a warp delta is additionally signaled in addition to the warp extended model. When this warp delta flag indicates true, the delta is signaled. When this warp delta flag indicates false, the warp extended model is used.
[0082] The range / accuracy / step size of the delta value for the parameters in the warp delta mode can be different when applying it to the normal warp mode and the warp extended mode.
[0083] FIG. 9 is a flowchart showing a method 910 for video coding executed by at least one process according to an embodiment.
[0084] In some implementations, one or more process blocks in FIG. 9 can be executed by computer 102. In some implementations, one or more process blocks in FIG. 9 may be executed by another device or group of devices separate from or included in computing environment 600.
[0085] As shown in FIG. 9, at operation 911, method 910 may include obtaining video data.
[0086] At operation 912, method 910 may include parsing the obtained video data into blocks.
[0087] In operation 913, method 910 can include generating a first warp model of a current block based on motion parameters of adjacent blocks, where the motion parameters of the adjacent blocks include at least block position information, motion vector information, and delta values.
[0088] In operation 914, method 910 may include selecting a warp model from among a first warp model of the current block and a second warp model associated with one of the adjacent blocks as a basis for a coding mode.
[0089] In operation 915, method 910 may include decoding video data in a coding mode based on the selected warp model.
[0090] FIG. 9 shows a block example of a method. However, in some implementations, the method may include additional blocks, fewer blocks, different blocks, or differently configured blocks compared to what is shown in FIG. 9. Additionally, or alternatively, two or more of the blocks of the method may be executed in parallel.
[0091] FIG. 10 is a block diagram 500 of internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment. It should be understood that FIG. 10 merely provides an illustration of one implementation and does not imply any limitation regarding environments in which different embodiments may be implemented. Numerous changes to the illustrated environment may be made based on design and implementation requirements.
[0092] Computer 102 (FIG. 1) and server computer 114 (FIG. 1) may each include a set 800A, B of respective internal components and a set 900A, B of external components. Each of the sets 800 of internal components includes one or more processors 820, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0093] Processor 820 may be implemented in hardware, software, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 820 includes one or more processors that can be programmed to execute functions. Bus 826 includes components that enable communication between internal components 800A, B.
[0094] One or more operating systems 828, software programs 108 (FIG. 1), and video encoding programs 116 (FIG. 1) on the server computer 114 (FIG. 1) can be stored in one or more of the respective computer-readable tangible storage devices 830 for execution by one or more of the respective processors 820 via one or more of the respective RAMs 822 (typically including cache memory). In the embodiment shown in FIG. 10, each of the computer-readable tangible storage devices 830 can be a magnetic disk storage device of an internal hard drive. In some embodiments, each of the computer-readable tangible storage devices 830 is a semiconductor storage device such as, for example, ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disk (CD), digital versatile disk (DVD), floppy disk (registered trademark), cartridge, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device capable of storing computer programs and digital information.
[0095] Each set 800A, B of internal components also includes an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936 such as, for example, CD-ROM, DVD, memory stick, magnetic tape, magnetic disk, optical disk, or semiconductor storage device. Software programs such as, for example, software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) are stored in one or more of the respective portable computer-readable tangible storage devices 936, read out via the respective R / W drives or interfaces 832, and can be loaded onto the respective hard drives such as, for example, storage device 830.
[0096] Each set 800A, B of internal components may also include a network adapter or interface 836, such as, for example, a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card, or other wired or wireless communication link. Software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) on server computer 114 (FIG. 1) may be downloaded from an external computer to computer 102 (FIG. 1) and server computer 114 via a network (e.g., the Internet, a local area network, or other wide area network) and respective network adapter or interface 836. From network adapter or interface 836, software program 108 and video encoding program 116 on server computer 114 may be loaded onto respective hard drives, such as, for example, storage device 830. The network may have copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0097] Each of the sets 900A, B of external components may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. External components 900A, B may also include a touch screen, virtual keyboard, touch pad, pointing device, and other human interface devices. Each of the sets 800A, B of internal components may also include a device driver 840 for interfacing with computer display monitor 920, keyboard 930, and computer mouse 934. Device driver 840, R / W drive or interface 832, and network adapter or interface 836 have hardware and software (stored in storage device 830 and / or ROM 824).
[0098] It should be understood in advance that this disclosure includes a detailed description of cloud computing, but the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, some embodiments can be implemented with any other type of computing environment, whether currently known or later developed.
[0099] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0100] The characteristics are as follows.
[0101] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed, without the need for human interaction with the service provider.
[0102] Broad network access: The capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0103] Resource pooling: The provider's computing resources are pooled to serve multiple users using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned according to requests. Users generally can specify their location at a higher level of abstraction (e.g., country, state, or data center), giving a sense of location independence in that they have no control or knowledge of the exact location of the resources provided.
[0104] Rapid adaptability: Functions can be made quickly and elastically available, in some cases automatically, to scale out rapidly, and can be released quickly to scale in. To the user, the capabilities available for provisioning often appear to be unlimited, and any amount can be purchased at any time.
[0105] Measured service: The cloud system automatically controls and optimizes resource utilization by using metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). It can monitor, control, and report resource usage, providing transparency to both the provider and the user of the services being utilized.
[0106] The service model is as follows.
[0107] Software as a Service (SaaS): The function provided to the user is to use the provider's application running on the cloud infrastructure. The application can be accessed from various client devices via a thin client interface such as a web browser (e.g., web-based email). The user generally does not manage or control the underlying cloud infrastructure, including the network, server, operating system, storage, or even individual application functions, except for limited user-specific application configuration settings.
[0108] Platform as a Service (PaaS): The function provided to the user is to deploy the application created or obtained by the user, which is created using the programming languages and tools supported by the provider, onto the cloud infrastructure. The user has control over the deployed application and optionally the application hosting environment settings, but does not manage or control the underlying cloud infrastructure, including the network, server, operating system, or storage.
[0109] Infrastructure as a Service (IaaS): The function provided to the user is to enable the user to deploy and run any software that may include an operating system and applications, using processing resources, storage resources, network resources, and other basic computing resources. The user has control over the operating system, storage, the deployed application, and limited control over optionally selecting network components (e.g., host firewall), but does not manage or control the underlying cloud infrastructure.
[0110] The deployment models are as follows.
[0111] Private cloud: The cloud infrastructure is operated only for a certain organization. This can be managed by that organization or a third party and may exist on-site or off-site.
[0112] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). This can be managed by those organizations or a third party and may exist on-site or off-site.
[0113] Public cloud: The cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services.
[0114] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public), and these two or more clouds remain distinct entities but are joined together by standardized or proprietary technologies that enable data and application portability (such as cloud bursting for load balancing between clouds).
[0115] The cloud computing environment is service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the center of cloud computing is an infrastructure with a network of interconnected nodes.
[0116] Referring to FIG. 11, an exemplary cloud computing environment 600 is shown that may be suitable for implementing a particular embodiment of the disclosed subject matter. As illustrated, cloud computing environment 600 has one or more cloud computing nodes 10 that may communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular telephone 54A, a desktop computer 54B, a laptop computer 54C, and / or an in-vehicle computer system 54N. The cloud computing nodes 10 may communicate with one another. They may be physically or virtually grouped (not shown) in one or more networks, such as, for example, private, community, public, or hybrid clouds as described hereinabove, or combinations thereof. This enables the cloud computing environment 600 to provide infrastructure, platform, and / or software as a service such that a cloud consumer need not maintain resources on a local computing device. It should be understood that the types of computing devices 54A - 54N shown in FIG. 11 are intended to be exemplary only, and that the cloud computing nodes 10 and the cloud computing environment 600 can communicate with any type of computerized device over any type of network and / or network addressable connection (e.g., using a web browser).
[0117] Referring to FIG. 12, a set of functional abstraction layers provided by cloud computing environment 600 (FIG. 11) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 12 are intended to be exemplary only and that embodiments are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0118] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0119] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual clients 75.
[0120] In one example, the management layer 80 can provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or invoicing for the consumption of these resources. In one example, these resources may have application software licenses. Security provides for the authentication of cloud users and tasks and for the protection of data and other resources. The user portal 83 provides access to the cloud computing environment for users and system administrators. Service level management 84 provides for the allocation and management of cloud computing resources such that the required service levels are met. Service level agreement (SLA) formulation and fulfillment 85 provides for the pre-provisioning and procurement of cloud computing resources for which future demands are predicted, in accordance with the SLA.
[0121] The workload layer 90 provides examples of functions for which a cloud computing environment can be utilized. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom lecture delivery 93, data analysis processing 94, transaction processing 95, and video encoding / decoding 96. Video encoding / decoding 96 can encode / decode video data using a delta angle derived from a nominal angle.
[0122] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of technical detail integration. The computer-readable media can include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to perform operations.
[0123] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks (registered trademark), mechanically encoded devices such as punch cards or raised structures in grooves that record instructions, and any suitable combination thereof. A computer-readable storage medium, as used herein, is not to be construed as being a transient signal per se, such as, for example, radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses passing through an optical fiber cable), or electrical signals transmitted through a wire.
[0124] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded via a network, such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network, from an external computer or an external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0125] The computer-readable program code / instructions for performing the operations can be in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in, for example, object-oriented programming languages such as Smalltalk, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing the aspect or operation.
[0126] These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which are executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions has a manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0127] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0128] It should be understood that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0129] These descriptions of various aspects and embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Even if combinations of features are recited in the claims and / or disclosed in the specification, those combinations are not intended to limit the possible implementations disclosed. In fact, many of these features may be combined in ways not specifically recited in the claims and / or not specifically disclosed in the specification. Each of the dependent claims listed below may, directly, depend on only one claim, but the possible implementations disclosed include each dependent claim in combination with all other claims in the claim set. Numerous changes and modifications will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or a technical improvement found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0130] FIG. 13 is a block diagram of an example of computer code 1300 for video coding according to an embodiment. In an embodiment, the computer code may be, for example, program code or computer program code. According to an embodiment of the present disclosure, a device / apparatus including at least one processor together with a memory storing the computer program code may be provided. The computer program code may be configured to execute any number of aspects of the present disclosure when executed by the at least one processor.
[0131] As shown in FIG. 13, the computer code 1300 includes an acquisition code 1310, an analysis code 1320, a generation code 1330, a selection code 1340, and a decoding code 1350.
[0132] The acquisition code 1310 is configured to cause the at least one processor to acquire video data.
[0133] The parsing code 1320 is configured to cause the at least one processor to parse the acquired video data into blocks.
[0134] The generation code 1330 is configured to cause the at least one processor to generate a first warp model of the current block based on the motion parameters of adjacent blocks, and the motion parameters of the adjacent blocks include at least block position information, motion vector information, and delta values.
[0135] The selection code 1340 is configured to cause the at least one processor to select a warp model from among the first warp model of the current block and a second warp model associated with one of the adjacent blocks as the basis of the coding mode.
[0136] The decoding code 1350 is configured to cause the at least one processor to decode the video data in the coding mode based on the selected warp model.
[0137] FIG. 13 shows an example block of code. However, in some implementations, the device / apparatus may include additional blocks, fewer blocks, different blocks, or blocks configured differently compared to those shown in FIG. 13. Additionally, or alternatively, two or more of the blocks of the device may be combined. In other words, although FIG. 13 shows separate blocks of code, these various code instructions need not be separate and may be combined.
[0138] The embodiments described herein may be used separately or combined in any order. Also, each of those methods (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium. For example, the term block may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU).
[0139] Although this disclosure describes several exemplary embodiments, there are changes, substitutions, and various equivalent alternatives that fall within the scope of the disclosure. Accordingly, it is understood that those skilled in the art can, although not explicitly illustrated or described herein, embody the principles of the disclosure and thus devise numerous systems and methods that are within its spirit and scope.
Claims
1. A method for video coding executed by at least one processor, comprising: obtaining a coded video bitstream; obtaining a plurality of blocks from the video bitstream, the plurality of blocks including a current block and one or more adjacent blocks; determining that a warp delta mode is used to predict the current block based on a syntax element in the video bitstream; determining positions and motion vectors of the one or more adjacent blocks; generating a first warp model of the current block based on the motion vectors of the one or more adjacent blocks; selecting a warp model from among the first warp model of the current block and a second warp model associated with one of the one or more adjacent blocks as a base model for a coding mode; decoding the plurality of blocks in the coding mode based on the base model; A method having the above steps.
2. The method further comprises: signaling a flag indicating whether to use the first warp model, the second warp model, or a third warp model as a base for the coding mode, wherein the third warp model is a global warp model; When an adjacent block indicated by an MVP index uses a local warp model, the flag indicates using the first warp model of the current block as the base for the coding mode. The method according to claim 1.
3. The method according to claim 1, further comprising: constructing a two-parameter translational warp model using motion vector information associated with the current block; and selecting the two-parameter translational warp model as a base for the coding mode when the first warp model is not generated.
4. The method according to claim 1, wherein a range of delta values included in motion parameters of the one or more adjacent blocks is smaller when using the first warp model as a base for the coding mode than when using the second warp model as a base for the coding mode.
5. The method according to claim 1, wherein a fourth warp model associated with adjacent blocks at a fixed position, based on the position information of the one or more adjacent blocks, is selected as the basis of the coding mode.
6. The method further comprises a step of selecting, as the basis of the coding mode, a set of warp models associated with a plurality of adjacent blocks, and an average of the set of warp models is used as the basis of the coding mode, according to claim 1.
7. The method further comprises a step of generating motion parameters of the current block based on the first warp model, the motion parameters of the current block are continuous with one or more of the motion parameters of the one or more adjacent blocks, and a pixel at the center of the current block coincides with the motion vector information of the motion vector of the current block signaled, or the motion parameters of the one or more adjacent blocks include at least block position information, motion vector information, and delta values, The method according to claim 1.
8. One or more processors; One or more memories storing a computer program; comprising the computer program causes the one or more processors to execute the method according to any one of claims 1 to 7. Apparatus.
9. A computer program causing a computer to execute the method according to any one of claims 1 to 7.