Improvement of Model Memory in Warp Expansion Mode and Warp Difference Mode
The local warp motion difference mode in video coding addresses inefficiencies in handling complex local motion by generating and applying warp models for improved decoding efficiency and compression performance.
Patent Information
- Application Number
- JP2024545953
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-03
- Filing Date
- 2022-11-08
- Publication Date
- 2025-07-25
AI Technical Summary
Existing video coding technologies, such as AV1 and VVC, face challenges in efficiently handling complex local motion variations in video frames, leading to suboptimal compression and increased computational complexity.
A video coding method that utilizes a local warp motion difference mode, where a set of parameters for a current block is generated and stored, allowing for the derivation of a warp model for other blocks, enhancing decoding efficiency by applying a warp model based on stored and derived parameters.
Improves decoding efficiency and compression performance by effectively modeling complex local motion variations, reducing computational complexity and enhancing video coding quality.
Smart Images

Figure 2025523735000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority based on and incorporates by reference in its entirety U.S. Provisional Patent Application No. 63 / 388,883, filed on July 13, 2022, and U.S. Patent Application No. 17 / 980,302, filed on November 3, 2022.
[0002] [Technical Field] The present disclosure generally relates to advanced image and video coding techniques, and more particularly, to model memory in warp extend mode and warp delta mode.
Background Art
[0003] AV1 (AOMedia Video 1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by AOMedia (Alliance for Open Media), a consortium established in 2015 that includes the semiconductor industry, video - on - demand providers, video content production companies, software development companies, and web browser vendors. Many of the components of the AV1 project were derived from previous research efforts by the consortium's members. Individual contributors had started experimental technical platforms several years earlier. Xiph / Mozilla's Daala had already made its code public in 2010, Google's experimental VP9 evolution project V10 was announced on September 12, 2014, and Cisco's Thor was made public on August 11, 2015. AV1, built on the VP9 codebase, incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was made public on April 7, 2016. The consortium released the AV1 bitstream specification on March 28, 2018, along with an encoder and decoder based on the reference software. On June 25, 2018, a verified version 1.0.0 of the specification was released. On January 8, 2019, a verified version 1.0.0 with errata sheet 1 was released. The AV1 bitstream specification includes the reference video codec.
[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3) and 2016 (version 4). They also investigated the potential need for standardization of future video coding technologies that may have significantly better performance than HEVC in terms of compression capabilities. In October 2017, a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP) was issued. By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for the 360° video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th JVET (Joint Video Exploration Team or Joint Video Expert Team) meeting. As a result of this meeting, JVET officially started the standardization of next-generation video coding beyond HEVC. The new standard was named VVC (Versatile Video Coding).
SUMMARY OF THE INVENTION
[0005] Embodiments of the present disclosure relate to a video coding method, device, and computer-readable medium for improving the local warp motion difference mode.
[0006] According to one aspect of one or more embodiments, a video coding method executed by at least one processor includes receiving a video bitstream including a plurality of blocks including at least a current block, extracting a syntax element from the video bitstream, the syntax element indicating that the current block should be predicted using a warp model, generating a set of parameters corresponding to the warp model for the current block, storing a first subset of the set of parameters corresponding to the warp model for the current block in a memory, at least one parameter from the set of parameters not being within the first subset, deriving at least one parameter corresponding to the warp model not within the first subset, generating a warp model for other blocks within the plurality of blocks based on the stored subset of parameters and the derived at least one parameter, and decoding the plurality of blocks based on the warp model.
[0007] According to another aspect of one or more embodiments, a device / apparatus and a non-transitory computer-readable medium that conform to the video coding method are also provided.
[0008] Further embodiments are described in the following description, some of which will be apparent from the description and / or may be realized by practicing the presented embodiments of the present disclosure.
Brief Description of the Drawings
[0009] To more clearly illustrate the technical solutions of the exemplary embodiments of the present disclosure, the following briefly introduces the accompanying drawings for explaining the exemplary embodiments. The accompanying drawings incorporated herein and constituting a part of this specification show embodiments that conform to the present disclosure and are used to explain the principles of the present disclosure together with this specification. Obviously, the accompanying drawings in the following description only show some embodiments. Furthermore, those skilled in the art will understand that the aspects of the exemplary embodiments may be combined together or implemented alone.
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Best Mode for Carrying Out the Invention
[0010] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0011] The above disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the implementation forms to the exact forms disclosed. Changes and modifications are possible in light of the above disclosure, or may be obtained from the implementation of the implementation. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with other embodiments (or one or more features of other embodiments). Furthermore, in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be (at least partially) executed simultaneously, and the order of one or more operations may be switched.
[0012] It is apparent that the systems and / or methods described herein may be implemented in different forms of hardware, software, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code. It is understood that the software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0013] Certain combinations of features are recited in the claims and / or disclosed in the specification, but these combinations are not intended to limit the disclosure of possible embodiments. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim combined with all other claims in the set of claims.
[0014] The proposed features described below may be used separately or combined in any order. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0015] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described. Also, as used in this specification, the singular is intended to include one or more items and may be used interchangeably with "one or more." When one item is intended, "one" or similar terms are used. Also, as used in this specification, terms such as "has," "have," "having," "include," "including," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Additionally, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0016] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It is understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0017] Referring now to FIG. 1, there is shown a functional block diagram of a network-connected computer environment for encoding and / or decoding video data according to an exemplary embodiment such as the embodiments described herein. A video coding system 100 (hereinafter referred to as "the system") is shown. It should be recognized that FIG. 1 is merely an illustration of one implementation and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0018] System 100 may include a computer 102 and a server computer 114. The computer 102 may communicate with the server computer 114 via a communication network 110 (hereinafter referred to as the "network"). The computer 102 may include a processor 104 and a software program 108 stored in a data storage device 106 and capable of interfacing with a user and communicating with the server computer 114. As will be described below with reference to FIG. 10, the computer 102 may include internal components 800A and external components 900A, respectively, and the server computer 114 may include internal components 800B and external components 900B, respectively. The computer 102 may be, for example, a mobile device, a telephone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of executing a program to access the network and access a database.
[0019] The server computer 114 may also operate in a cloud computing service model such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), as will be described below with reference to FIGS. 11 and 12. The server computer 114 may also be deployed in a cloud computing deployment model such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.
[0020] Server computer 114 that can be used to encode video data can execute a video encoding program 116 (hereinafter referred to as "program") that can interact with database 112. The video encoding program method will be described in more detail below with respect to FIG. 4. In one embodiment, computer 102 may operate as an input device including a user interface, while program 116 may operate mainly on server computer 114. In another embodiment, program 116 may operate mainly on one or more computers 102, while server computer 114 may be used for processing and storing data used by program 116. It should be noted that program 116 may be a stand-alone program or may be integrated into a larger video encoding program.
[0021] However, it should be noted that in some cases, the processing of program 116 may be shared in any ratio between computer 102 and server computer 114. In other embodiments, program 116 may operate on more than one computer, server computer, or some combination of computer and server computer, for example, multiple computers 102 that communicate with a single server computer 114 across network 110. In other embodiments, for example, program 116 may operate on multiple server computers 114 that communicate with multiple client computers across network 110. Alternatively, the program may operate on a network server that communicates with servers and multiple client computers across the network.
[0022] Network 110 may include a wired connection, a wireless connection, an optical fiber connection, or any combination thereof. Generally, Network 110 can be any combination of connections and protocols that support communication between Computer 102 and Server Computer 114. Network 110 can be, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as a Public Switched Telephone Network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, an optical fiber-based network, etc., and / or various types of networks such as combinations of the above or other types of networks.
[0023] The number and arrangement of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with different arrangements than those shown in FIG. 1. Further, two or more devices shown in FIG. 1 may be implemented within a single device, or a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Further alternatively or as an alternative, a set of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by other sets of devices of system 100.
[0024] As described above, AV1 is an open video coding format developed as a successor to VP9 and designed for video transmission over the Internet. As shown in FIG. 2A, VP9 uses a four-direction partition tree starting from the 64×64 level to the 4×4 level and has some further restrictions for blocks of 8×8 or less (as shown in the upper half of FIG. 2A). Note that the partition designated as R may be called recursive in that the same partition tree may be repeated at a lower scale until the partition reaches the lowest 4×4 level. As shown in FIG. 2B, AV1 not only extends the partition tree to a ten-direction structure but also increases the maximum size (referred to as a superblock in the terms of VP9 / AV1) to start from 128×128. Note that this may include 4:1 / 1:4 rectangular partitions that did not exist in VP9 (shown in FIG. 2A). None of the rectangular partitions can be further subdivided. Further, AV1 adds more flexibility to the use of partitions below the 8×8 level in the sense that 2×2 chroma inter prediction becomes possible in certain cases.
[0025] In HEVC, in order to adapt to various local characteristics, by using a quadtree (QT) structure called a coding tree, a coding tree unit (CTU) may be divided into coding units (CUs). The determination of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process may be applied, and the relevant information may be transmitted to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another QT structure such as the coding tree for the CU. One of the important features of the HEVC structure is having multiple partition concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square-shaped, but a PU can be square or rectangular for an inter-prediction block. One coding block may be further divided into four square sub-blocks, and transformation is performed on each sub-block (i.e., TU). Each TU may be further recursively divided into smaller TUs (using quadtree partitioning), which is called a residual quadtree (RQT). At the picture boundary, HEVC adopts implicit quadtree partitioning so that the block maintains quadtree partitioning until its size conforms to the picture boundary.
[0026] As shown in FIG. 3A, a coding tree unit (CTU) is first partitioned by a QT structure, and a QT leaf node is further partitioned by a binary tree (BT) structure to form a quad-tree plus binary tree (QTBT) block structure. The QTBT block structure removes the concept of multiple partition types. That is, the QTBT structure removes the separation of the concepts of CU, PU, and TU and supports more flexibility for the CU partition shape. In the QTBT block structure, the CU may have a square or rectangular shape. There are two types of BT partitions: symmetric horizontal partition and symmetric vertical partition. The BT leaf node is a CU, and its segmentation is used for prediction and transformation processing without further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. In the joint exploration model (JEM) used by JVET, the CU may be composed of coding blocks (CBs) of different color components. For example, in the case of a prediction (P) slice and a binary (B) slice in 4:2:0 chroma format, one CU may include one luma CB and two chroma CBs. The CU may also be composed of a single-component CB. For example, in the case of an I slice, one CU includes only one luma CB or only two chroma CBs.
[0027] In the QTBT partitioning method, the parameters include, but are not limited to, the CTU size (i.e., the size of the QT root node similar to the concept in HEVC), the minimum allowable QT leaf node size (i.e., MinQTSize), the maximum allowable BT root node size (i.e., MaxBTSize), the maximum allowable BT depth (i.e., MaxBTDepth), and the minimum allowable BT leaf node size (i.e., MinBTSize).
[0028] For example, the QTBT partitioning method may be as follows. The CTU size may be set to 128×128 luma samples having two corresponding 64×64 blocks of chroma samples, MinQTSize may be set to 16×16, MaxBTSize may be set to 64×64, MinBTSize (for both width and height) may be set to 4×4, and MaxBTDepth may be set to 4. The QT partition may be first applied to the CTU to generate QT leaf nodes. The QT leaf nodes may have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf QT node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), it is not further divided by BT. In some embodiments, the leaf QT node may be further partitioned by BT. Thus, the QT leaf node is also the root node of BT and has a BT depth of zero. When the BT depth reaches MaxBTDepth (i.e., 4), further division is not considered. When a BT node has a width equal to MinBTSize (i.e., 4), further horizontal division is not considered. Similarly, when a BT node has a height equal to MinBTSize, further vertical division is not considered. The leaf nodes of BT are further processed by prediction and transformation processing without further partitioning. In JEM, for example, the maximum CTU size is 256×256 luma samples.
[0029] FIG. 3A shows an example of block partitioning of the QTBT structure, and FIG. 3B shows the corresponding tree representation. The solid line indicates the QT partition, and the dotted line indicates the BT partition. In each split (i.e., non-leaf) node of BT, one flag may be signaled to indicate which split type (i.e., horizontal or vertical) is used. For example, as shown in FIG. 3B, 0 indicates a horizontal split, and 1 indicates a vertical split. In the case of the QT partition, since the QT partition always splits the block both horizontally and vertically to generate four sub-blocks having equal sizes, there is no need to indicate the split type.
[0030] The QTBT partitioning scheme supports flexibility for luma and chroma to have separate QTBT structures. Currently, for P slices and B slices, the luma CTB and chroma CTB within one CTU share the same QTBT structure. However, for I slices, the luma CTB is partitioned into CUs by the QTBT structure, and the chroma CTB is partitioned into chroma CUs by another QTBT structure. This means that the CUs within an I slice are composed of coding blocks of the luma component or coding blocks of two chroma components, while the CUs within a P slice or B slice are composed of coding blocks of all three color components.
[0031] In HEVC, in order to reduce memory access for motion compensation, the inter-prediction of small blocks is restricted, bi-prediction is not supported for 4×8 and 8×4 blocks, and inter-prediction is not supported for 4×4 blocks. In the QTBT implemented in JEM-7.0, these restrictions are removed.
[0032] Figure 4 shows the multi-type-tree (MTT) structure in VCC, which further adds (a) vertical center-bi tree partitioning and (b) horizontal center-bi tree partitioning in addition to QTBT. The advantages of the bi-tree partitioning shown in Figure 4 include complementing the quadtree and binary tree partitioning, that the quadtree and binary tree always divide along the block center while the bi-tree partitioning can capture the object located at the block center, and that the proposed width and height of the bi-tree partitions are always powers of 2 and thus no further transformation is required, but are not limited to these. The two-level tree design is mainly motivated by complexity reduction. Theoretically, the complexity of traversing the tree is T D where T represents the number of partition types and D represents the depth of the tree.
[0033] In merge mode, the implicitly derived motion information is directly used for generating the prediction samples of the current CU. Merge mode with motion vector differences (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the merge flag are sent to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, this may be further refined by the signaled MVD information. The MVD information may include a merge candidate flag, an index for specifying the magnitude of the motion, and an index for indicating the direction of the motion. In merge mode, one of the first two merge candidate flags in the merge list is selected to be used as the MV base. The merge candidate flag is signaled to specify which flag is used.
[0034] The distance index specifies the magnitude information of the motion and indicates a predetermined offset from the starting point. FIG. 5 shows the MMDV search points of two reference frames according to some embodiments. As shown in FIG. 5, the offset may be added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predetermined offset is specified in Table 1 below.
Table 1
[0035] The direction index represents the direction of the MVD with respect to the starting point. The direction index may represent one of four directions as shown in Table 2 below. Note that the meaning of the MVD code may vary according to the information of the starting MV. When the starting MV is a single-predicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., the POCs of both references are both greater than the POC of the current picture or the POCs of both references are both less than the POC of the current picture), the code in Table 2 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), and the difference in POC within list 0 (L0) is greater than the difference in POC within list 1 (L1), the code in Table 2 specifies the sign of the MV offset added to the L0 MV component of the starting MV, and the sign of the L1 MV has the opposite value. When the difference in POC within L1 is greater than that within L0, the code in Table 2 specifies the sign of the MV offset added to the L1 MV component of the starting MV, and the sign of the L0 MV has the opposite value.
[0036] The MVD is scaled according to the difference in POC in each direction. If the difference in POC within both lists is the same, scaling is not required. When the difference in POC within L0 is greater than the difference in L1, the MVD of L1 is scaled. When the POC difference in L1 is greater than that in L0, the MVD of L0 is scaled in the same way. When the starting MV is single-predicted, the MVD is added to the available MV.
Table 2
[0037] In VVC, in addition to the signaling of the MVD in the normal single-direction prediction and bi-direction prediction modes, a symmetric MVD mode for the signaling of the bi-direction MVD may be applied. In the symmetric MVD mode, motion information including both the reference picture indices of L0 and L1 and the MVD of L1 may be derived (not signaled).
[0038] The decoding process of the symmetric MVD mode is as follows. First, at the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived. For example, if mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. If the nearest reference picture in L0 and the nearest reference picture in L1 form a forward and backward pair of reference pictures, or a backward and forward pair of reference pictures, BiDirPredFlag is set to 1, and both the L0 reference picture and the L1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. Next, at the CU level, if the CU is bi-predicted coding and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true (e.g., equal to 1), only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of L0 and L1 are set equal to the pair of reference pictures, respectively. Finally, MVD1 is set equal to (-MVD0).
[0039] In AV1, for each coding block within an inter-frame, if the mode of the current block is an inter-coding mode rather than a skip mode, another flag is signaled to indicate whether a single-reference mode or a composite-reference mode is used for the current block. In the single-reference mode, the predicted block is generated by one motion vector. In the composite-reference mode, the predicted block is generated by the weighted average of two predicted blocks derived from two motion vectors. The modes that can be signaled for the single-reference case are detailed in Table 3 below.
Table 3
[0040] For the case of composite reference, the modes that can be signaled are detailed in Table 4 below. [Table 4]
[0041] AV1 enables 1 / 8 pixel motion vector accuracy (or precision). The syntax may be used to signal the motion vector difference in reference frame L0 or L1 as follows. The syntax mv_joint specifies which component of the motion vector difference is non-zero. A syntax mv_joint value of 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, a value of 1 indicates that there is a non-zero MVD along only the horizontal direction, a value of 2 indicates that there is a non-zero MVD along only the vertical direction, and a value of 3 indicates that there is a non-zero MVD along both the horizontal and vertical directions. The syntax mv_sign specifies whether the motion vector difference is positive or negative. The syntax mv_class specifies the class of the motion vector difference. As shown in Table 5 below, the higher the class, the greater the magnitude of the motion vector difference. The syntax Hhmv_bit specifies the integer part of the offset between the motion vector difference and the start magnitude of each MV class. The syntax mv_fr specifies the first two fractional bits of the motion vector difference. The syntax mv_hp specifies the third fractional bit of the motion vector difference. [Table 5]
[0042] For the NEW_NEARMV mode and the NEAR_NEWMV mode (shown in Table 4), the accuracy of the MVD depends on the relevant class and the size of the MVD. First, fractional MVDs are allowed only if the size of the MVD is 1 pixel or less. Next, when the value of the relevant MV class is MV_CLASS_1 or greater, only one MVD value is allowed, and the MVD values in each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). The allowed MVD values in each MV class are shown in Table 6.
Table 6
[0043] In some embodiments, when the current block is coded using the NEW_NEARMV or NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class. When the current block is not coded using the NEW_NEARMV or NEAR_NEWMV mode, another context is used to signal mv_joint or mv_class.
[0044] To indicate whether the MVDs for the two reference lists are signaled together, a new inter-coding mode (i.e., JOINT_NEWMV) may be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs for reference L0 and reference L1 are signaled together. Thus, only one MVD, named joint_mvd, may be signaled and transmitted to the decoder, and the differential MVs for reference L0 and reference L1 may be derived from joint_mvd. The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context may be added.
[0045] When the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for the reference L0 or reference L1 based on the POC distance. Specifically, the distance between the reference frame L0 and the current frame is shown as td0, and the distance between the reference frame L1 and the current frame is shown as td1. If td0 is greater than or equal to td1, the joint_mvd is directly used for reference L0, and the mvd for reference L1 is derived from the joint_mvd based on the following formula (1). derived_mvd = td1 / td0 * joint_mvd Formula (1)
[0046] If td1 is greater than or equal to td0, the joint_mvd is directly used for reference L1, and the mvd for reference L0 is derived from the joint_mvd based on the following formula (2). derived_mvd = td0 / td1 * joint_mvd Formula (2)
[0047] A new inter-coding mode (i.e., AMVDMV) may be added in the case of single reference. When the AMVDMV mode is selected, this indicates that AMVD is applied to the signal MVD. To indicate whether AMVD is applied to the joint MVD coding mode, a flag, for example, amvd_flag, may be added under the JOINT_NEWMV mode. When the adaptive MVD resolution is applied to the joint MVD coding mode named joint AMVD coding, the MVDs for the two reference frames are signaled together, and the accuracy of the MVD is implicitly determined by the magnitude of the MVD. The MVDs for two (or more) reference frames are signaled together, and the MVD coding is applied.
[0048] In the adaptive motion vector resolution (AMVR) first proposed in CWG-C012, a total of seven MV precisions (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, the AVM encoder explores all supported precision values and signals the best precision to the decoder. To reduce the encoder's execution time, two precision sets are supported. Each precision set contains four predetermined precisions. The precision sets are adaptively selected at the frame level based on the maximum precision value of the frame. Similar to AV1, the maximum precision is signaled in the frame header. Table 7 summarizes the precision values supported based on the frame-level maximum precision.
Table 7
[0049] In AVM software (similar to AV1), there is a frame-level flag to indicate whether the frame's MV includes sub-pel precision. AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the precision of a block is lower than the maximum precision, the motion model and interpolation filter are not signaled. When the precision of a block is lower than the maximum precision, the motion mode is assumed to be translational motion, and the interpolation filter is assumed to be the REGULAR interpolation filter. Similarly, when the precision of a block is either 4 pels or 8 pels, the inter-intra mode is not signaled and is assumed to be 0.
[0050] Motion compensation typically assumes a translational motion model between the reference block and the target block. However, warped motion utilizes an affine model. The affine motion model may be represented by Equation (3).
Equation
[0051] Here, [x, y] are the coordinates of the original pixel, and [x', y'] are the warped coordinates of the reference block. According to Equation (3), up to six parameters are required to specify the warped motion, where a3 and b3 specify the translational MV, a1 and b2 specify the scaling along the MV, and a2 and b1 specify the rotation.
[0052] In global warp motion compensation, global motion information including a global motion type and several motion parameters is signaled for each inter-reference frame. The global motion types and the number of associated parameters are listed in Table 9.
Table 8
[0053] After signaling the reference frame index, if a global motion is selected, the global motion type and parameters associated with a given reference frame for the current coding block are used.
[0054] In local warp motion compensation, local warp motion is allowed for an inter-coding block when the following conditions are met. First, the current block must use single-reference prediction. The width or height of the coding block must be 8 or more. Finally, at least one of the adjacent neighboring blocks must use the same reference frame as the current block.
[0055] When the local warp motion is used for the current block, the affine model parameters are estimated by minimizing the mean squared difference between the reference projection and the modeled projection based on the MVs of the current block and its adjacent neighboring blocks. To estimate the parameters of the local warp motion, when the adjacent blocks use the same reference frame as the current block, projection sample pairs of the central samples in the adjacent blocks and their corresponding samples in the reference frame are obtained. Subsequently, three additional samples are created by shifting the central position by a quarter sample in one or both dimensions. These additional samples may also be considered as projection sample pairs for ensuring the stability of the model parameter estimation process.
[0056] The MVs of the adjacent blocks used to derive the motion parameters are called motion samples. The motion samples are selected from adjacent blocks that use the same reference frame as the current block. Note that the warp motion prediction mode is only enabled for blocks that use a single reference frame.
[0057] Figure 6 shows exemplary motion samples used to derive the model parameters of a block using local warp motion prediction according to some embodiments. As shown in Figure 6, the MVs of adjacent blocks B0, B1, and B2 are referred to as MV0, MV1, and MV2, respectively. The current block is predicted using single prediction with reference frame Ref0. Adjacent block B0 is predicted using composite prediction with reference frames Ref0 and Ref1. Adjacent block B1 is predicted using single prediction with reference frame Ref0. Adjacent block B2 is predicted using composite prediction with reference frames Ref0 and Ref2. The motion vectors MV0Ref0 of B0, MV1Ref0 of B1, and MV2Ref0 of B2 may be used as motion samples for deriving the affine motion parameters of the current block.
[0058] In addition to translational motion, the AVM also supports warp motion compensation. Two types of warp motion models are supported, namely, the global warp model and the local warp model. The global warp model is associated with each reference frame, each of the four non-translational parameters has 12-bit accuracy, and the translational motion vector is coded with 15-bit accuracy. The coding block may choose to use it directly (a reference frame index is provided). The global warp model captures frame-level scaling and rotation. Thus, the global warp model mainly focuses on inflexible motion across the entire frame. The local warp model at the coding block level is also supported. In the local warp mode, also known as WARPED_CAUSAL, the warp parameters of the current block are derived by fitting the model to nearby motion vectors using least squares.
[0059] A new warp motion mode is called WARP_EXTEND. In the WARP_EXTEND mode, the motion of adjacent blocks is smoothly extended to the current block, but has some ability to change the warp parameters. This allows complex warping motions to be represented and spread across multiple blocks while minimizing blocking artifacts. To achieve this, the WARP_EXTEND mode applied to the NEWMV block constructs a new warp model based on two constraints. The per-pixel motion vector generated by the new warp model must be continuous with the per-pixel motion vector in the adjacent blocks, and the pixel at the center of the current block must have a per-pixel motion vector that matches the motion vector signaled for the block as a whole. FIG. 7 shows the motion vectors within a block using the warp extension mode according to some embodiments. For example, as shown in FIG. 7, when the adjacent block 710 to the left of the current block 720 is warped, a model that fits the motion vectors shown in FIG. 7 is used as the warp model.
[0060] The two constraints for constructing a new warp model mean specific equations that include the warp parameters of adjacent blocks and the current block. These equations may then be solved to calculate the warp model for the current block. For example, if (A,...,F) represents an adjacent warp model and (A',...,F') represents a new warp model, the first constraint is as follows for each point along the common edge:
Equation
[0061] Note that the points along the edge have different values of y, but they all have the same value of x. This means that the coefficients of y must be the same on both sides (i.e., B' = B and D' = D). On the other hand, the coefficients of x provide two equations regarding the other coefficients, defined by the following equations (5) - (8).
Equation
[0062] Here, in equations (7) - (8), x is the horizontal position of the vertical column of pixels and is thus, in effect, a constant.
[0063] The second constraint specifies that the motion vector at the center of the block must be equal to that signaled using the NEWMV mechanism. This provides two more equations, resulting in a system of six equations in six variables that has a unique solution. These equations can be efficiently solved both in software and in hardware. The solution can be obtained using basic addition, subtraction, multiplication, and division by powers of two. Thus, this mode does not become significantly more complex than the least - squares - based local warp mode.
[0064] Note that there may be multiple adjacent blocks that can be the source of expansion. Therefore, some method of selecting which block to expand from is required. This problem is similarly faced in motion vector prediction. In particular, there may be several possible motion vectors from nearby blocks, and one of them used must be selected as the basis for NEWMV coding. A solution for this may be extended to handle the need for WARP_EXTEND. This is done by tracking the source of each motion vector prediction. Then, WARP_EXTEND is enabled only if the selected motion vector prediction is taken from a directly adjacent block. That block is then used as a single "adjacent block" in the rest of the algorithm.
[0065] Note that in some cases, the adjacent warp model may be perfectly fine as it is without further modification. To code this case more inexpensively, WARP_EXTEND may be used for the NEARMV block. The selection of the adjacent is the same as in the case of NEWMV, except that the selection in NEWMV requires that the adjacent be warped (not just translated through translational motion). However, if this is true and WARP_EXTEND is selected, the adjacent warp model parameters are copied to the current block.
[0066] In some embodiments, a motion mode called WARP_DELTA may be used. In this mode, the warp model of a block is coded as the difference from a predicted warp model, similar to the way the motion vector is coded as the difference from a predicted motion vector. The prediction may be supplied from either a global motion model (if it exists) or an adjacent block.
[0067] To avoid having multiple ways to encode the same prediction warp model, restrictions may be applied. For example, when the mode is NEARMV or NEWMV, the same adjacent selection logic as described for WARP_EXTEND is used. If this results in a warped adjacent block, the model of that adjacent block (without applying the rest of the WARP_EXTEND logic) is used as the prediction. Otherwise, the global warp model is used as a basis. Other restrictions may be applied. This example is not intended to limit the scope of the embodiments. Then, the differences for each of the non-translational parameters may be coded. Finally, the translational part of the model is adjusted so that the motion vector for each pixel at the center of the block matches the overall motion vector of the block.
[0068] This tool (i.e., WARP_DELTA) involves explicitly coding the differences for each warp parameter and thus uses more bits to encode than other warp modes. Therefore, WARP_DELTA may be disabled for blocks smaller than 16×16. However, the decoding logic is extremely simple and thus may represent more complex motions that are not possible with other warp modes.
[0069] Merge with Motion Vector Difference (MMVD) may be used for either the skip mode or the merge mode using a motion vector representation method. MMVD re-uses merge candidates within VVC. A candidate is selected from among the merge candidates and may be further extended by the proposed motion vector representation method. MMVD provides a new motion vector representation using simple signaling. The representation method includes a starting point, a motion magnitude, and a motion direction. The MMVD technique uses the merge candidate list in VVC. However, only candidates with the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of MMVD. FIG. 8 shows, for example, the MMDV search process for the current frame using the two reference frames shown in FIG. 5. The base candidate index defines the starting point of the motion vector representation method. The base candidate index indicates the best candidate among the candidates in the list as shown in Table 10 below.
Table 9
[0070] When the number of base candidates is equal to 1, the base candidate IDX is not signaled. The distance index represents the motion magnitude information. The distance index indicates a predetermined distance from the starting point information. The predetermined distance may be as shown in Table 11 below.
Table 10
[0071] The direction index represents the direction of the MVD with respect to the starting point. The direction index may represent four directions as shown in Table 12.
Table 11
[0072] The MMVD flag may be signaled immediately after sending the skip and merge flags. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is in AFFINE mode. If the AFFINE flag is not equal to 1, the skip / merge index is parsed for the skip / merge mode in the reference software (e.g., VTM).
[0073] In the design of the warp extension mode and the warp differential mode (e.g., according to CWG-C050), both the warp extension mode and the warp differential mode need to store the warp model from adjacent blocks. However, the warp model may not need to be stored in the current AV1 or AOM design. The warp model can contain up to six parameters, and each parameter can have 16-bit precision, so the storage cost of the warp model of adjacent blocks can be very high. This cost is very expensive, for example, for line buffering.
[0074] In some embodiments, the warp model of the current block may be described using six parameters. That is, for example, the parameters [a, b, e; c, d, f]. To reduce memory and line buffering, a subset of the six parameters (e.g., less than six parameters) may be stored in memory. The memory may be, for example, the storage device 830 (as shown in FIG. 10). In one example, only four parameters [a, b; c, d] are stored, and the translation parameters e and f are derived using the coded MV of the current block. In other examples, only two parameters a and b are stored. Parameters c and d may be derived depending on the warp mode of the current block. Parameters c and d may be copied or projected, for example, when using the stored model in the warp expansion mode. As another example, parameters c and d from the model of the current block model may be used, and in the warp difference mode, the parameters may be quantized and signaled. Further, parameters e and f may be derived using any of the described methods.
[0075] In some embodiments, the stored warp model accuracy may be a lower value (or a reduced value from) the accuracy of the derived warp model accuracy for the current block, or the warp may be quantized. The reduced accuracy model method may reduce both normal model storage memory and expensive line buffering. In one example, instead of storing a 16-bit warp model, the warp model may be truncated to N bits (N < 16). When using this stored model, this is shifted up to 16 bits as the basis for the warp difference mode or the warp expansion mode. In one example, the stored bits may be determined by the type of adjacent block. For example, for spatially adjacent blocks, a relatively high (but less than 16) stored bit depth is used, and for temporally adjacent blocks, a relatively low stored bit depth is used.
[0076] In some embodiments, to reduce the line buffer of the warp model memory, when a block is located at the upper SB (CTU) boundary, the warp model from the upper adjacent block is prohibited from being used in both the warp difference model and the warp extension mode (or other potential modes that require the use of the spatially adjacent block warp model). In one example, when a block is located at the upper SB / CTU boundary, instead of using the warp model from the upper spatial adjacent block, regardless of where the MVP index points, the warp model from the left spatial adjacent block is used instead for the warp difference mode and the warp extension mode. In another example, when a block is located at the upper SB / CTU boundary, instead of using the warp model from the upper spatial adjacent block, the warp model of the MVP candidate having the lowest MVP index in the DRL list (which is from the left adjacent block) is used instead for the warp difference mode and the warp extension mode, regardless of where the MVP index points. In another example, when a block is located at the upper SB / CTU boundary, instead of the warp model from the upper spatial adjacent block, the warp model from the temporal adjacent block is used. In another example, when a block is located at the upper SB / CTU boundary, instead of using the warp model from the upper spatial adjacent, the global warp model is used. In another example, when a block is located at the upper SB / CTU boundary, instead of the warp model from the upper spatial adjacent block, the constructed translational warp model constructed using the MV of the adjacent block / current block is used.
[0077] FIG. 9 is a flowchart showing a method 910 for video coding executed by at least one process according to an embodiment.
[0078] In some implementations, one or more process blocks of FIG. 9 may be executed by computer 102. In some implementations, one or more process blocks of FIG. 9 may be executed by another device or group of devices included in a computing environment 600 separate from or included in computing environment 910.
[0079] As shown in FIG. 9, in operation 911, method 910 may include obtaining video data.
[0080] In operation 912, method 910 may include parsing the obtained video data into blocks.
[0081] In operation 913, method 910 may include generating a warp model for the current block based on a set of parameters of the current block, the set of parameters including at least block position information, motion vector information, and a difference value.
[0082] In operation 914, method 910 may include storing a subset of the parameters included in the set of parameters related to the current block.
[0083] In operation 915, method 910 may include selecting a first warp model for the current block based on the subset of parameters.
[0084] In operation 916, method 910 may include decoding the video data based on the warp model.
[0085] FIG. 9 shows exemplary blocks of a method, but in some implementations, the method may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in FIG. 9. Further or alternatively, two or more of the blocks of the method may be executed in parallel.
[0086] FIG. 10 is a block diagram 500 of the internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment. It should be recognized that FIG. 10 is only an illustration of one implementation and does not imply any limitation regarding the environments in which different embodiments may be implemented. Many modifications to the illustrated environment may be made based on design and implementation requirements.
[0087] The computer 102 (FIG. 1) and the server computer 114 (FIG. 1) may each include a respective set of internal components 800A, B and external components 900A, B. Each set of internal components 800 includes one or more processors 820, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.
[0088] The processor 820 may be implemented in hardware, software, or a combination of hardware and software. The processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor 820 includes one or more processors programmable to execute functions. The bus 826 includes components that enable communication between the internal components 800A, B.
[0089] One or more operating systems 828, software programs 108 (FIG. 1), and video encoding programs 116 (FIG. 1) on the server computer 114 (FIG. 1) may be stored in one or more of the respective computer-readable tangible storage devices 830 for execution by one or more of the respective processors 820 via one or more of the respective RAMs 822 (typically including cache memory). In the embodiment shown in FIG. 10, each of the computer-readable tangible storage devices 830 may be a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible storage devices 830 may be a semiconductor storage device such as a ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disc (CD), digital versatile disc (DVD), floppy disk, cartridge, magnetic tape, and / or other types of non-transitory computer-readable tangible storage devices capable of storing computer programs and digital information.
[0090] Each set of internal components 800A, B may also include an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible storage devices 936 such as a CD-ROM, DVD, memory stick, magnetic tape, magnetic disk, optical disk, or semiconductor storage device. Software programs such as software programs 108 (FIG. 1) and video encoding programs 116 (FIG. 1) may be stored in one or more of the respective portable computer-readable tangible storage devices 936, read via the respective R / W drives or interfaces 832, and loaded into the respective hard drives, e.g., storage devices 830.
[0091] Each set of internal components 800A, B may also include a network adapter or interface 836 such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card or other wired or wireless communication link. Software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) on server computer 114 (FIG. 1) can be downloaded from an external computer to computer 102 (FIG. 1) and server computer 114 via a network (e.g., the Internet, a local area network, or other wide area network) and respective network adapter or interface 836. From network adapter or interface 836, software program 108 and video encoding program 116 on server computer 114 may be loaded onto respective hard drives, e.g., storage device 830. The network may include copper wire, fiber optic, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0092] Each set of external components 900A, B may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. External components 900A, B may also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each set of internal components 800A, B may also include a device driver 840 for interfacing with computer display monitor 920, keyboard 930, and computer mouse 934. Device driver 840, R / W drive or interface 832, and network adapter or interface 836 include hardware and software (stored in storage device 830 and / or ROM 824).
[0093] This disclosure includes detailed descriptions regarding cloud computing, but it is to be understood in advance that the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, some embodiments are implementable with any other type of computing environment, whether currently known or later developed.
[0094] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0095] The characteristics are as follows.
[0096] On-demand self-service: Cloud consumers can unilaterally provision computing capabilities, such as server time and network storage, as needed, automatically without the need for human interaction with the service provider.
[0097] Broad network access: The capabilities are available over a network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0098] Resource pooling: The computing resources of a provider are pooled to serve multiple users using a multi-tenant model, and different physical and virtual resources are dynamically assigned and re-assigned on demand. Generally, users do not have control or awareness of the exact location of the resources provided to them, but there is a concept of location independence in that they can specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0099] Rapid elasticity: In some cases, functions can be provisioned quickly and elastically so that they are automatically released to scale out rapidly and scale in rapidly on demand. To the user, the functions available for provisioning often appear to be unlimited, and any amount can be purchased at any time.
[0100] Measured service: A cloud system automatically controls and optimizes resource usage by leveraging a metering function at some level of abstraction (e.g., storage, processing, bandwidth, and active user accounts) appropriate to the type of service. Resource usage is monitored, controlled, and reported, and can provide transparency to both the provider and the user of the service being utilized.
[0101] The service model is as follows.
[0102] Software as a Service (SaaS): The functions provided to the user are those that use the provider's applications running on the cloud infrastructure. The applications are accessible from various client devices through a client interface such as a web browser (e.g., web-based email). The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even the individual application functions, except in some cases where user-specific application configuration settings are an exception.
[0103] Platform as a Service (PaaS): The functions provided to users are to deploy applications created or obtained by users, which are created using programming languages and tools supported by the provider on the cloud infrastructure. Users do not manage or control the underlying cloud infrastructure including the network, servers, operating systems, or storage, but control the deployed applications and, in some cases, the configuration of the application hosting environment.
[0104] Infrastructure as a Service (IaaS): The functions provided to users are to provision processing, storage, networks, and other basic computing resources, and users can deploy and run any software that may include operating systems and applications. Users do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage, deployed applications, and, in some cases, limited control over selected network components (e.g., host firewalls).
[0105] The deployment models are as follows.
[0106] Private cloud: The cloud infrastructure is dedicatedly operated for an organization. This may be managed by the organization or a third party and may exist on-premises or off-premises.
[0107] Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by an organization or a third party and may exist on - premise or off - premise.
[0108] Public Cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.
[0109] Hybrid Cloud: The cloud infrastructure remains a distinct entity but is composed of two or more clouds (private, community, or public) joined together by standardized or proprietary technologies (e.g., cloud bursting for load - balancing between clouds) that enable data and application portability.
[0110] Cloud computing environments are service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure that includes a network of interconnected nodes.
[0111] Referring to FIG. 11, an exemplary cloud computing environment 600 that may be suitable for implementing certain embodiments of the subject matter of the present disclosure is shown. As illustrated, cloud computing environment 600 includes one or more cloud computing nodes 10, and local computing devices used by cloud consumers, such as personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, and / or automotive computer system 54N, may communicate with one or more cloud computing nodes 10. The cloud computing nodes 10 may communicate with each other. These may be physically or virtually grouped in one or more networks such as the private, community, public, or hybrid clouds as described above, or combinations thereof (not shown). This enables cloud computing environment 600 to provide infrastructure, platform, and / or software as a service where cloud consumers do not need to maintain resources on local computing devices. The types of computing devices 54A - 54N shown in FIG. 11 are only intended to be exemplary, and it is understood that cloud computing nodes 10 and cloud computing environment 600 can communicate with any type of computer device over any type of network and / or network addressable connection (e.g., using a web browser).
[0112] Referring to FIG. 12, a set of functional abstraction layers 700 provided by cloud computing environment 600 (FIG. 11) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 12 are only intended to be exemplary and embodiments are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0113] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0114] The virtualization layer 70 provides an abstraction layer from which examples of virtual entities such as virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75 may be provided.
[0115] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Metering and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and for charging or billing for the use of these resources. In one example, these resources may include application software licenses. Security provides for authentication for cloud users and tasks and for the protection of data and other resources. User portal 83 provides access to the cloud computing environment for users and system administrators. Service level management 84 provides for the allocation and management of cloud computing resources such that the required service level is met. Planning and enforcement of service level agreements (SLAs) 85 provides for the pre-placement and procurement of cloud computing resources for which future requirements are predicted in accordance with the SLA.
[0116] The workload layer 90 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and video encoding / decoding 96. The video encoding / decoding 96 may encode / decode video data using a differential angle derived from a nominal angle.
[0117] Some embodiments may relate to a system, method, and / or computer-readable medium in the integration of any possible level of technical detail. The computer-readable medium may include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions for causing a processor to execute operations.
[0118] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards, or raised structures in grooves having instructions recorded thereon, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a signal per se that is transient, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0119] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.
[0120] The computer-readable program code / instructions for performing the operations may be in any combination of source code or object code written in any one or more programming languages, including assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit to perform the aspects or operations.
[0121] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture including instructions which implement the manner of functioning specified in the flowchart and / or block diagram blocks.
[0122] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram blocks.
[0123] It is apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. It is understood that the actual special control hardware or software code used to implement these systems and / or methods is not a limitation of the implementation. Accordingly, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware may be designed based on the description herein to implement the systems and / or methods.
[0124] The descriptions of the various aspects and embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Combinations of features are described in the claims and / or disclosed herein, but these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim combined with all other claims in the claim set. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application, or a technical improvement over technologies found in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0125] FIG. 13 is a block diagram of an example of computer code 1300 for video coding according to an embodiment. In an embodiment, the computer code may be, for example, program code or computer program code. According to an embodiment of the present disclosure, there may be provided an apparatus / device including at least one processor having a memory storing computer program code. The computer program code may be configured to execute any number of aspects of the present disclosure when executed by the at least one processor.
[0126] As shown in FIG. 13, computer code 1300 includes received code 1310, extraction code 1320, first generation code 1330, storage code 1340, derivation code 1350, second generation code 1360, and decoding code 1370.
[0127] Received code 1310 is configured to cause at least one processor to receive a video bitstream including a plurality of blocks including at least a current block.
[0128] The extraction code 1320 is configured to cause at least one processor to extract a syntax element from a video bitstream, the syntax element indicating that the current block is to be predicted using a warp model.
[0129] The first generation code 1330 is configured to cause at least one processor to generate a set of parameters corresponding to a warp model for the current block.
[0130] The storage code 1340 is configured to cause at least one processor to store in memory a first subset of a set of parameters corresponding to a warp model for the current block, with at least one parameter from the set of parameters not being within the first subset.
[0131] The derivation code 1350 is configured to cause at least one processor to derive at least one parameter corresponding to a warp model that is not within the first subset.
[0132] The second generation code 1360 is configured to cause at least one processor to generate a warp model for other blocks within a plurality of blocks based on the stored subset of parameters and the at least one derived parameter.
[0133] The decoding code 1370 is configured to cause at least one processor to decode a plurality of blocks based on the warp model.
[0134] FIG. 13 shows exemplary code blocks, but in some implementations, the apparatus / device may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those shown in FIG. 13. Further or alternatively, two or more of the blocks of the apparatus may be combined. In other words, although FIG. 13 shows separate blocks of code, various code instructions need not be separate and may be intermixed.
[0135] The embodiments described herein may be used separately or may be combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium. For example, the term block may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU).
[0136] Although the present disclosure has described some exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized by those skilled in the art that many systems and methods, although not explicitly illustrated or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure, can be devised.
Claims
1. A method for video coding executed by at least one processor, comprising: receiving a video bitstream including a plurality of blocks including at least a current block; extracting a syntax element from the video bitstream, the syntax element indicating that the current block should be predicted using a warp model; generating a set of parameters corresponding to the warp model for the current block; storing in a memory a first subset of the set of parameters corresponding to the warp model for the current block, wherein at least one parameter from the set of parameters is not within the first subset; deriving the at least one parameter corresponding to the warp model that is not within the first subset; generating a warp model for other blocks within the plurality of blocks based on the stored first subset and the derived at least one parameter; decoding the plurality of blocks based on the warp model and a method including the above.
2. The method according to claim 1, wherein one or more parameters included in the set of parameters and not included in the first subset are determined based on the coded motion vector of the current block.
3. The method according to claim 1, wherein one or more parameters included in the set of parameters include at least one of block position information, motion vector information, or difference values.
4. One or more parameters included in the set of parameters and not included in the first subset are determined based on the warp mode of the current block, and projected using the stored first subset when the warp mode is a warp extension mode. The method according to claim 1.
5. The method according to claim 4, further comprising quantizing and signaling the one or more parameters included in the set of parameters and not included in the first subset when the warp mode is a warp difference mode.
6. storing the model accuracy of the warp model; A step of determining a bit size of the warp model based on a block type of an adjacent block, wherein the block type is a spatial block or a temporal block The method according to claim 1, further comprising. **Claim 7** A step of determining a position of the current block; A step of selecting an adjacent block warp model from adjacent blocks for decoding of the current block based on the position of the current block The method according to claim 1, further comprising. **Claim 8** A device for video coding, comprising: At least one memory configured to store computer program code; At least one processor configured to read the computer program code and operate as instructed by the computer program code The computer program code causes the at least one processor to execute the method according to any one of claims 1 to 7. **Claim 9** A computer program for causing at least one processor to execute the method according to any one of claims 1 to 7. **Claim 10** A method for video coding executed by at least one processor, comprising: Generating a set of parameters corresponding to a warp model for a current block; Storing in a memory a first subset of the set of parameters corresponding to the warp model for the current block, wherein at least one parameter from the set of parameters is not within the first subset; Deriving the at least one parameter corresponding to the warp model not within the first subset; Generating a warp model for other blocks within a plurality of blocks including the current block based on the stored first subset and the derived at least one parameter; Encoding the plurality of blocks based on the warp model; Transmitting a video bitstream including a syntax element indicating that the current block is to be predicted using a warp model and the plurality of blocks A method including.
Citation Information
Cited By
Method, computing system, device, and computer program for decoding video
JP2025536953A