Improvement of Local Distortion Motion Prediction Mode

JP2025522662A5Pending Publication Date: 2025-11-21TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024522542
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2022-11-14
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing video coding technologies, such as AV1 and VVC, face limitations in local distortion motion prediction modes, particularly in handling single and composite reference pictures, with inefficiencies in deriving motion parameters and flexibility in block splitting.

Method used

A method for video encoding that involves generating a distortion model based on motion vectors of current and adjacent blocks, applying this model to reference picture lists, and decoding frames to improve local distortion motion prediction.

Benefits of technology

Enhances the accuracy and flexibility of local distortion motion prediction, allowing for more efficient encoding and decoding of video data, particularly in handling multiple reference frames and block splitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An approach for encoding / decoding video data, which is executed by at least one processor and includes: a step of acquiring video data; a step of parsing the video data into blocks, where the blocks are related to a reference picture list; a step of generating a distortion model used for a first reference picture list and a second reference picture list of a current block included in the reference picture list based on a motion vector of the current block and motion vectors of blocks adjacent to the current block; and a step of decoding frames in the first reference picture list and the second reference picture list by applying the distortion model to the frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 358,735, filed on July 6, 2022, and U.S. Patent Application No. 17 / 983,975, filed on November 9, 2022. The entire disclosures of these are incorporated by reference.

[0002] This disclosure generally relates to advanced image and video coding techniques, and more particularly, to coding and / or decoding related to local warp motion modes.

Background Art

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. AV1 was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium established in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser suppliers. Many of the components of the AV1 project are sourced from the previous research efforts of AOMedia members. Individual contributors initiated a pilot technology platform several years ago. In 2010, Xiph / Mozilla's Daala had already released its code, on September 12, 2014, Google's experimental VP9 development project VP10 was announced, and on August 11, 2015, Cisco's Thor was announced. Based on the VP9 codebase, additional technologies were incorporated into AV1, some of which were developed in the context of that experimental format. On April 7, 2016, the first version 0.1.0 of the AV1 reference codec was announced. In addition to the software-based reference encoder and decoder, on March 28, 2018, AOMedia announced the release of the specification for the AV1 bitstream. On June 25, 2018, the verified version 1.0.0 of the specification was released. On January 8, 2019, the verified version 1.0.0 with errata 1 of the specification was released. The specification for the AV1 bitstream includes a reference video codec.

[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). The group has also investigated the possibility that future video coding technologies (technologies that can significantly exceed the compression capabilities of HEVC) may require standardization. In October 2017, the group issued a "Joint Call for Proposals on Video Compression with Capability beyond HEVC" (CfP). By February 15, 2018, submissions to the CfP for standard dynamic range (SDR) (a total of 22 submissions), submissions to the CfP for high dynamic range (HDR) (12 submissions), and submissions to the CfP for the 360-degree video category (12 submissions) were respectively submitted. In April 2018, all submissions to the CfP that were accepted were discussed at the 122nd MPEG / 10th JVET (Joint Video Exploration Team or Joint Video Expert Team) meeting. As a result of this meeting, JVET officially began standardizing the next generation of video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC).

Summary of the Invention

Problems to be Solved by the Invention

[0005] Embodiments of the present disclosure relate to a video coding method, device, and computer-readable medium for improving the local distortion motion Δ mode.

Means for Solving the Problems

[0006] According to aspects of one or more embodiments, a video encoding method is executed by at least one processor, and the method includes: obtaining video data including a plurality of blocks, where each block of the plurality of blocks is related to a first reference picture list and a second reference picture list; generating a distortion model used for the first reference picture list and the second reference picture list of a current block based on a motion vector of the current block and motion vectors of neighboring blocks adjacent to the current block; and decoding frames in the first reference picture list and the second reference picture list by applying the distortion model to the frames.

[0007] According to other aspects of one or more embodiments, a device / apparatus and a non-transitory computer-readable medium that conform to the video encoding method are also provided.

[0008] Further embodiments are shown in the following description, some of which will be apparent from the description and / or can be realized by implementing the disclosed embodiments.

[0009] To more clearly explain the technical solutions of the embodiments of the present disclosure, the accompanying drawings for explaining the embodiments are briefly introduced below. The accompanying drawings introduced here are incorporated into this specification and form a part of this specification, showing embodiments that conform to the present disclosure, and are used to explain the principles of the present disclosure together with this specification. It is obvious that the accompanying drawings in the following description only show some embodiments. Those skilled in the art will understand that the aspects of the embodiments may be combined or implemented alone.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following detailed description of the embodiments, reference is made to the accompanying drawings. The same or similar elements may be identified by the same numbers in different drawings.

[0012] Although the above disclosure provides examples and explanations, the above disclosure is not intended to be a restrictive enumeration or to limit implementation to exactly match the disclosed forms. In light of the above disclosure, modifications and variations are possible and can be obtained from the implementation of the examples. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). In addition to the above, it can be seen that in the flowcharts and descriptions of operations described below, one or more operations may be omitted, one or more operations may be added, one or more operations may be (at least partially) executed simultaneously, and the order of one or more operations may be swapped.

[0013] It is obvious that the systems and / or methods described in this disclosure may be implemented in individual forms of hardware, software, or a combination of hardware and software. The actual dedicated control hardware code or software code used to implement such systems and / or methods does not limit the implementation. Therefore, in this disclosure, the operations and functions of the systems and / or methods are described without referring to specific software codes. It can be seen that software and hardware can be designed to implement the systems and / or methods based on the description of this disclosure.

[0014] Even if a specific combination of features is recited in the claims and / or disclosed in the specification, such a combination is not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways that are not explicitly recited in the claims and / or not explicitly disclosed in the specification. Each of the dependent claims listed below can be directly dependent on only one claim, but combinations of each dependent claim with any other claim in the set of claims are included in the disclosure of possible implementations.

[0015] The proposed features described below may be used individually or in any combination in any order. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0016] Elements, operations, or instructions used in this application are not to be construed as being core or essential, unless so specified. Also, the articles “a” and “an” used in this application are intended to include one or more things and can also be used in the sense of “one or more”. Where only one thing is intended, the term “one” or a similar expression is used. Also, the terms “has”, “have”, “having”, “include”, “including”, etc. used in this disclosure are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “at least partially based on” unless otherwise specified. Additionally, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” are understood to include only A, only B, or both A and B.

[0017] In this disclosure, aspects are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It can be seen that each block of the flowchart and / or block diagram and combinations of blocks in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0018] Referring to FIG. 1, which is a functional block diagram of a networked computer environment showing a video encoding system 100 (hereinafter referred to as "the system") for encoding and / or decoding video data according to typical embodiments such as those described in the present disclosure. Note that FIG. 1 merely shows an example of one implementation and does not imply any limitation to the environment in which different embodiments can be implemented. Based on design requirements and implementation requirements, numerous modifications to the illustrated environment may be made.

[0019] System 100 may include a computer 102 and a server computer 114. Computer 102 may communicate with server computer 114 via a communication network 110 (hereinafter referred to as "the network"). Computer 102 may include a processor 104 and a software program 108 stored in a data storage device 106 and configured to interact with a user via an interface and communicate with server computer 114. As will be described later with reference to FIG. 10, computer 102 may include internal components 800A and external components 900A, respectively, and server computer 114 may include internal components 800B and external components 900B, respectively. Computer 102 may be, for example, a mobile device, a telephone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of executing programs, accessing a network, and accessing a database.

[0020] As will be described later with reference to FIGS. 11 and 12, the server computer 114 may also operate in cloud computing service models such as Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). The server computer 114 may also be deployed in cloud computing deployment models such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.

[0021] The server computer 114 may be used to encode video data and is provided with a function of executing a video encoding program 116 (hereinafter referred to as "program") that can communicate with the database 112. The video encoding program method will be described in more detail later with reference to FIG. 4. In one embodiment, the computer 102 may operate as an input device including a user interface in a state where the program 116 can be mainly executed on the server computer 114. In another embodiment, the program 116 may be mainly executed on one or more computers 102 in a state where the server computer 114 can be used for processing and storing data used by the program 116. It should be noted that the program 116 may be a stand-alone program or may be incorporated into a large-scale video encoding program.

[0022] However, it should be noted that the processing of program 116 may be shared among several instances in any ratio between computer 102 and server computer 114. In another embodiment, program 116 may operate on two or more computers, server computers, or some combinations of computers and server computers. For example, it may operate on a plurality of computers 102 communicating across a network 110 with one server computer 114. In another embodiment, for example, program 116 may operate on a plurality of server computers 114 communicating across network 110 with a plurality of client computers. Alternatively, the program may operate on a network server that communicates across a network with a server and a plurality of client computers.

[0023] Network 110 may include a wired connection, a wireless connection, an optical fiber connection, or some combination thereof. Generally speaking, network 110 can be any combination of connections and protocols that support communication between computer 102 and server computer 114. Network 110 may include various types of networks, such as local area networks (LANs), wide area networks (WANs) such as the Internet, telecommunication networks such as the public switched telephone network (PSTN), wireless networks, public switched networks, satellite networks, cellular networks (e.g., fifth generation (5G) networks, long term evolution (LTE) networks, third generation (3G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMNs), metropolitan area networks (MANs), private networks, ad hoc networks, intranets, optical fiber-based networks, etc., and / or combinations of the above networks or other networks.

[0024] The number and arrangement of the devices and networks shown in FIG. 1 are shown as an example. In reality, compared with what is shown in FIG. 1, there may be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with different arrangements. Further, two or more devices shown in FIG. 1 may be implemented in one device, or one device shown in FIG. 1 may be implemented as a plurality of distributed devices. In addition to or instead of the above, a set of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another set of devices of system 100.

[0025] As described above, AV1 is an open video coding format designed for video transmission that uses the Internet and was developed as a successor to VP9. As shown in FIG. 2A, VP9 uses four splitting trees from the 64×64 level to the 4×4 level starting from the 64×64 level, but there are some additional restrictions for blocks of 8×8 or less (shown in the upper half of FIG. 2A). Note that the split shown as R may be referred to as recursive in that the same split tree can be repeated on a smaller scale until the smallest 4×4 level is reached. As shown in FIG. 2B, AV1 not only expands the split tree to ten structures, but also expands the maximum size (referred to as a superblock in the VP9 / AV1 naming convention) to start from 128×128. Note that this can include 4:1 / 1:4 rectangular splits that did not exist in VP9 (shown in FIG. 2A). None of the rectangular splits are further subdivided. In addition to the above, in this case, higher flexibility is added by AV1 up to the use of splits less than the 8×8 level in the sense that 2×2 chroma intra prediction is possible in some examples.

[0026] In HEVC, a coding tree unit (CTU) can be divided into coding units (CUs) using a quadtree (QT) structure represented as a coding tree for adapting to various local features. A decision on whether to encode the picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CU level. Each CU can be further divided into one prediction unit (PU) according to the PU division type, or further divided into two PUs, or further divided into four PUs. The same prediction process may be applied within one PU, or relevant information may be sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU division type, the CU can be divided into a transform unit (TU) according to another QT structure such as the coding tree of the CU. One of the key features of the HEVC structure is the existence of multiple division concepts including CUs, PUs, and TUs. In HEVC, while the PU for an inter-predicted block may be square or rectangular in shape, only the CU or TU can be square in shape. One coding block may be further divided into four square sub-blocks, and a transform may be performed on each sub-block (i.e., TU). Each TU may be recursively further divided into smaller TUs (using quadtree division), which is called a Residual Quad-Tree (RQT). In HEVC, by using implicit quadtree division at the picture boundary, the block will continue to be divided into quadtree until it fits the picture boundary size.

[0027] As shown in FIG. 3A, first, a coding tree unit (CTU) is divided by a QT structure, and a QT leaf node is further divided by a binary tree (BT) structure to form a quad-tree plus binary tree (QTBT) block structure. Multiple split-type concepts are removed by the QTBT block structure. That is, the separation of the CU, PU, and TU concepts is removed by the QTBT structure, and higher flexibility of the CU split shape is supported. In the QTBT block structure, the shape of the CU may be square or rectangular. There are two types of splits for BT splitting: horizontal symmetric splitting and vertical symmetric splitting. The BT leaf node is a CU, and the split is used for prediction processing and conversion processing without any further splitting. This means that in the QTBT coding block structure, the block sizes of the CU, PU, and TU are the same. In the joint exploration model (JEM) used by JVET, a CU may consist of coding blocks (CBs) of different color components. For example, in the case of a prediction (P) slice and a binary (B) slice in a 4:2:0 chroma format, one CU may include one luma CB and two chroma CBs. Also, a CU may consist of a one-component CB. For example, in the case of an I slice, one CU may include only one luma CB or only two chroma CBs.

[0028] In the QTBT splitting method, the parameters include, but are not limited to, the CTU size (i.e., the size of the root node of the QT, similar to the concept of HEVC), the minimum allowable QT leaf node size (i.e., MinQTSize), the maximum allowable BT root node size (i.e., MaxBTSize), the maximum allowable BT depth (i.e., MaxBTDepth), and the minimum allowable BT leaf node size (i.e., MinBTSize).

[0029] For example, the QTBT splitting method may be as follows. The CTU size may be set to 128×128 luma samples, there are two corresponding 64×64 blocks of chroma samples, MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize is set to 4×4 (set for both width and height), and MaxBTDepth is set to 4. First, QT splitting may be applied to the CTU to generate QT leaf nodes. The size of the QT leaf node may be from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the leaf QT node is 128×128, since its size exceeds MaxBTSize (i.e., 64×64), the leaf QT node will not be further split by the BT. In some embodiments, the leaf QT node may be further split by the BT. Thus, the QT leaf node is also the root node of the BT, and its BT depth is zero. When the BT depth reaches MaxBTDepth (i.e., 4), no further splitting is considered. When the width of the BT node is equal to MinBTSize (i.e., 4), no further horizontal splitting is considered. Similarly, when the height of the BT node is equal to MinBTSize, no further vertical splitting is considered. The leaf nodes of the BT are further processed by prediction processing and transformation processing without any further splitting. In JEM, for example, the maximum CTU size is 256×256 luma samples.

[0030] Figure 3A shows an example of block splitting of the QTBT structure, and Figure 3B shows the corresponding tree representation. The solid line indicates QT splitting, and the dotted line indicates BT splitting. At each splitting (i.e., non-leaf splitting) node of the BT, one flag may be signaled to indicate which splitting type is used (i.e., horizontal or vertical). For example, as shown in Figure 3B, 0 indicates horizontal splitting and 1 indicates vertical splitting. For QT splitting, since QT splitting always splits the block into four equal-sized sub-blocks both horizontally and vertically, there is no need to indicate the splitting type.

[0031] The QTBT splitting method supports the flexibility that luma and chroma have separate QTBT structures. Currently, in the case of P slices and B slices, the luma CTB and chroma CTB in one CTU share the same QTBT structure. In contrast, in the case of I slices, the luma CTB is divided into CUs by the QTBT structure, and the chroma CTB is divided into chroma CUs by another QTBT structure. This means that the CU in the I slice consists of a coding block of the luma component or coding blocks of two chroma components, while the CU in the P slice or B slice consists of coding blocks of all three color components.

[0032] In HEVC, the inter prediction of small blocks is restricted and the memory access of motion compensation is suppressed so that bi-prediction is not supported for 4×8 blocks and 8×4 blocks, and inter prediction is not supported for 4×4 blocks. In QTBT implemented in JEM-7.0, such restrictions are removed.

[0033] Figure 4 shows the multi-type-tree (MTT) structure of VCC. In this structure, at the vertex of QTBT, (a) vertical central bilateral ternary splitting and (b) horizontal central bilateral ternary splitting are further added. The advantages of the ternary splitting shown in Figure 4 include complementing quadtree splitting and binary splitting, being able to capture objects located at the center of the block while quadtree and binary splitting always split along the center of the block, and since the width and height of the proposed ternary splitting are always powers of 2, there is no need for further transformation, but it is not limited to these. The two-level tree design motivation is mainly activated by alleviating complexity. Theoretically, the complexity across the tree is T D where T represents the number of splitting types and D represents the depth of the tree.

[0034] In merge mode, the implicitly derived motion information is directly used to generate the prediction samples of the current CU. Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after sending the skip flag and the merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it may be further narrowed down by the signaled MVD information. The MVD information may include a merge candidate flag, an index indicating the motion magnitude, and an index for indicating the motion direction. In merge mode, one of the first two merge candidate flags in the merge list is selected and used as the MV basis. The merge candidate flag is signaled to specify which flag is used.

[0035] The distance index indicates the motion magnitude information and indicates a default offset from the origin. FIG. 5 shows the MMDV search locations for two reference frames according to some embodiments. As shown in FIG. 5, an offset may be added to either the horizontal or vertical component of the origin MV. The relationship between the distance index and the default offset is shown in Table 1 below.

[0036]

Table 1

[0037] The direction index represents the direction of the MVD with respect to the starting point. The direction index may represent one of the four directions shown in Table 2 below. Note that the meaning of the sign of the MVD may be in a different form depending on the information of the starting MV. When the starting MV is a single-prediction MV or a dual-prediction MV, and both lists point to the same side of the current picture (i.e., both POCs of the two references are greater than the POC of the current picture, or both POCs of the two references are less than the POC of the current picture), the sign in Table 2 indicates the sign of the MV offset added to the starting MV. When the starting MV is a dual-prediction MV, the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the difference in POC in list 0 (L0) exceeds that in list 1 (L1), the sign in Table 2 indicates the sign of the MV offset added to the L0 MV component of the starting MV, and the sign of the L1 MV indicates the opposite value. When the difference in L1 in the POC exceeds L0, the sign in Table 2 indicates the sign of the MV offset added to the L1 MV component of the starting MV, and the sign of the L0 MV indicates the opposite value.

[0038] Scaling is performed on the MVD according to the difference in POC in each direction. When the difference in POC in both lists is the same, scaling is not necessary. When the difference in POC in L0 exceeds that in L1, scaling is performed on the MVD of L1. When the POC difference in L1 exceeds L0, scaling is similarly performed on the MVD of L0. When the starting MV is single-predicted, the MVD is added to the available MV.

[0039]

Table 2

[0040] In VVC, in addition to the normal single - direction prediction and bi - direction prediction mode MVD signaling, the symmetric MVD mode of bi - direction MVD signaling may also be applied. In the symmetric MVD mode, motion information including both reference picture indexes of L0 and L1 and the MVD of L1 may be derived (not signaled).

[0041] The decoding process of the symmetric MVD mode is as follows. First, variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived at the slice level. For example, if mvd_l1_zero_flag is 1, BiDirPredFlag is set to 0. If the closest reference picture in L0 and the closest reference picture in L1 form a pair of a reference picture in the front and a reference picture in the back or a pair of a reference picture in the back and a reference picture in the front, BiDirPredFlag is set to 1, and both reference pictures of L0 and L1 are short - term reference pictures. Otherwise, BiDirPredFlag is set to 0. Next, at the CU level, if the CU is encoded in bi - prediction and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. If the symmetric mode flag is true (e.g., equal to 1), only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of L0 and L1 are set to be equal to the pair of reference pictures. Finally, MVD1 is set to be equal to (-MVD0).

[0042] In AV1, for each block encoded inter - frame, when the mode of the current block is not the skip mode but the inter - coding mode, another flag is signaled to notify whether a single - reference mode or a composite - reference mode is used for the current block. In the single - reference mode, a predicted block may be generated by one motion vector. In the composite - reference mode, a predicted block is generated by the weighted average of two predicted blocks derived from two motion vectors. The modes that can be signaled in the case of single - reference are shown in detail in Table 3 below.

[0043]

Table 3

[0044] The modes that can be signaled in the case of compound references are shown in detail in Table 4 below.

[0045]

Table 4

[0046] In AV1, 1 / 8 pixel motion vector accuracy (or precision) is possible. The following syntax may be used to signal the motion vector differences in reference frame L0 or L1. The syntax mv_joint indicates which components of the motion vector difference are non-zero. A syntax mv_joint value of 0 indicates that there is no non-zero MVD along either the horizontal or vertical direction, a value of 1 indicates that there is a non-zero MVD along only the horizontal direction, a value of 2 indicates that there is a non-zero MVD along only the vertical direction, and a value of 3 indicates that there are non-zero MVDs along both the horizontal and vertical directions. The syntax mv_sign indicates whether the motion vector difference is positive or negative. The syntax mv_class indicates the class of the motion vector difference. As shown in Table 5 below, the higher class means that the magnitude of the motion vector difference is larger. The syntax Hhmv_bit indicates the integer part of the offset between the motion vector difference and the starting point of the magnitude of each MV class. The syntax mv_fr indicates the first two fractional bits of the motion vector difference. The syntax mv_hp indicates the third fractional bit of the motion vector difference.

[0047]

Table 5

[0048] In the case of the NEW_NEARMV mode and the NEAR_NEWMV mode (shown in Table 4), the accuracy of the MVD depends on the class related to the MVD and the size of the MVD. First, fractional MVDs are allowed only when the MVD size is 1 pixel or less. Next, when the value of the related MV class is MV_CLASS_1 or higher, only one MVD value is allowed, and the MVD values for each MV class are derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), MV class 2 (MV_CLASS_2), MV class 3 (MV_CLASS_3), MV class 4 (MV_CLASS_4), or MV class 5 (MV_CLASS_5). The allowable MVD values for each MV class are shown in Table 6.

[0049]

Table 6

[0050] In some embodiments, when the current block is encoded using the NEW_NEARMV or NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class. When the current block is not encoded using the NEW_NEARMV or NEAR_NEWMV mode, a different context is used to signal mv_joint or mv_class.

[0051] It may also be indicated whether the MVDs of two reference lists are signaled together by applying a new inter-encoding mode, i.e., JOINT_NEWMV. When the inter-prediction mode is equivalent to the JOINT_NEWMV mode, the MVDs of reference L0 and reference L1 are signaled together. Therefore, it may be possible to signal only one MVD (referred to as joint_mvd) and send it to the decoder, and the ΔMVs of reference L0 and reference L1 may be derived from joint_mvd. The JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context may be required.

[0052] When the JOINT_NEWMV mode is signaled and the POC distances between the current frame and the two reference frames are different, scaling of the MVD is performed for reference L0 or reference L1 based on the POC distance. Specifically, let the distance between reference frame L0 and the current frame be denoted as td0, and the distance between reference frame L1 and the current frame be denoted as td1. When td0 is greater than or equal to td1, joint_mvd is directly used for reference L0, and the mvd of reference L1 is derived from joint_mvd based on the following formula (1).

Number

[0053] When td1 is greater than or equal to td0, joint_mvd is directly used for reference L1, and the mvd of reference L0 is derived from joint_mvd based on the following formula (2).

Number

[0054] A new inter-coding mode (i.e., AMVDMV) may be added to the single-reference example. When the AMVDMV mode is selected, this indicates that AMVD is applied to the signal MVD. A flag, such as amvd_flag, may be added during the JOINT_NEWMV mode to notify whether AMVD is applied to the co-transmission MVD coding mode. When the adaptable MVD resolution is applied to the co-transmission MVD coding mode (referred to as co-transmission AMVD coding), the MVDs of the two reference frames are signaled together, and the accuracy of the MVD is implicitly determined by the MVD magnitude. The MVDs of two (or three or more) reference frames are signaled together, and the conventional MVD coding is applied.

[0055] In the adaptive motion vector resolution (AMVR) initially proposed by CWG-C012, a total of seven MV precisions (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, the AVM encoder explores all supported precision values and signals the highest precision to the decoder. To reduce the encoder execution time, two precision sets are supported. Each precision set contains four default precisions. The precision set is adaptively selected at the frame level based on the maximum precision value of the frame. Similar to AV1, the maximum precision is signaled in the frame header. Table 7 summarizes the supported precision values based on the maximum precision at the frame level.

[0056]

Table 7

[0057] In AVM software (software similar to AV1), there is a frame-level flag to indicate whether the MV of the frame contains precision less than 1 pixel. AMVR is effective only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the precision of a block is less than the maximum precision, the motion model and the interpolation filter are not signaled. When the precision of a block is less than the maximum precision, the motion mode is presumed to be a translational motion, and the interpolation filter is presumed to be a REGULAR interpolation filter. Similarly, when the precision of a block is either 4 pixels or 8 pixels, the inter-intra mode is not signaled and is presumed to be 0.

[0058] In motion compensation, usually, a translational motion model between the reference block and the target block is assumed. On the other hand, an affine model is used for warped motion. The affine motion model can be represented by Equation (3).

Equation

[0059] Here, [x, y] are the coordinates of the original pixel, and [x’, y’] are the distorted coordinates of the reference block. According to Equation (3), up to six parameters are required to identify the distorted motion. a3 and b3 represent the conventional translational MV, a1 and b2 represent the scaling along the MV, and a2 and b1 represent the rotation.

[0060] In global warped motion compensation, global motion information is signaled for each inter-reference frame, which includes the global motion type and the number of motion parameters. The global motion types and the number of associated parameters are listed in Table 9.

[0061] [Table 9]

[0062] After signaling the reference frame index, if global motion is selected, the global motion type and the parameters associated with a given reference frame are used for the current coding block.

[0063] In local warped motion compensation, local warped motion is allowed for an inter-coding block if the following conditions are met. First, single-reference prediction must be used in the current block. The width or height of the coding block must be 8 or more. Finally, the same reference frame as the current block must be used in at least one of the adjacent neighboring blocks.

[0064] When local distortion with motion is used for the current block, the affine model parameters are estimated by minimizing the mean squared difference between the reference and the modeled mapping based on the MVs of the current block and its neighboring blocks. To estimate the local distortion with motion parameters, when using the same reference frame as the current block for neighboring blocks, a set of mapping samples between the central sample in the neighboring blocks and the corresponding sample in the reference frame is obtained. Then, three additional samples are created by shifting the central position by 1 / 4 sample in one or both dimensions. Such additional samples can also be considered as a set of mapping samples for ensuring the stability of the model parameter estimation process.

[0065] The MVs of the neighboring blocks used to derive the motion parameters are referred to as motion samples. The motion samples are selected from neighboring blocks that use the same reference frame as the current block. Note that the distortion with motion prediction mode is valid only for blocks using a single reference frame.

[0066] Figure 6 shows typical motion samples used to derive the model parameters of a block using local distortion with motion prediction according to some embodiments. As shown in Figure 6, the MVs of neighboring blocks B0, B1, and B2 are denoted as MV0, MV1, and MV2 respectively. The current block is predicted using single prediction with reference frame Ref0. Neighboring block B0 is predicted using composite prediction with reference frames Ref0 and Ref1. Neighboring block B1 is predicted using single prediction with reference frame Ref0. Neighboring block B2 is predicted using composite prediction with reference frames Ref0 and Ref2. The motion vector MV0 of B0 Ref0 , the MV1 of B1 Ref0 and the MV2 of B2 Ref0 may be used as motion samples for deriving the affine motion parameters of the current block.

[0067] Either a skip mode or a merge mode using a motion vector expression method may use merge with motion vector difference (MMVD). MMVD also uses the merge candidates of VVC. Candidates may be selected from the merge candidates, and may further be extended by the proposed motion vector expression method. In MMVD, signaling is simplified for the new motion vector expression. The expression method includes a starting point, a motion magnitude, and a motion direction. The MMVD technique uses the VVC merge candidate list. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for the extension of MMVD. FIG. 7 shows, for example, the MMVD search process for the current frame using the two reference frames shown in FIG. 5. The starting point of the motion vector expression method is determined by the basic candidate index. The basic candidate index indicates the best candidate among the candidates in the following list as shown in Table 10.

[0068]

Table 10

[0069] When the number of basic candidates is equal to 1, the basic candidate index is not signaled. The distance index represents the motion magnitude information. The distance index indicates a default distance from the starting point information. The default distance may be as follows in Table 11.

[0070]

Table 11

[0071] The direction of the MVD with respect to the starting point is represented by the direction index. The direction index may represent four directions as shown in Table 12.

[0072]

Table 12

[0073] The MMVD flag may be signaled immediately after sending the skip flag and the merge flag. If the skip and merge flags are true, the MMVD flag is parsed. If the MMVD flag is equal to 1, the MMVD syntax is parsed. On the other hand, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, this case is the AFFINE mode. If the AFFINE flag is not equal to 1, the skip / merge index is parsed in the reference software (e.g., VTM) for the skip / merge mode.

[0074] There are several limitations in the design of local distortion motion in AV1 and the local distortion motion mode. For example, the local distortion present motion mode is only applicable to the case of a single reference picture, and the local distortion present motion model parameters must be derived from the motion samples related to the current block and neighboring blocks without considering the motion model derived from neighboring blocks. Even if the accuracy is not as high as that used to derive the local distortion present motion model parameters, no adjustment is made to the motion samples. In some embodiments, the distortion motion mode may be applied to a composite prediction mode (i.e., multiple reference pictures). In some embodiments, the distortion motion parameters used for the reference pictures in L0 and L1 may be individually derived using the motion vectors of neighboring blocks and the motion vectors of the current block related to the reference pictures in L0 and L1. In some embodiments, one or more distortion models may be derived using the MVs of neighboring blocks and the current block. In one example, when one distortion model is derived, the MV of a neighboring block having the same reference picture as the reference picture in L0 of the current block or the reference picture in L1 of the current block is used. Subsequently, the derived distortion model may be used to generate the warp predictors of the distortions of both reference pictures in L0 and L1. In another example, when two distortion models are derived, the MV of a neighboring block having the same reference picture as the reference picture in L0 of the current block is used to derive the distortion model of the reference picture in L0. Further, the MV of a neighboring block having the same reference picture as the reference picture in L1 of the current block is used to derive the distortion model of the reference picture in L1. Accordingly, two distortion models are used to generate the distortion prediction values of the reference pictures in L0 and L1 correspondingly.

[0075] In some embodiments, a weighted average of the distortion predictions of the reference pictures L0 and L1 may be used as the final prediction value.

[0076] In some embodiments, the distortion mode parameter may be derived for only one of the plurality of reference pictures.

[0077] In some embodiments, whether the distortion motion model is applied may be signaled separately for the reference pictures in L0 and the reference pictures in L1. Whether the distortion motion model is applied to the reference picture list may be implicitly derived by encoded information including, but not limited to, the neighboring block MV, the neighboring block motion model (whether the distortion motion model is applied, whether the translational motion model is applied, whether the optical flow motion model is applied), etc.

[0078] In some embodiments, the distortion motion model may be shared between the reference picture L0 and the reference picture L1. That is, the distortion motion model may have a mirror-image relationship between the reference picture L0 and the reference picture L1.

[0079] In some embodiments, the distortion motion may be applied to only one of the reference frames (or pictures) (either the reference frame in L0 or the reference frame in L1). In some embodiments, the translational motion is applied to other reference lists (reference lists to which the distortion motion is not applied), and a weighted average of the prediction sample obtained from the distortion motion and the prediction sample obtained from the translational motion is used as the final prediction of the current block. In some embodiments, a weighted average of the distortion motion and the intra prediction is used as the final prediction of the current block. The intra prediction mode used for the current block may be included in the bitstream as a prefix, or the intra prediction mode may be signaled and included in the bitstream. In one example, when used as a prefix, the intra prediction mode is used as the DC prediction mode, i.e., the smooth prediction mode. In some embodiments, other averaging methods, such as a wedge composite mode of the distortion-with-motion and the translational motion / intra prediction, are used for the final prediction of the current block.

[0080] In some embodiments, the distortion motion model may be inherited from neighboring blocks in a single reference picture or blocks encoded in a local distortion mode. A distortion model list may be created on both the encoder side and the decoder side of video encoding. The entries of the distortion model list are the distortion models of the blocks neighboring the current block. The neighboring blocks may be, but are not limited to, those that are adjacent and spatially neighboring, those that are not adjacent but spatially neighboring, those that are temporally neighboring, etc. The size of the distortion model list may be invariant (it may be predefined or signaled with a high-level flag) or may vary according to the situation. An index indicating which entry is used for distortion prediction may be signaled and included in the bitstream. The distortion model generated using the surrounding MVs and the current MV may also be an entry of the distortion model list. One or more generated distortion models (generated depending on different rules) may be elements of the distortion model list. To insert a distortion model candidate into the distortion model list, for example, the reference picture of the neighboring block must be the same as that of the current block. Unnecessary ones may be removed from those inserted into the distortion model list so that the same model is not allowed in the distortion model list. A model bank may be created on the encoder side and / or the decoder side. The distortion model bank may be updated in an on-the-fly manner after encoding or after decoding the encoded blocks. Candidates obtained from the bank may be inserted into the distortion model list.

[0081] In some embodiments, when a neighboring block is encoded using a distorted motion mode, the related distorted motion model parameters may be carried over to the current block. The carried-over motion model may be used as a predicted value of the distorted motion model of the current block (i.e., the warp motion predictor (WMP)). According to some embodiments, the carried-over motion model may be used as a predicted value, or the predicted value may be directly used as the distorted motion model of the current block. According to some embodiments, the carried-over motion model may be used as a predicted value, and the difference between the actual distorted motion model parameters and the motion model predicted value (i.e., the warp motion difference (WMD)) may be signaled for the current block.

[0082] In some embodiments, when the distorted motion model is carried over from a block encoded in a local distortion mode in one reference picture, the block is identified by a displacement vector that points to the block in the reference picture from the current block. This displacement vector may be signaled or derived implicitly. When multiple blocks are used to derive the carried-over distorted motion model, the selection of the blocks may be signaled or determined based on a predefined rule.

[0083] In some embodiments, a correction value (i.e., a Δ value) is signaled for local distortion motion model parameters derived from motion samples. For example, signaling when there are a total of N parameters in the local distortion motion model (N includes 2, 4, 6, Δ values in addition to the selected M parameters out of the N parameters, but is not limited thereto) may notify the actual parameters used in the local distortion motion model. The Δ value signaled for a selected distortion motion model parameter may be used as a predicted value of the Δ value signaled for another selected distortion motion model parameter. In some embodiments, when local distortion motion is applied for a composite prediction mode, the Δ value signaled for a certain reference picture may be used as a predicted value of the Δ value for another reference picture.

[0084] In some embodiments, the predicted value may be a Δ value in a mirrored relationship and / or a Δ value in a mirrored relationship multiplied by a scaling factor. The scaling factor may depend on the distance of the reference picture from the current picture.

[0085] In some embodiments, motion vector differences may be signaled for one or more of the motion samples, and the related motion vectors are added with the motion vector differences signaled before being used as motion samples to derive the local distortion motion model. The motion vector differences may be signaled using the MMVD approach. In some embodiments, the motion vector differences are signaled for sub-block motion vectors or sample-based motion vectors derived based on the distortion model used at that time.

[0086] FIG. 8 is a flowchart showing a method 810 for video encoding executed by at least one processing means according to an embodiment.

[0087] In some implementation examples, one or more processing blocks in FIG. 8 may be executed by computer 102. In some implementation examples, one or more processing blocks in FIG. 8 may be executed by another device or a group of devices separated from computing environment 600, or may be executed by another device or a group of devices included in computing environment 600.

[0088] As shown in FIG. 8, in operation 811, method 810 may include obtaining video data.

[0089] In operation 812, method 810 may include parsing and blocking the video data, where the blocks are related to a reference picture list.

[0090] In operation 813, method 810 may include generating a distortion model for the first reference picture list and the second reference picture list of the current block included in the reference picture list based on the motion vector of the current block and the motion vectors of the blocks adjacent to the current block.

[0091] In operation 814, method 810 may include decoding the frames in the first reference picture list and the second reference picture list by applying the distortion model to the frames.

[0092] Although FIG. 8 shows an example of method blocks, in some implementation examples, the method may include more blocks, fewer blocks, different blocks, or blocks in a different arrangement compared to those illustrated in FIG. 8. In addition to or instead of the above, two or more of the method blocks may be executed in parallel.

[0093] FIG. 9 is a block diagram of an example of computer code 910 for video encoding according to an embodiment. In an embodiment, the computer code may be, for example, program code or computer program code. According to an embodiment of the present disclosure, an apparatus / device including at least one processor together with a memory storing the computer program code may be provided. The computer program code may be configured to execute any number of aspects of the present disclosure when executed by at least one processor.

[0094] As shown in FIG. 9, the computer code 910 includes an acquisition code 911, a parse code 912, a generation code 913, and a decoding code 914.

[0095] The acquisition code 911 is configured to cause at least one processor to acquire video data.

[0096] The parse code 912 is configured to cause at least one processor to parse the acquired video data into blocks. Each of the blocks may be associated with a reference picture list.

[0097] The generation code 913 is configured to cause at least one processor to generate a distortion model used for a first reference picture list and a second reference picture list of the current block included in the reference picture list based on the motion vector of the current block and the motion vectors of the blocks adjacent to the current block.

[0098] The decoding code 914 is configured to cause at least one processor to decode the frames in the first reference picture list and the second reference picture list by applying the distortion model to the frames.

[0099] FIG. 9 shows an example of a block of code. However, in some implementations, the apparatus / device may include a greater number of blocks, a fewer number of blocks, different blocks, or blocks in a different arrangement compared to what is illustrated in FIG. 9. In addition to or instead of the above, two or more of the blocks of the apparatus may be combined. In other words, although FIG. 9 shows separate blocks of code, various code instructions need not be separate and may be intermixed.

[0100] The embodiments described in this description may be used individually or combined in any order. Further, each of the method (i.e., embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. For example, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU).

[0101] FIG. 10 is a block diagram 500 of the internal and external components of the computer illustrated in FIG. 1 according to an embodiment. It should be understood that FIG. 10 merely shows an example of one implementation and does not imply any limitation to the environment in which different embodiments may be implemented. Based on design requirements and implementation requirements, numerous modifications to the illustrated environment may be made.

[0102] Computer 102 (FIG. 1) and server computer 114 (FIG. 1) may each include a respective set of internal components 800A,B and external components 900A,B. Each set of internal components 800 includes one or more processors 820, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824 on one or more buses 826, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.

[0103] Processor 820 may be implemented in hardware, software, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processor. In some implementations, processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that enable communication between internal components 800A, B.

[0104] One or more operating systems 828, software program 108 (FIG. 1), and video encoding program 116 (FIG. 1) on server computer 114 (FIG. 1) may be stored in one or more of respective computer-readable tangible storage devices 830 for execution by one or more of respective processors 820 via one or more of respective RAMs 822 (which typically includes cache memory). In the embodiment shown in FIG. 10, each of computer-readable tangible storage devices 830 may be a magnetic disk storage device of an internal hard disk. In some embodiments, each of computer-readable tangible storage devices 830 may be a semiconductor storage device such as ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disk (CD), digital versatile disk (DVD), floppy disk, cartridge, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device capable of storing computer programs and digital information.

[0105] Each set of internal components 800A, B may also include an R / W drive or interface 832 for reading from and writing to one or more portable computer-readable tangible memory devices 936 such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor memory devices. Software programs such as software program 108 (FIG. 1) and video encoding program 116 (FIG. 1) can be stored in one or more of the respective portable computer-readable tangible memory devices 936, read out via the respective R / W drives or interfaces 832, and loaded onto respective hard disks, such as storage device 830.

[0106] Each set of internal components 800A, B may also include a network adapter or interface 836 such as a TCP / IP adapter card, a wireless Wi-Fi interface card, or a 3G, 4G, or 5G wireless interface card, and other wired or wireless communication links. The software program 108 (FIG. 1) and the video encoding program 116 (FIG. 1) on the server computer 114 may be downloaded from an external computer to the computer 102 (FIG. 1) and the server computer 114 via a network (such as the Internet, a local area network, or other wide area network) and the respective network adapters or interfaces 836. The software program 108 and the video encoding program 116 on the server computer 114 may be loaded from the network adapter or interface 836 onto respective hard disks, such as storage device 830. The network may include copper wires, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers.

[0107] Each of the sets of external components 900A, B may include a computer display monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, B may also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each of the sets of internal components 800A, B may also include a device driver 840 that interfaces with the computer display monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 comprise hardware and software (stored in the storage device 830 and / or the ROM 824).

[0108] This disclosure includes a detailed description of cloud computing, but it will be understood prior to the description that the implementation of the teachings described herein is not limited to a cloud computing environment. Further, some embodiments can be implemented in cooperation with any other type of computing environment, whether currently known or later developed.

[0109] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing means, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0110] The characteristics are as follows.

[0111] On-demand self-service: Without the need for human interaction with the service provider, the cloud consumer can automatically and as needed provision computing capabilities such as server time and network storage unilaterally.

[0112] Broad network access: The functions are available using a network, and access to the functions is through a standard mechanism that facilitates use by heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0113] Resource pooling: The provider's computing resources are pooled and services are provided to multiple consumers using a multi-tenant model with various physical and virtual resources that are dynamically allocated and reallocated in response to requests. The consumer has little or no control over or knowledge of the exact location of the provided resources, but may be able to specify the location at a higher level of abstraction (e.g., country, state, data center), meaning it is location-independent.

[0114] Rapid elasticity: The functions can be provisioned quickly and with high scalability, and in some cases automatically, to scale out rapidly, be released quickly, and scale in rapidly. In many cases, the functions available for provisioning appear to be unlimited for the consumer, and any number can be purchased at any time.

[0115] Measured service: By leveraging measurement functions at a predetermined level of abstraction suitable for the type of service, the cloud system automatically controls and optimizes the use of resources (e.g., storage, processing means, bandwidth, number of active user accounts). It monitors, controls, and reports on how resources are used, achieving transparency for both the provider and consumer of the services being utilized.

[0116] The service model is as follows.

[0117] Software as a Service (SaaS): The functionality provided to the consumer uses the provider's applications running on the cloud infrastructure. The applications are accessible through a thin client interface such as a web browser from various client devices (e.g., web - based email). The underlying cloud infrastructure includes the network, servers, operating systems, and storage, and even includes individual application functionality (except for special application configuration settings specific to the user), but the consumer does not manage or control the underlying cloud infrastructure.

[0118] Platform as a Service (PaaS): The functionality provided to the consumer is deployed on a cloud infrastructure created by the consumer using programming languages and tools supported by the provider or by existing applications. The consumer does not manage or control the underlying cloud infrastructure that includes the network, servers, operating systems, or storage, but controls the deployed applications and, in some cases, the application hosting environment configuration.

[0119] Infrastructure as a Service (IaaS): The functions provided to consumers perform the provisioning of processing means, storage, networks, and other basic computing resources. Consumers can deploy and execute any software that can include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but control the operating systems, storage, and deployed applications, and in some cases perform limited control of selected network components (such as host firewalls).

[0120] The deployment models are as follows.

[0121] Private cloud: The cloud infrastructure is operated only for an organization. The private cloud can be managed by the organization or a third party and can exist on-premises or off-premises.

[0122] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community that shares considerations (such as roles, security requirements, policies, and compliance considerations). The community cloud can be managed by the organization or a third party and can exist on-premises or off-premises.

[0123] Public cloud: The cloud infrastructure is available to the public or large industrial groups and is owned by the organization that sells cloud services.

[0124] Hybrid Cloud: The infrastructure is composed of two or more clouds (private, community, or public), and these two or more clouds continue to be independent entities, but are combined by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0125] The cloud computing environment is a service-oriented environment centered around statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure with a network of interconnected nodes.

[0126] Referring to FIG. 11, an illustrative view of a cloud computing environment 600 that may be suitable for implementing some embodiments of the disclosed protection present is shown. As illustrated, the cloud computing environment 600 includes one or more cloud computing nodes 10, and local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA), a cellular telephone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automotive computer system 54N, etc., may communicate with the one or more cloud computing nodes 10. The cloud computing nodes 10 may communicate with each other. The cloud computing nodes 10 may be physically or virtually grouped within one or more networks such as the private cloud, community cloud, public cloud, hybrid cloud, or combinations thereof (not shown). Thereby, the cloud computing environment 700 can provide infrastructure, platform, and / or software as services such that cloud consumers do not need to maintain resources on local computing devices. It is intended that the types of computing devices 54A-N shown in FIG. 11 are merely illustrative, and it can be seen that the cloud computing nodes 10 and the cloud computing environment 600 can communicate with any type of device that is computerized using any type of network and / or connection capable of specifying a network address (e.g., one using a web browser).

[0127] Referring to FIG. 12, a set of functional abstraction layers 700 provided by the cloud computing environment 600 (FIG. 11) is shown. It is intended that the components, layers, and functions shown in FIG. 12 are merely illustrative, and it will be readily understood from the description prior to any explanation that embodiments are not limited thereto. As shown, the following layers and corresponding functions are provided.

[0128] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0129] The virtualization layer 70 provides an abstraction layer from which examples of virtual entities such as virtual server 71, virtual storage 72, virtual network 73 (including virtual private network), virtual application and operating system 74, and virtual client 75 can be provided.

[0130] In one example, the management layer 80 can implement the following functions. Resource provisioning 81 realizes the dynamic procurement of computing resources and other resources used to execute tasks within the cloud computing environment. Metered pricing 82 realizes the tracking of costs when resources are used within the cloud computing environment and the charging and billing for the consumption of the resources. In one example, these resources may be equipped with application software licenses. Security realizes the identification of cloud consumers and tasks and the protection of data and other resources. The user portal 83 realizes the access of consumers and system administrators to the cloud computing environment. Service level management 84 realizes the allocation and management of cloud computing resources so that the required service level is met. Service Level Agreement (SLA) planning fulfillment 85 realizes the pre-configuration and procurement of cloud computing resources that are expected to be the targets of future requirements according to the SLA.

[0131] The workload layer 90 implements examples of functions that can utilize a cloud computing environment. Examples of workloads and functions that can be implemented by this layer include mapping and navigation 91, software development and life cycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and video encoding / decoding 96. Video encoding / decoding 96 can encode / decode video data using a delta angle derived from a nominal angle.

[0132] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible technical detail level of integration. The computer-readable media may include one computer-readable non-transitory storage medium (or multiple media) that stores computer-readable program instructions for causing a processor to perform operations.

[0133] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. Exemplary listings of more specific examples of computer-readable storage media include portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, punched cards, mechanically encoded devices such as raised structures within grooves that record instructions, and any suitable combination of the foregoing. A computer-readable storage medium as used in this disclosure is not considered, by its nature, to be a transient signal such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0134] The computer-readable program instructions described in this disclosure can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded from an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network, and the network adapter card or network interface transfers the computer-readable program instructions for storage in a computer-readable storage medium in each respective computing / processing device.

[0135] The computer-readable program code / instructions for performing the operations may be source code or object code described in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit setting data, and object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language and similar programming languages. The entire computer-readable program instructions may be executed on the user's computer, or a part of the computer-readable program instructions may be executed on the user's computer, or the computer-readable program instructions may be executed as a stand-alone software package, or a part of the computer-readable program instructions may be executed on the user's computer and a part on a remote computer, or the entire computer-readable program instructions may be executed on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet service provider). In some embodiments, for example, by using an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA), the computer-readable program instructions may be executed by customizing the electronic circuit using the state information of the computer-readable program instructions in order to perform an aspect or operation.

[0136] Such computer-readable program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine that implements the functions / operations shown in one or more blocks of the flowchart and / or block diagram through the processor of the computer or other programmable data processing apparatus. A computer-readable storage medium storing the instructions may store the above computer-readable program instructions in a computer-readable storage medium that can be instructed to function in a specific manner in a computer, a programmable data processing apparatus, and / or other devices so as to include a product including instructions for implementing the modes of the functions / operations shown in one or more blocks of the flowchart and / or block diagram.

[0137] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, and other devices so that operations and steps are performed on the computer, other programmable apparatus, and other devices in accordance with the instructions executed on the computer, other programmable apparatus, and other devices, thereby generating a computer-implemented process, such that the functions / operations shown in one or more blocks of the flowchart and / or block diagram are implemented.

[0138] It will be apparent that the systems and / or methods described in this disclosure may be implemented in individual forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware code or software code used to implement such systems and / or methods is not limiting. Thus, in this disclosure, the operations and actions of the systems and / or methods are described without reference to specific software code, and it can be seen that software and hardware can be designed to implement the systems and / or methods based on the description of this disclosure.

[0139] Although the description has been presented for purposes of illustration of various aspects and embodiments, the description is not intended to be limiting or to be limited to the disclosed embodiments. Even if a combination of features is recited in a claim and / or disclosed herein, that combination is not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not expressly recited in a claim and / or not expressly disclosed in the specification. Each of the dependent claims listed hereinafter can depend directly on only one claim, but combinations of each dependent claim with any other claim in the set of claims are included in the disclosure of possible implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used in this disclosure are chosen to best explain the principles of the embodiments, the practical application, or a technical improvement beyond the technologies found in the marketplace, and to enable others of ordinary skill in the art to understand the embodiments disclosed in this disclosure.

[0140] Although several exemplary embodiments have been described in this disclosure, there are variations, substitutions, and various alternative equivalents that are within the scope of this disclosure. Thus, it can be seen that those of ordinary skill in the art can implement the principles of this disclosure and thus envision numerous systems and methods within its spirit and scope without being explicitly shown or described in this application.

Description of Reference Numerals

[0141] 10 cloud computing nodes, 54A cellular phone, 54B desktop computer, 54C laptop computer, 54N automotive computer system, 60 software layer, 61 mainframe, 62 architecture-based server, 63 server, 64 blade server, 65 memory device, 66 network component, 67 network application server software, 68 database software, 70 virtualization layer, 71 virtual server, 72 virtual storage, 73 virtual network, 74 operating system, 75 virtual client, 80 management layer, 81 resource provisioning, 82 metering pricing, 83 user portal, 84 service level management, 85 plan fulfillment, 90 workload layer, 91 navigation, 92 life cycle management, 93 virtual classroom education delivery, 94 data analysis processing, 95 transaction processing, 96 decoding, 100 video encoding system, 102 computer, 104 processor, 106 data storage device, 108 software program, 110 communication network, 112 database, 114 server computer, 116 video encoding program, 500 block diagram, 600 cloud computing environment, 700 cloud computing environment, 700 set of functional abstraction layers, 800A internal component, 800B internal component, 810 method, 820 processor, 822 computer-readable RAM, 824 computer-readable ROM, 826 bus, 828 operating system, 830 computer-readable tangible storage device, 832 R / W drive or interface, 836 network adapter or interface, 840 device driver, 900A external component, 900B external component, 910 computer code, 911 acquisition code, 912 parse code, 913 generation code, 914 decoding code, 920 computer display monitor, 930 keyboard, 934 computer mouse, 936 portable computer-readable tangible storage device

Claims

1. 1. A method for video encoding executed by at least one processor, the method comprising: obtaining video data comprising a plurality of blocks, each block of the plurality of blocks associated with a first reference picture list and a second reference picture list; generating distortion models to be used for the first reference picture list and the second reference picture list of the current block based on the motion vector of the current block and the motion vectors of neighboring blocks adjacent to the current block; and decoding frames in the first reference picture list and the second reference picture list by applying the distortion model to the frames.

2. determining whether the motion vector of the neighboring block has the same reference picture as a reference picture in the first reference picture list of the current block or a reference picture in the second reference picture list of the current block; generating the distortion model based on the same reference picture based on determining that the motion vectors of the neighboring blocks have the same reference picture; The method of claim 1 further comprising:

3. The method of claim 1 , further comprising generating distortion model predictions for the first and second reference picture lists of the current block.

4. applying the distortion model separately to a first frame in the first reference picture list and a second frame in the second reference picture list; applying the distortion model to a frame in one of the first reference picture list or the second reference picture list; or applying another distortion model obtained from one of the neighboring blocks to the frames in the first reference picture list and the second reference picture list of the current block. The method of claim 1 , further comprising at least one of:

5. The method of claim 1 , wherein the distortion model is applied to the video data in response to determining that the video data is predicted in a mixed prediction mode.

6. The method of claim 1 , further comprising receiving a signal indicative of a Δ value used to correct parameters of the distortion model, the Δ value being derived from the motion vector.

7. generating a distortion model list containing distortion models of the neighboring blocks; adding the generated distortion model to the distortion model list in association with the current block; The method of claim 1 further comprising:

8. A device configured to perform the method of any one of claims 1-7.

9. A computer program for causing a computer to carry out the method according to any one of claims 1 to 7.