Global motion for merge mode candidates in inter prediction

By constructing a merge candidate list and adding a global motion vector in the video codec, the problems of low video compression efficiency and high decoding complexity in the existing technology are solved, and more efficient video compression and quality improvement are achieved.

CN114073083BActive Publication Date: 2025-10-03DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080046624.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-25
Filing Date
2020-04-24
Publication Date
2025-10-03
Estimated Expiration
2040-04-24

AI Technical Summary

Technical Problem

Existing video coding methods suffer from information loss during the compression process, resulting in the decompressed video quality being lower than the original video quality. In addition, the encoding and decoding algorithms are highly complex, making it difficult to effectively utilize global motion information for efficient compression.

Method used

When constructing the merge candidate list in the decoder, the global motion vector is added to improve compression efficiency. By signaling the global motion information in the header and using the global motion vector as the prediction candidate, the number of bits of the motion vector difference is reduced, reducing the coding complexity.

Benefits of technology

By using global motion vectors to improve motion vector coding, the bit rate is reduced, the compression efficiency is improved, the computational complexity is reduced, and the decompressed video quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114073083B_ABST
    Figure CN114073083B_ABST
Patent Text Reader

Abstract

A decoder includes circuitry configured to: receive a bitstream; use the bitstream to determine whether merge mode is enabled for a current block; construct a merge candidate list, including adding a global motion vector to the motion vector candidate list; and reconstruct pixel data for the current block using the motion vector candidate list. Related devices, systems, techniques, and articles of manufacture are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 838,618, filed on April 25, 2019, entitled “GLOBAL MOTION FOR MERGE MODE CANDIDATES IN INTER PREDICTION,” which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present invention generally relates to the field of video compression and, in particular, to global motion of merge mode candidates in inter-frame prediction. Background Art

[0004] A video codec may include electronic circuitry or software that compresses or decompresses digital video. A video codec can convert uncompressed video into a compressed format, or it can decompress compressed video back into an uncompressed format. In the case of video compression, the device that compresses the video (and / or performs certain functions thereof) may generally be referred to as an encoder, and the device that decompresses the video (and / or performs certain functions thereof) may be referred to as a decoder.

[0005] The format of the compressed data can conform to standard video compression specifications. Compression can be lossy, as some information present in the source video is missing from the compressed video. Consequently, the decompressed video may be of lower quality than the original, uncompressed video, as there is insufficient information to accurately reconstruct the original video.

[0006] There may be a complex relationship between video quality, the amount of data used to represent the video (e.g., determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, end-to-end delay (e.g., latency), etc.

[0007] Motion compensation can include methods that predict a video frame or portion of a video frame based on a given reference frame, such as a previous frame and / or a future frame, by taking into account the motion of the camera and / or objects in the video. This method can be used in encoding and decoding video data for video compression, such as in encoding and decoding using the Moving Picture Experts Group (MPEG)-2 (also known as Advanced Video Coding (AVC) and H.264) standards. Motion compensation can describe a picture based on a transformation from a reference picture to the current picture. The reference picture can be a picture that precedes the current picture in time or a picture that is in the future compared to the current picture. Compression efficiency can be improved when an image can be accurately synthesized from previously transmitted and / or stored images. Summary of the Invention

[0008] In one aspect, a decoder includes circuitry configured to: receive a bitstream; use the bitstream to determine whether merge mode is enabled for a current block; construct a merge candidate list, wherein constructing the merge candidate list further includes adding a global motion vector to the motion vector candidate list; and reconstruct pixel data for the current block using the motion vector candidate list.

[0009] In another aspect, a method includes receiving a bitstream by a decoder; using the bitstream to determine whether a merge mode is enabled for a current block; constructing a merge candidate list, wherein constructing the merge candidate list further includes adding a global motion vector to the motion vector candidate list; and reconstructing pixel data of the current block using the motion vector candidate list.

[0010] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below.Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] For the purpose of illustrating the invention, the drawings show aspects of one or more embodiments of the invention. It should be understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown in the drawings, in which:

[0012] Figure 1 is a diagram illustrating motion vectors for an example frame with global and local motion;

[0013] Figure 2 illustrates three example motion models that may be used for global motion, including index values ​​(0, 1, or 2) for the three example motion models;

[0014] Figure 3 is a process flow diagram according to some example implementations of the current subject matter;

[0015] Figure 4 is a system block diagram of an example decoder according to some example implementations of the current subject matter;

[0016] Figure 5 is a process flow diagram according to some example implementations of the current subject matter;

[0017] Figure 6 is a system block diagram of an example encoder according to some example implementations of the current subject matter; and

[0018] Figure 7 is a block diagram of a computing system that may be used to implement any one or more of the methods disclosed herein, and any one or more portions thereof.

[0019] The accompanying drawings are not necessarily drawn to scale and may be illustrated by dotted lines, diagrammatic representations, and partial views. In some cases, details that are not necessary for understanding the embodiments or that make other details difficult to perceive may have been omitted. The same reference numerals in the various figures represent the same elements. DETAILED DESCRIPTION

[0020] Global motion in a video refers to motion that occurs across the entire frame. Global motion can be caused by camera motion; for example, a camera pan and zoom can produce motion within a frame that typically affects the entire frame. Motion that exists in certain parts of a video can be referred to as local motion. Local motion can be caused by moving objects in the scene. For example, an object moves from left to right across the scene. A video may contain a combination of local and global motion. Some implementations of the current subject matter can provide efficient methods for communicating global motion to a decoder and using global motion vectors to improve compression efficiency.

[0021] Figure 1 1 is a diagram illustrating motion vectors for an example frame 100 having global motion and local motion. Frame 100 may include a plurality of pixel blocks illustrated as squares and motion vectors associated with the plurality of pixel blocks illustrated as arrows. Squares (e.g., pixel blocks) with arrows pointing to the upper left indicate that motion in those blocks is considered global motion, and squares (indicated by 104) with arrows pointing in other directions indicate that motion in those blocks is considered local motion. Figure 1 In the illustrated example of , many blocks have the same global motion. Global motion is signaled in a header such as a picture parameter set (PPS) and / or sequence parameter set (SPS), and using this signaled global motion can reduce the motion vector information required for the blocks and can achieve improved prediction. Although the examples described below for illustrative purposes refer to determining and / or applying global motion vectors or local motion vectors at the block level, global motion vectors can be determined and / or applied for any region of a frame and / or picture, including: a region consisting of multiple blocks; a region bounded by any geometric shape, such as, but not limited to, a region defined by geometric and / or exponential coding, where one or more lines and / or curves defining the shape can be angled and / or curved; and / or an entire frame and / or picture. Although signaling is described herein as being performed at the frame level and / or in a header and / or parameter set of a frame, signaling can alternatively or additionally be performed at a sub-picture level, where a sub-picture can include any region of a frame and / or picture as described above.

[0022] As an example, still refer to Figure 1, a simple translational motion can be described using a motion vector (MV) having two components MVx, MVy, which describes the displacement of blocks and / or pixels in the current frame. More complex motions such as rotation, scaling and / or distortion can be described using affine motion vectors, wherein the "affine motion vector" used in this disclosure describes a uniform displacement of a group of pixels or points in a video picture and / or picture, such as the illustrated group of pixels showing that during motion, an object moves across the view in the video but its apparent shape does not change. Some video encoding and / or decoding methods may use a four-parameter affine model or a six-parameter affine model for motion compensation in inter-picture coding.

[0023] Further references Figure 1 And as an example, a six-parameter affine motion can be described as:

[0024] x'=ax+by+c

[0025] y'=dx+ey+f

[0026] The four-parameter affine motion can be described as:

[0027] x'=ax+by+c

[0028] y'=-bx+ay+f

[0029] Where (x, y) and (x', y') are the pixel positions in the current image and the reference image respectively; a, b, c, d, e, and f are the parameters of the affine motion model.

[0030] Still refer to Figure 1 , parameters describing affine motion can be signaled to the decoder to apply affine motion compensation at the decoder. In some methods, the motion parameters can be signaled explicitly or by signaling translation control point motion vectors (CPMVs) and then deriving the affine motion parameters from the translation motion vectors. The affine motion parameters for a four-parameter affine motion model can be derived using two control point motion vectors (CPMVs), and the parameters for a six-parameter motion model can be obtained using three control point translation motion vectors (CPMVs). Signaling the affine motion parameters using control point motion vectors can allow the use of efficient motion vector coding methods to signal the affine motion parameters.

[0031] In the embodiment, still refer to Figure 1The sps_affine_enabled_flag in the PPS and / or SPS can specify whether affine-based motion compensation can be used for inter-frame prediction. If sps_affine_enabled_flag is equal to 0, the syntax can be constrained so that affine-based motion compensation is not used in the later coded video sequence (CLVS), and the inter_affine_flag and cu_affine_type_flag will not be present in the coding unit syntax of the CLVS. Otherwise (sps_affine_enabled_flag is equal to 1), affine-based motion compensation can be used in the CLVS.

[0032] Continue to refer to Figure 1 The sps_affine_type_flag in the PPS and / or SPS can specify whether motion compensation based on a 6-parameter affine model can be used for inter prediction. If sps_affine_type_flag is equal to 0, the syntax can be constrained so that motion compensation based on a 6-parameter affine model is not used in CLVS, and the cu_affine_type_flag will not be present in the coding unit syntax in CLVS. Otherwise (sps_affine_type_flag is equal to 1), motion compensation based on a 6-parameter affine model can be used in CLVS. When not present, the value of sps_affine_type_flag can be inferred to be equal to 0.

[0033] Continue to refer to Figure 1 , processing the MV prediction candidate list can be a step performed in some compression methods that utilize motion compensation at the decoder. Some methods may define the use of spatial motion vector candidates and temporal motion vector candidates. Global motion signaled in a header such as SPS or PPS may indicate the presence of global motion in the video. Such global motion may be expected to be common to most blocks in a frame. Motion vector coding may be improved and bit rate may be reduced by using global motion as a prediction candidate. The candidate MVs added to the MV prediction list may be selected based on the motion model used to represent global motion and the motion model used in inter-frame coding.

[0034] Still refer to Figure 1 Some implementations use one or more control point motion vectors (CPMVs) to describe global motion, depending on the motion model used. Thus, depending on the motion model used, one to three control point motion vectors may be available and used as candidates for prediction. In some implementations, all available CPMVs may be added to the list as prediction candidates. Adding the general case of all available CPMVs can increase the likelihood of finding a good motion vector prediction and improve compression efficiency.

[0035] Table 1:

[0036]

[0037]

[0038] Continue to refer to Figure 1 , creating a motion vector (MV) prediction candidate list is a step in performing motion compensation at the decoder.

[0039] Still refer to Figure 1 Global motion, signaled in headers such as the SPS, can indicate the presence of global motion in a video. This global motion is likely to be present in many blocks within a frame. Therefore, a given block is likely to have motion similar to the global motion. Motion vector coding can be improved and bitrate reduced by using global motion as a prediction candidate. Candidate MVs can be added to an MV prediction list and selected based on the motion model used to represent global motion and the motion model used in inter-frame coding.

[0040] Still refer to Figure 1 Depending on the motion model used, one or more control point motion vectors (CPMVs) may be used to describe the global motion. Therefore, depending on the motion model used, one to three control point motion vectors may be available and used as candidates for prediction. The list of MV prediction candidates can be reduced by selectively adding a CPMV to the list. Reducing the list size can reduce computational complexity and improve compression efficiency.

[0041] In some implementations, reference is still made to Figure 1 , the CPMV selected as a candidate may be based on a predefined mapping, such as the mapping shown in Table 3 below:

[0042]

[0043]

[0044] Continue to refer to Figure 1 , the use of selective prediction candidates from global motion can be signaled in the picture parameter set of the sequence parameter set, thereby reducing encoding and decoding complexity.

[0045] Still refer to Figure 1 , since the block is likely to have motion similar to the global motion, adding the global motion vector as the first candidate in the prediction list can reduce the bits necessary to signal the prediction candidate used and to encode the motion vector difference.

[0046] Continue to refer to Figure 1 , creating a motion vector (MV) prediction candidate list can be a step in performing motion compensation at a decoder.

[0047] Still refer to Figure 1 Global motion, signaled in headers such as the SPS, can indicate the presence of global motion in a video. This global motion is likely present in many blocks within a frame. Therefore, a given block is likely to have motion similar to the global motion. Using global motion as a prediction candidate can improve motion vector coding and reduce bitrate.

[0048] Further references Figure 1 , for example, the candidate MV to be added to the MV prediction list may be adaptively selected according to which control point motion vector (CPMV) is selected as a prediction candidate in a neighboring block (eg, prediction unit (PU)).

[0049] For example, still referring to Figure 1 If a neighboring PU uses a specific CPMV as the prediction MV for a specific control point, the CPMV can be added to the MV candidate list. If a neighboring PU uses more than one CPMV as the prediction MV (for example, the left PU uses CPMV0 and the top PU uses CPMV1), all CPMVs used as prediction candidates can be added to the list. If the neighboring PU does not use a CPMV, no CPMV is added to the prediction list. In some implementations, global motion information can be added as the first candidate in the prediction list.

[0050] Continue to refer to Figure 1 , the use of adaptive prediction candidates from global motion can be signaled in a header such as a picture parameter set (PPS) or a sequence parameter set (SPS), thereby reducing encoding and decoding complexity.

[0051] Still refer to Figure 1 Since blocks may have motion similar to the global motion, adding the global motion vector as the first candidate in the prediction list can reduce the bits required to signal the prediction candidate used and encode the motion vector difference. Creating a motion vector (MV) prediction candidate list can be a step in performing motion compensation at the decoder.

[0052] Still refer to Figure 1Global motion, signaled in headers such as PPS or SPS, can indicate the presence of global motion in a video. This global motion is likely present in many blocks within a frame. Therefore, a given block is likely to have motion similar to the global motion. Using global motion as a prediction candidate can improve motion vector coding and reduce bitrate. Furthermore, using global motion models and control points as merge candidates can improve inter-frame coding and reduce bitrate.

[0053] In some implementations, further reference is made to Figure 1 When global motion is signaled in a header such as a PPS or SPS, the motion model and control point motion vector (CPMV) (collectively referred to as a global motion vector) can be added to the merge candidate list. Using global motion merge candidates can improve compression performance.

[0054] In some implementations, reference is still made to Figure 1 , strong global motion is likely to result in most blocks having global motion. In such cases, compression efficiency can be improved by signaling the use of global_merge_mode for a given block (e.g., prediction unit (PU)) within a frame. When global_merge_mode is signaled, global motion parameters can be used for motion compensation in the block (e.g., PU) and no merge candidate list can be created. This approach can also reduce computational complexity at the decoder. Figure 2 Three example motion models that may be used for global motion are illustrated, including index values ​​(0, 1, or 2) for the three example motion models.

[0055] Still refer to Figure 2 , PPS can be used to signal parameters that can change between pictures in a sequence. Parameters that remain the same for a sequence of pictures can be signaled in a sequence parameter set to reduce the size of the PPS and reduce the video bitrate. An example picture parameter set (PPS) is shown in Table 2:

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] Additional fields can be added to the PPS to signal global motion. In the case of global motion, the presence of global motion parameters in the picture sequence can be signaled in the SPS and the PPS can reference the SPS via the SPS ID. The SPS in some decoding methods can be modified to add fields to signal the presence of global motion parameters in the SPS. For example, a one-bit field can be added to the SPS. If the global_motion_present bit is 1, global motion related parameters can be expected in the PPS. If the global_motion_present bit is 0, global motion parameter related fields cannot exist in the PPS. For example, the PPS of Table 2 can be extended to include the global_motion_present field, for example, as shown in Table 3:

[0062]

[0063]

[0064] Similarly, the PPS may include a pps_global_motion_parameters field for a frame, for example as shown in Table 4:

[0065]

[0066] In more detail, the PPS may include a field for representing global motion parameters using control point motion vectors, for example, as shown in Table 5:

[0067]

[0068]

[0069] As a further non-limiting example, Table 6 below may represent an exemplary SPS:

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079] The SPS table above can be expanded as described above to incorporate a global motion presence indicator as shown in Table 7:

[0080]

[0081]

[0082] Additional fields may be incorporated into the SPS to reflect further indicators as described in this disclosure.

[0083] In the embodiment, still refer to Figure 2 The sps_affine_enabled_flag in the PPS and / or SPS can specify whether affine-based motion compensation can be used for inter-frame prediction. If sps_affine_enabled_flag is equal to 0, the syntax can be constrained so that affine-based motion compensation is not used in the later coded video sequence (CLVS), and the inter_affine_flag and cu_affine_type_flag will not be present in the coding unit syntax of the CLVS. Otherwise (sps_affine_enabled_flag is equal to 1), affine-based motion compensation can be used in the CLVS.

[0084] Continue to refer to Figure 2 The sps_affine_type_flag in the PPS and / or SPS can specify whether motion compensation based on the six-parameter affine model can be used for inter prediction. If sps_affine_type_flag is equal to 0, the syntax can be constrained so that motion compensation based on the six-parameter affine model is not used in CLVS, and the cu_affine_type_flag will not be present in the coding unit syntax in CLVS. Otherwise (sps_affine_type_flag is equal to 1), motion compensation based on the six-parameter affine model can be used in CLVS. When not present, the value of sps_affine_type_flag can be inferred to be equal to 0.

[0085] Still refer to Figure 2 , the translation CPMV data can be signaled in the PPS. The control points can be predefined. For example, control point MV 0 can be relative to the upper left corner of the picture, MV 1 can be relative to the upper right corner, and MV 3 can be relative to the lower left corner of the picture. Table 5 illustrates an example method for signaling CPMV data based on the motion model used.

[0086] In an exemplary embodiment, still referring to Figure 2 , the array amvr_precision_idx signaled in a coding unit, coding tree, etc. can specify the resolution AmvrShift of the motion vector difference, and the resolution AmvrShift can be defined as a non-limiting example shown in Table 8 below. The array indices x0, y0 can specify the position (x0, y0) of the upper-left luminance sample of the coding block under consideration relative to the upper-left luminance sample of the picture; when amvr_precision_idx[x0][y0] does not exist, it can be inferred to be equal to 0. When inter_affine_flag[x0][y0] is equal to 0, the variables MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdL1[x0][y0][0], MvdL1[x0][y0][1] corresponding to the modulated vector difference of the block under consideration can be modified by shifting the following values by AmvrShift, for example using

[0087] MvdL0[x0][y0][0] = MvdL0[x0][y0][0] << AmvrShift;

[0088] MvdL0[x0][y0][1] = MvdL0[x0][y0][1] << AmvrShift;

[0089] MvdL1[x0][y0][0] = MvdL1[x0][y0][0] << AmvrShift; and

[0090] MvdL1[x0][y0][1] = MvdL1[x0][y0][1] << AmvrShift.

[0091] Where inter_affine_flag[x0][y0] is equal to 1, the variables MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0] and MvdCpL0[x0][y0][2][1] can be modified by shifting, for example as follows:

[0092] MvdCpL0[x0][y0][0][0] = MvdCpL0[x0][y0][0][0] << AmvrShift;

[0093] MvdCpL1[x0][y0][0][1] = MvdCpL1[x0][y0][0][1] << AmvrShift;

[0094] MvdCpL0[x0][y0][1][0] = MvdCpL0[x0][y0][1][0] << AmvrShift;

[0095] MvdCpL1[x0][y0][1][1] = MvdCpL1[x0][y0][1][1] << AmvrShift;

[0096] MvdCpL0[x0][y0][2][0] = MvdCpL0[x0][y0][2][0] << AmvrShift; and

[0097] MvdCpL1[x0][y0][2][1] = MvdCpL1[x0][y0][2][1] << AmvrShift.

[0098]

[0099] Figure 3 is a process flow diagram illustrating an exemplary embodiment of a process 300 for constructing a motion vector candidate list, where constructing a merge candidate list includes adding a global motion vector to the motion vector candidate list.

[0100] In step 305, a bitstream is received by a decoder. The current block may be included within the bitstream received by the decoder. The bitstream may include, for example, data found in a bitstream that is an input to a decoder when using data compression. The bitstream may include information required to decode a video. Receiving may include extracting and / or parsing the block and associated signaling information from the bitstream. In some implementations, the current block may include a coding tree unit (CTU), a coding unit (CU), and / or a prediction unit (PU). In step 310, the bitstream is used to determine that the merge mode is enabled for the current block. In step 315, a merge candidate list is constructed; constructing includes adding a global motion vector to the motion vector candidate list. In step 340, the motion vector candidate list may be used to reconstruct the pixel data of the current block.

[0101] Figure 4 is a system block diagram illustrating an example decoder 400 capable of decoding a bitstream, including constructing a merge candidate list (including adding a global motion vector to the motion vector candidate list). The decoder 400 may include an entropy decoder processor 404, an inverse quantization and inverse transform processor 408, a deblocking filter 412, a frame buffer 416, a motion compensation processor 420, and / or an intra prediction processor 424.

[0102] In operation, still refer to Figure 4 , the bitstream 428 can be received by the decoder 400 and input to the entropy decoder processor 404, which can entropy decode portions of the bitstream into quantized coefficients. The quantized coefficients can be provided to the inverse quantization and inverse transform processor 408, which can perform inverse quantization and inverse transform to create a residual signal, which can be added to the output of the motion compensation processor 420 or the intra-frame prediction processor 424 depending on the processing mode. The output of the motion compensation processor 420 and the intra-frame prediction processor 424 may include block predictions based on previously decoded blocks. The sum of the prediction and the residual can be processed by the deblocking filter 412 and stored in the frame buffer 416.

[0103] Figure 5 is a process flow diagram illustrating an example process 500 for encoding video that can reduce encoding complexity while increasing compression efficiency according to some aspects of the present subject matter, including adding a global motion vector to a motion vector candidate list to construct a merge candidate list. In step 505, a video frame can undergo initial block partitioning, for example using a tree-structured macroblock partitioning scheme that can include partitioning a picture frame into CTUs and CUs. In step 510, global motion can be determined. In step 515, the block can be encoded and included in the bitstream. Encoding can include constructing a merge candidate list, including adding the global motion vector to the motion vector candidate list. Encoding can include utilizing, for example, inter-frame prediction mode and intra-frame prediction mode.

[0104] Figure 6 is a system block diagram illustrating an example video encoder 600 capable of constructing a merge candidate list, including adding a global motion vector to the motion vector candidate list. The example video encoder 600 can receive an input video 604 that can be initially partitioned or divided according to a processing scheme such as a tree-structured macroblock partitioning scheme (e.g., a quadtree plus a binary tree). An example of a tree-structured macroblock partitioning scheme can include partitioning a picture frame into large block elements called coding tree units (CTUs). In some implementations, each CTU can be further partitioned into multiple sub-blocks called coding units (CUs). The final result of this partitioning can include a set of sub-blocks that can be called prediction units (PUs). Transform units (TUs) can also be used.

[0105] Still refer to Figure 6The example video encoder 600 may include an intra prediction processor 608, a motion estimation / compensation processor 612, a transform / quantization processor 616, an inverse quantization / inverse transform processor 620, a loop filter 624, a decoded picture buffer 628, and / or an entropy coding processor 632. The motion estimation / compensation processor 612 may also be referred to as an inter prediction processor and may be capable of constructing a motion vector candidate list, including adding a global motion vector candidate to the motion vector candidate list. The bitstream parameters may be input to the entropy coding processor 632 for inclusion in the output bitstream 636.

[0106] Continue to refer to Figure 6 In operation, for each block of a frame of input video 604, a determination may be made as to whether the block is to be processed via intra picture prediction or using motion estimation / compensation. The block may be provided to either an intra prediction processor 608 or a motion estimation / compensation processor 612. If the block is to be processed via intra prediction, the intra prediction processor 608 may perform processing to output a prediction value. If applicable, if the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 612 may perform processing including constructing a motion vector candidate list (including adding a global motion vector candidate to the motion vector candidate list).

[0107] Further references Figure 6 , a residual can be formed by subtracting the prediction value from the input video. The residual can be received by the transform / quantization processor 616, which can perform a transform process (e.g., a discrete cosine transform (DCT)) to produce coefficients that can be quantized. The quantized coefficients and any associated signaling information can be provided to the entropy coding processor 632 for entropy coding and included in the output bitstream 636. The entropy coding processor 632 can support the encoding of signaling information related to encoding the current block. In addition, the quantized coefficients can be provided to the inverse quantization / inverse transform processor 620, which can reproduce the pixels, which can be combined with the prediction value and processed by the loop filter 624. The output of the loop filter 624 can be stored in the decoded picture buffer 628 for use by the motion estimation / compensation processor 612, which is used to construct a motion vector candidate list, including adding the global motion vector candidate to the motion vector candidate list.

[0108] Continue to refer to Figure 6 Although some variations have been described in detail above, other modifications or additions are possible. For example, in some implementations, the current block can include any symmetric block (8x8, 16x16, 32x32, 64x64, 128x128, etc.) and any asymmetric block (8x4, 16x8, etc.).

[0109] Still refer to Figure 6 In some implementations, a quadtree plus binary decision tree (QTBT) can be implemented. In QTBT, at the coding tree unit level, the partitioning parameters of the QTBT can be dynamically derived to adapt to local characteristics without any transmission overhead. Subsequently, at the coding unit level, the joint classifier decision tree structure can eliminate unnecessary iterations and control the risk of wrong predictions. In some implementations, the LTR frame block update mode can be used as an additional option available at each leaf node of the QTBT.

[0110] Still refer to Figure 6 In some implementations, additional syntax elements can be signaled at different hierarchical levels of the bitstream. For example, a flag can be enabled for the entire sequence by including an enable flag encoded in a sequence parameter set (SPS). Additionally, a coding tree unit (CTU) flag can be encoded at the CTU level.

[0111] It should be noted that any one or more aspects and embodiments described herein can be conveniently implemented using digital electronic circuits, integrated circuits, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof, such as implemented and / or implemented in one or more machines (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices, such as document servers, etc.) programmed according to the teachings of this specification, as will be apparent to one of ordinary skill in the computer arts. These various aspects or features can include implementations in one or more computer programs and / or software that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be either special-purpose or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to a storage system, at least one input device, and at least one output device. Appropriate software coding can be readily prepared by a skilled programmer based on the teachings of this disclosure, as will be apparent to one of ordinary skill in the software arts. The aspects and implementations discussed above that employ software and / or software modules may also include appropriate hardware to assist in implementing the machine-executable instructions of the software and / or software modules.

[0112] Such software can be a computer program product using a machine-readable storage medium. A machine-readable storage medium can be any medium that can store and / or encode an instruction sequence executed by a machine (e.g., a computing device) and enable the machine to execute any one of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state storage devices, EPROMs, EEPROMs, programmable logic devices (PLDs), and / or any combination thereof. As used herein, a machine-readable medium is intended to include a single medium and a physically separated set of media, e.g., one or more hard disk drives and optical disks combined with a computer memory. As used herein, a machine-readable storage medium does not include a temporary form of signal transmission.

[0113] Such software may also include information (e.g., data) carried as a data signal on a data carrier such as a carrier wave. For example, it may include machine-executable information as a data-carrying signal embodied in a data carrier, where the signal encodes a sequence of instructions or a portion thereof for execution by a machine (e.g., a computing device), as well as any related information (e.g., data structures and data) that causes the machine to perform any of the methods and / or embodiments described herein.

[0114] Examples of computing devices include, but are not limited to, electronic book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smartphones, etc.), network devices, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions specifying an action to be taken by the machine, and any combination thereof. In one example, the computing device may include and / or be included in an information kiosk.

[0115] Figure 7A diagrammatic representation of one embodiment of a computing device in a computer system 700 in exemplary form is shown within which a set of instructions for causing a control system to perform any one or more aspects and / or methods of the present disclosure may be executed. It is also contemplated that multiple computing devices may be used to implement a set of specifically configured instructions for causing one or more devices to perform any one or more aspects and / or methods of the present disclosure. The computer system 700 includes a processor 704 and a memory 708 that communicate with each other and with other components via a bus 712. The bus 712 may include any of several types of bus structures, including but not limited to a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures.

[0116] The memory 708 may include various components (e.g., machine-readable media), including, but not limited to, random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 716 (BIOS) including basic routines that facilitate the transfer of information between components within the computer system 700, such as during startup, may be stored in the memory 708. The memory 708 may also include instructions (e.g., software) 720 that implement any one or more aspects and / or methods of the present disclosure (e.g., stored on one or more machine-readable media). In another example, the memory 708 may also include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.

[0117] The computer system 700 may also include a storage device 724. Examples of storage devices (e.g., storage device 724) include, but are not limited to, a hard drive, a magnetic disk drive, a combination of an optical disk drive and optical media, a solid-state storage device, and any combination thereof. The storage device 724 may be connected to the bus 712 via an appropriate interface (not shown). Example interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE), and any combination thereof. In one example, the storage device 724 (or one or more of its components) may be removably interactive with the computer system 700 (e.g., via an external port connector (not shown)). In particular, the storage device 724 and associated machine-readable media 728 may provide non-volatile and / or volatile storage for machine-readable instructions, data structures, program modules, and / or other data for the computer system 700. In one example, the software 720 may reside entirely or partially within the machine-readable media 728. In another example, the software 720 may reside entirely or partially within the processor 704.

[0118] The computer system 700 may also include an input device 732. In one example, a user of the computer system 700 can enter commands and / or other information into the computer system 700 via the input device 732. Examples of the input device 732 include, but are not limited to, an alphanumeric input device (e.g., a keyboard), a pointing device, a joystick, a game controller, an audio input device (e.g., a microphone, a voice response system, etc.), a cursor control device (e.g., a mouse), a touchpad, an optical scanner, a video capture device (e.g., a still camera, a video camera), a touch screen, and any combination thereof. The input device 732 can be connected to the bus 712 via any of a variety of interfaces (not shown), including, but not limited to, a serial interface, a parallel interface, a game port, a USB interface, a FIREWIRE interface, a direct interface to the bus 712, and any combination thereof. The input device 732 may include a touch screen interface, which may be part of or separate from the display 736, as will be discussed further below. The input device 732 may be used as a user selection device for selecting one or more graphical representations in the graphical interface described above.

[0119] The user can also input commands and / or other information to the computer system 700 via storage devices 724 (e.g., removable disk drives, flash drives, etc.) and / or network interface devices 740. Network interface devices such as network interface device 740 can be used to connect the computer system 700 to one or more of a variety of networks, such as network 744, and one or more remote devices 748 connected to these networks. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, enterprise networks), local area networks (e.g., networks associated with an office, building, campus, or other relatively small geographic space), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof. Networks, such as network 744, can employ wired and / or wireless communication modes. In general, any network topology can be used. Information (eg, data, software 720 , etc.) can be transferred to and / or from computer system 700 via network interface device 740 .

[0120] The computer system 700 may also include a video display adapter 752 for transmitting displayable images to a display device such as a display device 736. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. Display adapter 752 and display device 736 may be used in conjunction with processor 704 to provide a graphical representation of aspects of the present disclosure. In addition to the display device, the computer system 700 may include one or more other peripheral output devices, including but not limited to audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to bus 712 via a peripheral interface 756. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, FIREWIRE connections, parallel connections, and any combination thereof.

[0121] The above is a detailed description of an illustrative embodiment of the present invention. Various modifications and additions may be made without departing from the spirit and scope of the present invention. The features of each embodiment in the various embodiments described above may be appropriately combined with the features of the other described embodiments to provide a variety of feature combinations in associated new embodiments. In addition, although a plurality of separate embodiments have been described above, what is described herein is merely an illustration of the application of the principles of the present invention. In addition, although the specific methods herein may be illustrated and / or described as being performed in a particular order, within the ordinary skill of implementing the embodiments disclosed herein, the order is highly variable. Therefore, this description is intended to be illustrative only and not to limit the scope of the present invention in any other way.

[0122] In the above description and claims, phrases such as "at least one of" or "one or more of" may be followed by a list of elements or features. The term "and / or" may also appear in a list of two or more elements or features. Unless explicitly or implicitly contradicted by the context in which it is used, a phrase refers to any individual element or feature of the listed elements or features, or to any one of the listed elements or features combined with any one of the other listed elements or features. For example, each of the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" means "only A, only B, or both A and B." A similar interpretation applies to lists containing three or more items. For example, each of the phrases "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, and / or C" means "only A, only B, only C, A and B together, A and C together, B and C together, or A and B and C together." Furthermore, the term "based on" used above and in the claims means "based at least in part on," such that features or elements not mentioned are also allowable.

[0123] The subject matter described herein may be embodied in systems, devices, methods and / or articles depending on the desired configuration. The implementations described in the foregoing description do not represent all implementations consistent with the subject matter described herein. On the contrary, these implementations are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those set forth herein. For example, the implementations described above may involve various combinations and subcombinations of the disclosed features, and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows described in the accompanying drawings and / or herein do not necessarily require the specific order shown, or sequence, to achieve the desired results. Other implementations may fall within the scope of the appended claims.

Claims

1. A decoder comprising a circuit configured to: Receiving a bitstream comprising a coded picture, the coded picture comprising a first contiguous region and a second contiguous region, the first contiguous region comprising a first plurality of coded blocks, the second contiguous region comprising a second plurality of coded blocks, the first contiguous region containing only global motion, and the second contiguous region containing local motion; Decoding a first continuous region of the coded picture to reconstruct the global motion by: For each coding block in a first plurality of coding blocks in a first continuous region, using a motion model, the motion model being a global motion model for all first coding blocks in the first continuous region, the global motion model being one of a translational motion, a 4-parameter affine motion, or a 6-parameter affine motion; If the global motion model is translational motion, constructing a merge candidate list for each of a first plurality of coding blocks in the first continuous region, the merge candidate list including a first candidate, the first candidate being a motion vector of a neighboring block in the picture, and decoding each of the first plurality of coding blocks in the first continuous region using the merge candidate list by selecting the first candidate for translational motion compensation; If the global motion model is 4-parameter affine motion, constructing a merge candidate list for each of a first plurality of coding blocks in the first continuous region, the merge candidate list including a second candidate, the second candidate including two control point motion vectors, each control point motion vector being a motion vector of a neighboring block in the picture, and decoding each of a first plurality of decoded blocks of the first continuous region using the merge candidate by selecting the second candidate for 4-parameter affine motion compensation; If the global motion model is 6-parameter affine motion, constructing a merge candidate list for each of the first plurality of coding blocks in the first continuous region, the merge candidate list including a third candidate, the third candidate including three control point motion vectors, each control point motion vector being a motion vector of a neighboring block in the picture, and decoding each of the first plurality of coding blocks in the first continuous region using the merge candidate list by selecting the third candidate for 6-parameter affine motion compensation; as well as A second consecutive region of the coded picture is decoded to reconstruct the local motion.

2. The decoder according to claim 1, wherein The bitstream signals a motion vector difference for use with one of the first candidate, the second candidate, or the third candidate to decode each of the first plurality of coding blocks of the first continuous region.

3. The decoder according to claim 1, wherein The first plurality of coding blocks in the first continuous area are all 64x64 or all 128x128.

4. The decoder according to claim 1, wherein The global motion in the first continuous region is caused by camera motion.

5. The decoder according to claim 1, wherein The local motion in the second continuous region is caused by the motion of objects in the scene.

6. A video encoder comprising a circuit configured to: Encoding a bitstream decoded by a compatible decoder, the encoded bitstream comprising a coded picture, the coded picture comprising a first contiguous region and a second contiguous region, the first contiguous region comprising a first plurality of coded blocks, the second contiguous region comprising a second plurality of coded blocks, the first contiguous region containing only global motion, and the second contiguous region containing local motion; A decoder receiving the encoded bit stream is configured to: receiving the encoded bit stream; Decoding a first continuous region of the coded picture to reconstruct the global motion by: For each coding block in a first plurality of coding blocks in a first continuous region, using a motion model, the motion model being a global motion model for all first coding blocks in the first continuous region, the global motion model being one of a translational motion, a 4-parameter affine motion, or a 6-parameter affine motion; If the global motion model is translational motion, constructing a merge candidate list for each of a first plurality of coding blocks in the first continuous region, the merge candidate list including a first candidate, the first candidate being a motion vector of a neighboring block in the picture, and decoding each of the first plurality of coding blocks in the first continuous region using the merge candidate list by selecting the first candidate for translational motion compensation; If the global motion model is 4-parameter affine motion, constructing a merge candidate list for each of a first plurality of coding blocks in the first continuous region, the merge candidate list including a second candidate, the second candidate including two control point motion vectors, each control point motion vector being a motion vector of a neighboring block in the picture, and decoding each of a first plurality of decoded blocks of the first continuous region using the merge candidate by selecting the second candidate for 4-parameter affine motion compensation; If the global motion model is 6-parameter affine motion, constructing a merge candidate list for each of the first plurality of coding blocks in the first continuous region, the merge candidate list including a third candidate, the third candidate including three control point motion vectors, each control point motion vector being a motion vector of a neighboring block in the picture, and decoding each of the first plurality of coding blocks in the first continuous region using the merge candidate list by selecting the third candidate for 6-parameter affine motion compensation; as well as A second consecutive region of the coded picture is decoded to reconstruct the local motion.

7. The encoder according to claim 6, wherein The bitstream signals a motion vector difference for use with one of the first candidate, the second candidate, or the third candidate to decode each of the first plurality of coding blocks of the first continuous region.

8. The encoder according to claim 7, wherein The first plurality of coding blocks in the first continuous area are all 64x64 or all 128x128.

9. The encoder according to claim 6, wherein The global motion in the first continuous region is caused by camera motion.

10. The encoder according to claim 8, wherein The local motion in the second continuous region is caused by the motion of objects in the scene.

Citation Information

Patent Citations

  • Method and apparatus for global motion compensation in video coding system

    CN108293128A