Merge candidate reordering based on global motion vector

By constructing and reordering a merged candidate list based on global motion vectors, the problems of high coding complexity and high bit rate in existing video coding technologies are solved, improving video compression efficiency, especially the coding efficiency when processing video frames with mixed global and local motion.

CN121985142APending Publication Date: 2026-05-05DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2020-06-03
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing video coding techniques suffer from high coding complexity, high bit rate, and low coding efficiency during compression. In particular, when processing video frames with global and local motion, existing methods fail to effectively utilize global motion vectors for candidate merging and sorting optimization.

Method used

By constructing a merged candidate list based on global motion vectors and prioritizing candidates with representations of global motion vectors, the motion vector candidate list is reordered, reducing the number of bits sent for signals with motion vector differences and improving coding efficiency.

Benefits of technology

By optimizing the sorting of the candidate list, the encoding complexity and bit rate are reduced, and the efficiency of video compression is improved, especially when processing video frames with mixed global and local motion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985142A_ABST
    Figure CN121985142A_ABST
Patent Text Reader

Abstract

A decoder, the decoder comprising a circuit configured to receive a bit stream; constructing, for the current block, a motion vector candidate list including motion vector candidates having motion information describing a global motion vector; reordering the motion vector candidate list, so that the motion vector candidates with the motion information representing the global motion vector are arranged at the first position in the reordered motion vector candidate list; and reconstructing pixel data of the current block and using the reordered motion vector candidate list. Related equipment, systems, techniques, and articles are also presented herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application filed on June 3, 2020, with application number 202080053096.1, applicant OP Solutions LLC, and invention title "Merging Candidate Reordering Based on Global Motion Vector". Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 856,339, filed June 3, 2019, entitled “Merged Candidate Reordering Based on Global Motion Vectors,” which is incorporated herein by reference in its entirety. Technical Field

[0003] This invention generally relates to the field of video compression. In particular, this invention relates to merging candidate reordering based on global motion vectors. Background Technology

[0004] A video codec can include electronic circuitry or software that compresses or decompresses digital video. It can convert uncompressed video into a compressed format and vice versa. In the case of video compression, the device that compresses the video (and / or performs some of the functions of that device) is generally called an encoder, while the device that decompresses the video (and / or performs some of the functions of that device) is called a decoder.

[0005] The format of compressed data can conform to standard video compression specifications. Compression may be lossy because compressed video lacks some information present in the original video. Such results may include decompressed video having lower quality than the original uncompressed video because there is not enough information to accurately reconstruct the original video.

[0006] There can be complex relationships between video quality, the amount of data used to represent the video (e.g., determined by bit rate), the complexity of encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, end-to-end latency (e.g., waiting time), and so on.

[0007] Motion compensation can include a method that predicts a video frame or a portion of a given reference frame (e.g., a previous and / or future frame) by taking into account the motion of objects in the camera and / or video. It can be used in the encoding and decoding of video data for video compression, such as encoding and decoding using the Moving Picture Experts Group (MPEG)-2 (also known as High-Level Video Coding (AVC) and H.264) standards. Motion compensation can describe an image based on the transformation from a reference image to the current image. The reference image may be temporally previous when compared to the current image, and may be from the future when compared to the current image. Compression efficiency can be improved when images can be accurately synthesized from previously transmitted and / or stored images. Summary of the Invention

[0008] In one aspect, a decoder includes circuitry configured to receive a bitstream; construct a motion vector candidate list for a current block, the motion vector candidate list containing motion vector candidates having motion information characterizing a global motion vector; reorder the motion vector candidate list such that a motion vector candidate is ranked first in the reordered motion vector candidate list, the motion vector candidate having motion information characterizing the global motion vector; and reconstruct pixel data of the current block using the reordered motion vector candidate list.

[0009] In another approach, a method includes: receiving a bitstream by a decoder; constructing a motion vector candidate list for the current block, the motion vector candidate list containing motion information characterizing a global motion vector; reordering the motion vector candidate list such that a motion vector candidate is ranked first in the reordered motion vector candidate list, the motion vector candidate having motion information characterizing the global motion vector; and reconstructing the pixel data of the current block using the reordered motion vector candidate list.

[0010] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will become apparent from the specification, the drawings, and the claims. Attached Figure Description

[0011] To illustrate the invention, the accompanying drawings show various aspects of one or more embodiments of the invention. However, it should be understood that the invention is not limited to the precise arrangements and mechanisms shown in the drawings, wherein: Figure 5 It is a flowchart of the process implemented based on some examples of the current topic; Figure 6 This is a system block diagram of an example decoder implemented based on some examples of the current topic; Figure 7It is a flowchart of the process implemented based on some examples of the current topic; Figure 8 This is a system block diagram of an example encoder implemented based on some examples of the current topic; Figure 1 It is a diagram illustrating the motion vectors of an instance frame with global and local motion; Figure 4 This describes three instance motion models that can be used for global motion, including the index values ​​(0, 1, or 2) of these three instance motion models. Figure 2 This is a block diagram illustrating the spatial candidates considered in the merging pattern approach; Figure 3 This is a block diagram illustrating the global motion vector, which considers spatial candidates and correlations in the merging mode approach; and Figure 9 This is a block diagram of a computing system that can be used to implement any one or more of the methods disclosed herein, and any one or more of its components.

[0012] The drawings are not necessarily drawn to scale and may be illustrated by dashed lines, diagrams, and partial views. In some cases, details that are not essential for understanding the embodiments or that would make other details difficult to understand have been omitted. The same reference numerals denote the same elements in the various figures. Detailed Implementation

[0013] Global motion in video refers to motion occurring throughout the entire frame. Global motion can be caused by camera movement; for example, camera panning and zooming produce motion within a frame that typically affects the entire frame. Motion present in parts of the video can be called local motion. Local motion can be caused by moving objects in the scene, such as objects moving from left to right. Video can contain a combination of local and global motion. Some implementations of the current topic provide a way to construct a merged candidate list based on global motion vectors, which can improve compression by reducing the number of bits required to signal candidates and encode differences in motion vectors.

[0014] Figure 1 This is a diagram illustrating the motion vectors of an example frame 100 with global and local motion. Frame 100 may include multiple pixel blocks, shown as squares, and motion vectors associated with them, shown as arrows. Squares (e.g., pixel blocks) with arrows pointing upwards and to the left indicate blocks with what can be considered global motion, while squares with arrows pointing in other directions (indicated by 104) indicate blocks with local motion. Figure 1In the example shown, many blocks share the same global motion. Sending a global motion signal in the header, such as as an image parameter set (PPS) or a sequence parameter set (SPS), and using the signaled global motion reduces the amount of motion vector information required for each block and can lead to improved predictions.

[0015] As an example, and continue to refer to Figure 1 It is possible to use MV with two components. x MV y The motion vector (MV) is used to describe simple translational motion. x MV y This describes the displacement of blocks and / or pixels in the current frame. Affine motion vectors can be used to describe more complex motions such as rotation, scaling, and warping, where, as used in this disclosure, an "affine motion vector" is a vector describing the uniform displacement of a set of pixels or points represented in a video image and / or picture, for example, describing a set of pixels that indicate the movement of an object in a view within the video without changing its appearance shape during the motion. Some methods of video encoding and / or decoding can use four-parameter or six-parameter affine models for motion compensation in inter-image coding.

[0016] For example, a six-parameter affine motion model can be described as: x' = ax + by + c y' = dx + ey + f Four-parameter affine motion can be described as: x' = ax + by + c y' = -bx + ay + f Where (x, y) and (x', y') are the pixel positions in the current and reference images, respectively; a, b, c, d, e, and f are the parameters of the affine motion model.

[0017] Continue to refer to Figure 1 This can be used alternatively or additionally to apply block-based and / or sub-block-based affine transformation motion compensation prediction. The affine motion region of a block and / or sub-block can be described by motion information from two control points (four parameters) or three control point motion vectors (six parameters). In the four-parameter affine motion model, the motion vector at the sample position (x, y) in the block can be derived as: ; For a six-parameter affine motion model, the motion vector at the sample position (x, y) in the block can be derived as: ; Among them (mv) 0x ,mv 0y(mv) is the motion vector of the top left control point. 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point.

[0018] To simplify motion-compensated prediction, block-based affine transformation prediction can be applied. As an illustrative example, to derive the motion vector for each 4×4 luma sub-block, the motion vector of the center sample of each sub-block can be calculated according to the above equation and rounded to 1 / 16 fractional accuracy. A motion-compensated interpolation filter can then be applied to generate the prediction for each sub-block using the derived motion vector. Continuing with this example, the sub-block size for the chroma component can also be set to 4×4. The motion vector of the 4×4 chroma sub-block can be calculated as the average of the MV of the four corresponding 4×4 luma sub-blocks.

[0019] Similar to inter-frame prediction for translational motion, there are two other inter-frame prediction modes for affine motion: affine merging mode and affine AMVP mode. (See also...) Figure 1 The parameters describing the affine motion can be sent to the decoder for application of affine motion compensation. In some methods, motion parameters can be sent explicitly or by signaling translational control point motion vectors (CPMVs) and deriving the affine motion parameters from these vectors. Two control point motion vectors (CPMVs) can be used to derive the affine motion parameters for a four-parameter affine motion model, and three control point translational motion vectors (CPMVs) can be used to obtain the parameters for a six-parameter motion model. Signaling affine motion parameters using control point motion vectors allows for the use of efficient motion vector encoding methods.

[0020] Continue to refer to Figure 1 Some blocks can share the same motion vector information. For example, two blocks corresponding to an object moving on the screen can share the same motion vector because they are both associated with the same object. In such cases, some motion compensation methods can utilize merging modes, where adjacent blocks can share motion vectors that allow the first block's motion information to be encoded in the bitstream, and a second block can inherit motion information from the first block (e.g., merge with the first block). During encoding, a merge list containing available merge candidates can be constructed. Merge candidates can be selected from the constructed merge list, and the index to the merge list can be signaled in the bitstream. During decoding, a merge list can be constructed again from the available merge candidates, and the signaled index in the bitstream can be used to indicate from which block the current block will inherit motion information (e.g., merge with which block).

[0021] Figure 2A block diagram 200 illustrates an exemplary embodiment of a spatial candidate considered in a typical approach such as a merging mode for an HEVC implementation. The current block 204 may include coding units or prediction units. Spatial merging candidates may include A0, A1, B0, B1, and B2. A0, A1, B0, and B2 may include adjacent prediction and / or coding units. When creating a merging candidate list, the list can be constructed by considering up to four spatial merging candidates derived from five spatially adjacent blocks, such as... Figure 2 As shown in the example. In this instance, thresholds for five spatial candidates can be applied. Besides considering... Figure 2 In addition to the spatial candidates described, additional candidates that can be considered for addition to the merging list may include a temporal merging candidate, which can be derived from two temporally colocalized blocks, combined bidirectional prediction candidates, and zero motion vector candidates.

[0022] Still referencing Figure 2 Spatial merge candidates can be added to the merge list in response to determining their availability. In a quadtree plus binary decision tree (QTBT) partition, some of the adjacent blocks may be asymmetric blocks, and therefore such adjacent blocks may not be considered spatial merge candidates because they may be asymmetric partitions, since partitions (e.g., prediction units) do not share similar motion information.

[0023] As described above, and continue to refer to Figure 2 In some video coding methods, a list of merging candidates can be constructed based on the following candidates: up to four spatial merging candidates derived from five spatially adjacent blocks; a temporal merging candidate derived from two temporally colocalized blocks; additional merging candidates including combined bidirectional prediction candidates; and zero motion vector candidates.

[0024] Still referencing Figure 2 In order to derive a list of spatial candidates, (a) it is possible to check whether adjacent blocks are available and whether they contain motion information, and (b) to perform redundancy checks to avoid having candidates with redundant motion data in the list.

[0025] Continue to refer to Figure 2 When N is the number of spatial merging candidates, a complete redundancy check can consist of N x (N-1) / 2 motion data comparisons. With five potential merging candidates, ten motion data comparisons could be used, ensuring that all candidates in the merging list have different motion data. This could increase the complexity of the decoder.

[0026] In some video encoding methods, and still referencing Figure 2To improve coding efficiency, after constructing the merge candidate list (processing spatial candidate positions in the order A1, B1, B0, A0, B2), the order of each merge candidate is adjusted based on the template matching cost. The template matching cost can be measured as the sum of the absolute differences (SADs) between the current coding unit's (CU) neighboring samples and their corresponding reference samples. For example, and not limited to, the merge candidates can be ordered in ascending order of the SADs calculated using the merge candidates. The number of merge candidates selected using the template matching cost can be finite. For example, a set of four lowest-cost candidates from five originally generated and / or provided candidates can be selected.

[0027] Still referencing Figure 2 Some implementations of the current topic can further improve coding efficiency by using global motion vectors to reorder and merge candidates. As used in this invention, global motion in a video refers to motion that occurs throughout the entire frame. Global motion is often caused by camera motion, which affects things like camera panning and zooming throughout the frame.

[0028] Still referencing Figure 2 Some implementations of the current topic can create a list of merging candidates based on motion vectors signaled to the decoder. If global motion is represented by signals, then this global motion can be expected to be shared by many blocks in the frame. For example, as... Figure 3 As illustrated for illustrative purposes, three of the five spatial merging candidates (B1, B2, and A1) can be transmitted based on global motion signals. Based on the signal transmission, at the decoder, the decoder can create the following list of merging candidates. They are ordered such that the global motion candidate is listed first, as shown in Table 1.

[0029] Updated (reordered) merge candidate list B1 - GMV1 B2 - GMV2 A1 - GMV3 B0 A0 See also Figure 2 Since blocks may have motions similar to the global motion, modifying the list so that the global motion vector is the first candidate in the list reduces the number of bits required to signal the predicted candidates and encode the differences in motion vectors. In this way, motion vector encoding can be improved and the bit rate reduced, which will increase compression efficiency.

[0030] In some implementations, and continue to refer to Figure 2 Global motion signal transmission can be included in headers such as PPS or SPS. Global motion can vary with the image. Motion vectors represented by signals in the image header describe motion relative to previously decoded frames. In some implementations, global motion can be translational or affine. The motion mode used, such as multiple parameters, whether the model is affine, translational, or similar, can also be transmitted by signals in the image header. Figure 4Three exemplary embodiments of motion model 600 are shown, which can be used for global motion and includes their index values ​​(0, 1 or 2).

[0031] Still referencing Figure 4 In PPS, translation CPMV can be sent using signals. Control points can be predefined. For example, control point MV0 can be associated with the top left corner of the image, MV1 with the top right corner, and MV3 with the bottom left corner.

[0032] Continue to refer to Figure 4 Global motion can be correlated with previously encoded frames. When only one set of global motion parameters exists, motion can be relative to the frame immediately preceding the current frame.

[0033] Figure 5 This is a process flowchart illustrating an exemplary embodiment of a process 500 for merging candidate reordering based on global motion vectors.

[0034] In step 505, and still referring to Figure 5 The decoder receives a bitstream including the current block. The current block may be contained within the bitstream received by the decoder. The bitstream may include data found, for example, in the bitstream that serves as input to the decoder when data compression is used. The bitstream may contain information necessary for decoding the video. Receiving may include extracting and / or parsing blocks and associated signal transmission information from the bitstream. In some implementations, the current block may include a coding tree unit (CTU), a coding unit (CU), or a prediction unit (PU).

[0035] At step 510, and further see Figure 5 This allows the construction of a candidate list of motion vectors for the current block, containing motion information representing global motion vectors. Global motion vectors can be represented by a header of the bitstream, which contains the image parameter set (PPS) and / or the sequence parameter set (SPS).

[0036] At step 515, and continue to see Figure 5 The motion vector candidate list is reordered so that the motion vector candidate with motion information representing the global motion vector is placed first in the reordered list. Reordering may involve inserting the first global motion vector candidate into the merged candidate list. In some implementations, the construction of the motion vector candidate list may include reordering.

[0037] In step 520, and further refer to Figure 5 The pixel data of the current block can be reconstructed using a reordered list of motion vector candidates.

[0038] Still referencing Figure 5In some implementations, the decoder may be configured to determine the current frame containing the current block as an indication of global motion. Global motion vectors may include control point motion vectors. Control point motion vectors may include translational motion vectors. Control point motion vectors may contain vectors from a four-parameter affine motion model or a six-parameter affine motion model.

[0039] Figure 6 This is a system block diagram illustrating an example decoder 600 capable of decoding a bitstream based on a merged candidate reordering of global motion vectors. Decoder 600 may include an entropy decoder processor 604, an inverse quantization and inverse transform processor 608, a deblocking filter 612, a frame buffer 616, a motion compensation processor 620, and / or an intra-frame prediction processor 624.

[0040] During operation, and still referencing Figure 6 The bitstream 628 can be received by decoder 600 and input to entropy decoder processor 604, which decodes a portion of the bitstream's entropy into quantization coefficients. The quantization coefficients can be provided to inverse quantization and inverse transform processor 608, which performs inverse quantization and inverse transform to create a residual signal. This residual signal can be added to the output of motion compensation processor 620 or intra-frame prediction processor 624, depending on the processing mode. The outputs of motion compensation processor 620 and intra-frame prediction processor 624 can contain block predictions based on previously decoded blocks. The sum of the predictions and residuals can be processed by deblocking filter 612 and stored in frame buffer 616.

[0041] Figure 7 The flowchart illustrating an exemplary embodiment of process 700 shows how process 700 encodes video based on a global motion vector-based merge candidate reordering of some aspects of the current topic. This process can reduce encoding complexity while improving compression efficiency. In step 705, video frames may undergo initial block partitioning, for example, using a tree-structured macroblock partitioning scheme that may include partitioning image frames into CTUs and CUs.

[0042] In step 710, a candidate list can be determined. The candidate list may be based on the global motion of the current block. The candidate list may contain motion vector candidates that have motion information characterizing the global motion vector. The motion vector candidate list may be reordered such that the motion vector candidate with motion information characterizing the global motion vector is placed first in the reordered motion vector candidate list. Reordering may involve inserting the first global motion vector candidate into the merged candidate list. In some implementations, the construction of the motion vector candidate list may involve reordering.

[0043] In step 715, the block can be encoded and included in the bitstream. As a non-limiting example, the encoding may include the use of inter-frame prediction modes and intra-frame prediction modes. More specifically, the indices of a reordered candidate list may be included and / or encoded into the bitstream for use by the decoder.

[0044] Figure 8 To illustrate the system block diagram of an example video encoder 800, which is capable of encoding video based on a merged candidate reordering of global motion vectors. The example video encoder 800 can receive an input video 804, which can be initially partitioned or divided according to a tree-structured macroblock partitioning scheme (e.g., quadtree plus binary tree). Examples of tree-structured macroblock partitioning schemes may include dividing image frames into large blocks of elements called coding tree units (CTUs). In some implementations, each CTU may be further divided into one or more sub-blocks called coding units (CUs). The final result of this partitioning may include a set of sub-blocks called prediction units (PUs). Transform units (TUs) may also be used.

[0045] Still referencing Figure 8 The example video encoder 800 may include an intra-frame prediction processor 808, a motion estimation / compensation processor 812 (which may also be referred to as an inter-frame prediction processor), and is capable of constructing a motion vector candidate list, including adding global motion vector candidates to the motion vector candidate list, a transform / quantization processor 816, an inverse quantization / inverse transform processor 820, an in-loop filter 824, a decoded image buffer 828, and / or an entropy coding processor 832. Bitstream parameters may be input to the entropy coding processor 832 to be included in the output bitstream 836.

[0046] During operation, and continue to refer to Figure 8 For each block of a frame in the input video 804, it can be determined whether the block will be processed by intra-frame prediction or by motion estimation / compensation. The block can be provided to either the intra-frame prediction processor 808 or the motion estimation / compensation processor 812. If the block is to be processed by intra-frame prediction, the intra-frame prediction processor 808 can perform processing to output predicted values. If the block is to be processed by motion estimation / compensation, the motion estimation / compensation processor 812 can perform processing including constructing a list of motion vector candidates, and, where applicable, adding global motion vector candidates to the list of motion vector candidates.

[0047] Further reference Figure 8The residual can be formed by subtracting the predictor from the input video. The residual can be received by a transform / quantization processor 816, which performs transform processing (e.g., Discrete Cosine Transform (DCT)) to produce quantizable coefficients. The quantization coefficients and any associated signal transmission information can be provided to an entropy coding processor 832 for entropy coding and included in the output bitstream 836. The entropy coding processor 832 can support the encoding of signal transmission information related to the encoding of the current block. Furthermore, the quantization coefficients can be provided to an inverse quantization / inverse transform processor 820, which can reproduce pixels that can be combined with the predictor and processed by a loop filter 824. The output of the loop filter 824 can be stored in a decoded image buffer 828 for use by a motion estimation / compensation processor 812 capable of constructing a list of motion vector candidates, including adding global motion vector candidates to the motion vector candidate list.

[0048] Continue to refer to Figure 8 Although some variations have been described in detail above, other modifications or additions are possible. For example, in some implementations, the current block may include any symmetric block (8x8, 16x16, 32x32, 64x64, 128x128, etc.) as well as any asymmetric block (8x4, 16x8, etc.).

[0049] In some implementations, and still referencing Figure 8 This allows for the implementation of a quadtree plus binary decision tree (QTBT). In QTBT, at the encoding tree unit layer, the partitioning parameters of the QTBT are dynamically derived to adapt to local characteristics without transmitting any overhead. Subsequently, at the encoding unit layer, the joint classifier decision tree structure can eliminate unnecessary iterations and control the risk of mispredictions. In some implementations, the LTR frame block update mode can be used as an additional option available at each leaf node of the QTBT.

[0050] In some implementations, and still referencing Figure 8 Additional syntax elements can be signaled at different levels of the bitstream. For example, an enable flag can be included in the Sequence Parameter Set (SPS) to enable the entire sequence. Furthermore, the CTU flag can be encoded at the Code Tree Unit (CTU) level.

[0051] It should be noted that, as will be apparent to those skilled in the art of computers, any one or more aspects and embodiments described herein can be readily implemented using digital electronic circuits, integrated circuits, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof implemented in one or more machines programmed according to the teachings of this specification (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices such as document servers, etc.). These aspects or features may include implementations in one or more computer programs and / or software executable and / or interpretable on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to the storage system, at least one input device, and at least one output device. It will be apparent to those skilled in the art of software that a skilled programmer can readily prepare appropriate software code based on the teachings of this disclosure. The aspects and implementations employing software and / or software modules discussed above may also include appropriate hardware for assisting in the implementation of machine-executable instructions for the software and / or software modules.

[0052] Such software can be a computer program product that employs a machine-readable storage medium. A machine-readable storage medium can be any medium capable of storing and / or encoding a sequence of instructions executable by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory “ROM” devices, random access memory “RAM” devices, magnetic cards, optical cards, solid-state storage devices, EPROMs, EEPROMs, programmable logic devices (PLDs), and / or any combination thereof. As used herein, machine-readable media are intended to include both single media and collections of physically separate media, such as collections of optical disks or one or more hard disk drives combined with computer memory. As used herein, machine-readable storage media do not include transient signal transmissions.

[0053] Such software may also include information (e.g., data) carried as data signals on a data carrier such as a carrier wave. For example, machine-executable information may include data-bearing signals contained in a data carrier, wherein the signals encode a sequence of instructions or a portion thereof for execution by a machine (e.g., a computing device), and any relevant information (e.g., data structures and data) that causes the machine to perform any of the methods and / or embodiments described herein.

[0054] Examples of computing devices include, but are not limited to, e-book readers, computer workstations, terminal computers, server computers, handheld devices (e.g., tablets, smartphones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions specifying an action to be taken by the machine, and any combination thereof. In one example, a computing device may include and / or be included in a kiosk.

[0055] Figure 9 An illustration of one embodiment is shown, illustrating a computing device of exemplary form, a computer system 900 in which a set of instructions can be executed to cause a control system to perform any one or more aspects and / or methods of this disclosure. It is also contemplated that multiple computing devices may be utilized to implement specially configured sets of instructions for causing one or more of the aspects and / or methods of the present invention to perform. The computer system 900 includes a processor 904 and a memory 908 that communicate with each other and with other components via a bus 912. The bus 912 may include any of several types of bus architectures, including but not limited to memory buses, memory controllers, peripheral buses, local buses, and any combinations thereof using any of various bus architectures.

[0056] Memory 908 may include various components (e.g., machine-readable media), including but not limited to random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 916 (BIOS) may be stored in memory 908, including, for example, basic routines that facilitate the transfer of information between elements within computer system 900 during startup. Memory 908 may also include (e.g., stored on one or more machine-readable media) instructions (e.g., software) 920 embodying any one or more aspects and / or methods of this disclosure. In another example, memory 908 may also include any number of program modules, including but not limited to an operating system, one or more application programs, other program modules, program data, and any combination thereof.

[0057] Computer system 900 may also include storage device 924. Examples of storage devices (e.g., storage device 924) include, but are not limited to, hard disk drives, disk drives, optical disk drives combined with optical media, solid-state storage devices, and any combination thereof. Storage device 924 may be connected to bus 912 via a suitable interface (not shown). Exemplary interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FireWire), and any combination thereof. In one example, storage device 924 (or one or more components of storage device 924) may be removably interfaced to computer system 900 (e.g., via an external port connector (not shown)). In particular, storage device 924 and associated machine-readable medium 928 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 900. In one example, software 920 may reside wholly or partially within machine-readable medium 928. In another example, software 920 may reside wholly or partially within processor 904.

[0058] Computer system 900 may also include input device 932. In one example, a user of computer system 900 may input commands and / or other information into computer system 900 via input device 932. Examples of input device 932 include (but are not limited to) alphanumeric input devices (e.g., keyboard), positioning devices, joysticks, game controllers, audio input devices (e.g., microphone, voice response system, etc.), cursor control devices (e.g., mouse), touchpads, optical scanners, video capture devices (e.g., still cameras, video cameras), touchscreens, and any combination thereof. Input device 932 may be connected to bus 912 via any of a variety of interfaces (not shown), including but not limited to serial interfaces, parallel interfaces, game ports, USB interfaces, firewire interfaces, direct interfaces to bus 912, and any combination thereof. Input device 932 may include a touchscreen interface, which may be part of or separate from display 936, as will be discussed further below. Input device 932 may be used as a user selection device for selecting one or more graphical representations in the graphical interface described above.

[0059] Users can also input commands and / or other information to computer system 900 via storage device 924 (e.g., removable disk drive, flash drive, etc.) and / or network interface device 940. Network interface devices, such as network interface device 940, can be used to connect computer system 900 to one or more networks, such as network 944, and one or more remote devices 948 connected to network 944. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, enterprise networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical areas), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof. Networks such as network 944 can employ wired and / or wireless communication modes. Typically, any network topology can be used. Information (e.g., data, software 920, etc.) can be transmitted to and / or from computer system 900 via network interface device 940.

[0060] Computer system 900 may also include a video display adapter 952 for transmitting displayable images to a display device, such as display device 936. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tube (CRTs), plasma displays, light-emitting diode (LED) displays, and any combination thereof. Display adapter 952 and display device 936 may be used in conjunction with processor 904 to provide graphical representations of various aspects of this disclosure. In addition to display devices, computer system 900 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to bus 912 via peripheral interface 956. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, firewire connections, parallel connections, and any combination thereof.

[0061] The foregoing is a detailed description of illustrative embodiments of the present invention. Various modifications and additions can be made without departing from the spirit and scope of the invention. Features of each of the above embodiments can be suitably combined with features of other described embodiments to provide multiple combinations of features in related new embodiments. Furthermore, although several individual embodiments have been described above, the content described herein is merely an illustration of the application of the principles of the invention. Additionally, although specific methods herein may be shown and / or described as being performed in a particular order, this order is highly variable to those skilled in the art in implementing the embodiments disclosed herein. Therefore, this specification is by way of example only and is not intended to limit the scope of the invention.

[0062] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear, followed by a combined list of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless implied or explicitly contradicted to the context in which it is used, such phrases are intended to mean any element or feature listed individually, or any referenced element or feature in combination with any other referenced element or feature. For example, the phrases “at least one of A and B;”, “one or more of A and B;”, and “A and / or B” each respectively mean “A alone, B alone, or A and B together.” Similar interpretations are also intended for lists comprising three or more items. For example, the phrases “at least one of A, B, and C;”, “one or more of A, B, and C;”, and “A, B, and / or C” each respectively mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Additionally, the term “based on” as used above and in the claims is intended to mean “at least partially based on,” allowing for the inclusion of unreferenced features or elements.

[0063] Depending on the desired configuration, the subject matter described herein can be embodied in systems, devices, methods, and / or articles. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely examples consistent with aspects relating to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, additional features and / or variations may be provided in addition to those set forth herein. For example, the foregoing embodiments may involve various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of several other features disclosed above. Furthermore, the logical flows depicted in the figures and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other embodiments are within the scope of the appended claims.

Claims

1. A video encoder, comprising circuitry configured to: Receive video signals; Generate an encoded bitstream including an encoded image, the encoded image comprising a first region having a first consecutive plurality of coding units and a second region having a second consecutive plurality of coding units, the bitstream being decodeable by a decoder configured to receive the bitstream, and further configured to: For each coding unit in the first region, a motion vector candidate list is constructed, and each of the motion vector candidate lists has a common motion vector, wherein, The order of motion vector candidates in each of the motion vector candidate lists is determined such that the common motion vector is ranked first; The first plurality of coding units are decoded using common motion vectors from the motion vector candidate list, thereby reconstructing an image region with common motion in the first region; Determine an independently determined motion vector for each coding unit in the second region from the bit stream, wherein adjacent coding units in the second region have different independently determined motion vectors, and each independently determined motion vector is either a translational motion vector or a control point motion vector for affine motion; as well as The second plurality of coding units are decoded using the independently determined motion vectors to reconstruct the local motion in the second region.

2. The encoder according to claim 1, wherein, The decoder that receives the bitstream is configured to determine whether global motion is indicated for the encoded image.

3. The encoder according to claim 1, wherein, The common motion vector includes the control point motion vector.

4. The encoder according to claim 3, wherein, The motion vector of the control point is a translational motion vector.

5. The encoder according to claim 3, wherein, The motion vector of the control point is a vector of a four-parameter affine motion model.

6. The encoder according to claim 3, wherein, The control point motion vector is a vector of a six-parameter affine motion model.

7. A decoder, the decoder comprising circuitry configured to: Receive a bitstream including an encoded image, the encoded image including a first region with common motion having a first consecutive plurality of coding units, a second region with local motion having a second consecutive plurality of coding units, and intra-frame prediction coding units; For each coding unit in the first region, a motion vector candidate list is constructed, and each motion vector candidate list has a common motion vector, wherein, The candidate list of motion vectors is sorted such that the common motion vector is ranked first; The first plurality of coding units are decoded using common motion vectors from the motion vector candidate list, thereby reconstructing an image region with common motion in the first region; Determine an independently determined motion vector for each coding unit in the second consecutive plurality of coding units from the bit stream, wherein adjacent coding units in the second consecutive plurality of coding units have different independently determined motion vectors, and each independently determined motion vector is one of a translational motion vector for translational motion or a control point motion vector for four-parameter or six-parameter affine motion; The second plurality of coding units are decoded using the independently determined motion vectors to reconstruct the local motion in the second region; and The intra-frame predictive coding unit is decoded using the prediction of the previously decoded coding unit in the image.

8. A video encoder, comprising circuitry configured to: Receive video signals; Generate an encoded bitstream including an encoded image, the encoded image including a first region with common motion having a first consecutive plurality of coding units, a second region with local motion having a second consecutive plurality of coding units, and intra-frame prediction coding units, the bitstream being encoded and decoded by a decoder, the decoder being configured to receive the encoded bitstream and execute a decoding method, the decoding method including: A motion vector candidate list is constructed for each coding unit in the first region, and each motion vector candidate list has a common motion vector, wherein the motion vector candidate lists are sorted such that the common motion vector is ranked first; The first plurality of coding units are decoded using common motion vectors from the motion vector candidate list, thereby reconstructing an image region with common motion in the first region; Determine an independently determined motion vector for each coding unit in the second consecutive plurality of coding units from the bit stream, wherein adjacent coding units in the second consecutive plurality of coding units have different independently determined motion vectors, and each of the independently determined motion vectors is either a translational motion vector for translational motion or a control point motion vector for four-parameter or six-parameter affine motion; The second plurality of coding units are decoded using the independently determined motion vectors to reconstruct the local motion in the second region; and The intra-frame predictive coding unit is decoded using the prediction of the previously decoded coding unit in the image.

9. A method for transmitting a video signal as an coded bitstream, the method comprising: Receive source video signal; Generate an coded bitstream representing the source video signal, the coded bitstream including an coded image, the coded image including a first region having a first plurality of coding units and a second region having a second plurality of coding units, the coded bitstream being further configured to be decoded by a decoding method, the decoding method including: A motion vector candidate list is constructed for each coding unit in the first region, and each motion vector candidate list has a common motion vector, wherein the motion vector candidate lists are sorted such that the common motion vector is ranked first; The first plurality of coding units are decoded using common motion vectors from the motion vector candidate list, thereby reconstructing an image region with common motion in the first region; Determine an independently determined motion vector for each coding unit in the second region from the bitstream, wherein adjacent coding units in the second region have different independently determined motion vectors, each independently determined motion vector being either a translational motion vector or a control point motion vector for four-parameter or six-parameter affine motion; and The second plurality of coding units are decoded using the independently determined motion vectors to reconstruct the local motion in the second region; and The encoded bit stream is transmitted through a communication channel.

10. The method according to claim 9, wherein, The common motion vector includes the control point motion vector.

11. The method according to claim 10, wherein, The motion vector of the control point is a translational motion vector.

12. The method according to claim 10, wherein, The motion vector of the control point is a vector of a four-parameter affine motion model.

13. The method according to claim 10, wherein, The control point motion vector is a vector of a six-parameter affine motion model.