Global motion constrained motion vector in inter prediction
By identifying global motion model parameters in the frame header and using global motion vectors for motion compensation, the problem of low coding efficiency in inter-frame prediction is solved, achieving more efficient video compression and reducing device complexity, making it suitable for low-power devices and real-time applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-24
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video coding techniques fail to effectively utilize global motion information in inter-frame prediction, resulting in low coding efficiency and high decoding complexity.
By identifying the parameters of the global motion model in the frame header and using global motion vectors for motion compensation, the motion model of the block is restricted to reduce complexity and improve compression efficiency.
It improves the compression efficiency of video encoding, reduces the complexity of encoders and decoders, and is suitable for low-power devices and real-time applications.
Smart Images

Figure CN121842382A_ABST
Abstract
Description
[0001] Divisional Application This application is a divisional application of the application entitled “GLOBAL MOTION CONSTRAINED MOTION VECTOR IN INTER PREDICTION” having an application date of April 24, 2020, application number 202080046912.6. Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application entitled “GLOBAL MOTION CONSTRAINED MOTION VECTOR IN INTER PREDICTION” having application number 63 / 838 563, filed on April 25, 2019, which is incorporated by reference in its entirety. TECHNICAL FIELD
[0003] The present invention relates generally to the field of video compression. In particular, the present invention relates to global motion constrained motion vector in inter prediction. BACKGROUND
[0004] A video codec can include electronic circuits or software that compress and decompress digital video. It can convert uncompressed video into a compressed format, and vice versa. In the context of video compression, a device that compresses video (and / or performs some function thereof) can generally be referred to as an encoder, while a device that decompresses video (and / or performs some function thereof) can be referred to as a decoder.
[0005] The format of the compressed data can conform to a standard video compression specification. Compression can be lossy, where some information present in the original video is lacking in the compressed video. As a result of not having enough information to accurately reconstruct the original video, the result of compression can include that the quality of the decompressed video can be lower than the original uncompressed video.
[0006] There can be a complex relationship between video quality, amount of data used to represent the video (e.g., determined by bit rate), complexity of encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, end-to-end delay (e.g., latency), etc.
[0007] Motion compensation can include methods for predicting video frames or portions thereof, given a reference frame (e.g., previous and / or future frames), by taking into account camera motion and / or the motion of objects in the video. It can be applied to the encoding and decoding of video data for video compression, for example, to encoding and decoding using the Motion Picture Experts Group (MPEG)-2 (also known as High-Level Video Coding (AVC) and H.264) standards. Motion compensation can describe an image based on the transformation from a reference image to the current image. The reference image can be earlier in time compared to the current image; or it can be future in time compared to the current image. Compression efficiency can be improved when images can be accurately synthesized from previously transmitted and / or stored images. Summary of the Invention
[0008] In one aspect, a decoder includes circuitry configured to: receive a bitstream; extract a frame header associated with a current frame, the frame header including a signal characterizing that global motion has been enabled, and a signal further characterizing parameters of a global motion model; and decode the current frame, the decoding including: for each current block, using a motion model with a complexity less than or equal to that of the global motion model.
[0009] On the other hand, one approach includes receiving a bitstream via a decoder. This approach includes extracting a frame header associated with the current frame, the frame header including signals indicating that global motion has been enabled, and signals further characterizing parameters of the motion model. This approach includes decoding the current frame, the decoding including, for each current block, using a motion model with a complexity less than or equal to that of the global motion model.
[0010] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, as well as from the claims. Attached Figure Description
[0011] To illustrate the invention, the accompanying drawings show aspects of one or more embodiments of the invention. However, it should be understood that the invention is not limited to the precise arrangements and means shown in the drawings, wherein: Figure 3 This is a process flowchart based on some example implementations of the current topic; Figure 4 This is a system block diagram of an example decoder based on some example implementations of the current topic; Figure 5 This is a process flowchart based on some example implementations of the current topic; Figure 6This is a system block diagram of an example decoder based on some example implementations of the current topic; Figure 1 This is a graph showing the motion vectors of an example frame with global and local motion; Figure 2 Three example motion models that can be used for global motion are shown, including their index values (0, 1, or 2); and, Figure 7 This is a block diagram of a computing system that can be used to implement any one or more methods disclosed herein and any one or more parts thereof.
[0012] The accompanying drawings are not necessarily drawn to scale and may be illustrated by dashed lines, diagrams, and partial views. In some cases, details that are not essential for understanding the embodiments or that make other details difficult to perceive may have been omitted. The same reference numerals in the figures denote the same elements. Detailed Implementation
[0013] Global motion in a video refers to motion that occurs throughout the entire frame. Global motion may be caused by camera movement; for example, camera panning and zooming produce motion within a frame that typically affects the entire frame. Motion present in portions of a video can be called local motion. Local motion can be caused by moving objects in the scene, such as, but not limited to, objects moving from left to right in the scene. Video may contain a combination of local and global motion. Some embodiments of the present topic can provide efficient methods for transmitting global motion to the decoder, as well as efficient methods for using global motion vectors to improve compression efficiency.
[0014] Figure 1 This is a diagram illustrating the motion vectors of an example frame 100 with global and local motion. Frame 100 includes: a plurality of pixel blocks illustrated as squares, and their associated motion vectors illustrated as arrows. Squares with upward and leftward arrows (e.g., pixel blocks) represent blocks with motion that can be considered global motion, while squares with arrows pointing in other directions (represented by 104) represent blocks with local motion. Figure 1In the example shown, many blocks share the same global motion. Signaling the global motion in the frame header, such as a Picture Parameter Set (PPS) or Sequence Parameter Set (SPS), and using this signaled global motion, reduces the amount of motion vector information required for each block, thus improving prediction. While the examples described below for illustrative purposes involve determining and / or applying global or local motion vectors at the block level, global motion vectors can be determined and / or applied to any region of a frame and / or picture, including regions consisting of multiple blocks, regions defined by any geometry (e.g., but not limited to, regions defined by geometric and / or exponential coding, where one or more straight lines and / or curves defining the shape may be angled and / or curved), and / or the entire frame and / or picture. Although signaling is described herein as being performed at the frame level and / or in the frame header and / or parameter set, signaling may also be performed alternatively or additionally at the sub-picture level, where sub-pictures may include any region of the frame and / or picture as described above.
[0015] As an example, and continuing to refer to Figure 1 Simple translational motion can be described using motion vectors (MVs), which have two components, MVx and MVy, to describe the displacement of blocks and / or pixels in the current frame. More complex motions, such as rotation, scaling, and warping, can be described using affine motion vectors; where, as used in this disclosure, an "affine motion vector" is a vector describing the uniform displacement of a set of pixels or points represented in a video picture and / or a picture, for example, a set of pixels representing the movement of an object's view in a video without changing its appearance during the motion. Some video encoding and / or decoding methods may use 4-parameter or 6-parameter affine models for motion compensation in inter-picture encoding.
[0016] For example, six-parameter affine motion can be described as: x' = ax + by + c y' = dx + ey + f Four-parameter affine motion can be described as: x' = ax + by + c y' = -bx + ay + f Where (x,y) and (x',y') are the pixel positions in the current image and the reference image, respectively; a, b, c, d, e, and f are the parameters of the affine motion model.
[0017] Still refer to Figure 1The parameters describing affine motion can be used to signal the decoder so that affine motion compensation can be applied at the decoder. In some methods, the motion parameters may be sent explicitly; or, they may be sent via translational control point motion vectors (CPMVs), from which the affine motion parameters are derived. Two control point motion vectors (CPMVs) can be used to derive the affine motion parameters of a four-parameter affine motion model, and three control point translational motion vectors (CPMVs) can be used to obtain the parameters of a six-parameter motion model. Using control point motion vectors to signal the affine motion parameters allows for the use of efficient motion vector encoding methods to transmit these signals.
[0018] In some implementations, reference continues. Figure 1 Global motion signaling can be included in the frame header, such as PPS or SPS. Global motion may vary from image to image. The motion vector identified in the image's frame header can describe the motion relative to previously decoded frames. In some implementations, global motion can be translational or affine. The motion model used (e.g., the number of parameters, whether the model is affine, translational, or otherwise) can also be identified in the image's frame header. Figure 2 Three example motion models 200 that can be used for global motion are shown, including their index values (0, 1, or 2).
[0019] Still refer to Figure 2 PPS can be used to identify parameters that can change between picture sequences. Parameters that remain consistent across picture sequences can be identified in a sequence parameter set to reduce the size of the PPS and lower the video bitrate. Table 1 shows an example picture parameter set (PPS):
[0020] Additional fields can be added to the PPS to identify global motion. In the case of global motion, the presence of global motion parameters in the image sequence can be identified in the SPS; the PPS can reference the SPS via the SPS ID. In some decoding methods, the SPS can be modified by adding a field to identify the presence of global motion parameters in the SPS. For example, a one-bit field can be added to the SPS. If the bit global_motion_present is 1, global motion-related parameters can be expected in the PPS; if the bit global_motion_present is 0, fields related to global motion parameters may not exist in the PPS. For example, the PPS in Table 1 can be expanded to include the global_motion_present field, as shown in Table 2:
[0021] Similarly, PPS can include the pps_global_motion_parameters field in the frame, for example, as shown in Table 3:
[0022] More specifically, PPS may include fields that characterize global motion parameters using control point motion vectors, for example, as shown in Table 4:
[0023] As a further non-limiting example, Table 5 below can represent exemplary SPS:
[0024] The SPS table can be extended as described above to incorporate global motion presence indicators as shown in Table 6:
[0025] Additional fields may be incorporated into the SPS to reflect additional indicators as described in this disclosure.
[0026] In one embodiment, still referencing Figure 2In PPS and / or SPS, the `sps_affine_enabled_flag` specifies whether affine-based motion compensation can be used for inter-frame prediction. If `sps_affine_enabled_flag` equals 0, the syntax is constrained so that affine-based motion compensation is not used in code-later video sequences (CLVS), and `inter_affine_flag` and `cu_affine_type_flag` are not required in the CLVS coding unit syntax. Otherwise (`sps_affine_enabled_flag` equals 1), affine-based motion compensation can be used in CLVS.
[0027] Continue to refer to Figure 2 The `sps_affine_enabled_flag` in PPS and / or SPS specifies whether motion compensation based on a 6-parameter affine model can be used for inter-frame prediction. If `sps_affine_enabled_flag` equals 0, the syntax can be constrained so that motion compensation based on a 6-parameter affine model is not used in CLVS, and `cu_affine_type_flag` may not exist in the CLVS coding unit syntax. Otherwise (`sps_affine_enabled_flag` equals 1), motion compensation based on a 6-parameter affine model can be used in CLVS. When `sps_affine_type_flag` does not exist, it can be inferred that its value is equal to 0.
[0028] Still refer to Figure 2 Translation CPMV can be identified in PPS. Control points can be predefined. For example, control point MV0 can be relative to the top left corner of the image, MV1 can be relative to the top right corner, and MV3 can be relative to the bottom left corner. Table 4 shows an example method for identifying CPMV data based on the motion model used.
[0029] In one exemplary embodiment, still referring to Figure 2, the array amvr_precision_idx can be identified in coding units, coding trees, etc. The array amvr_precision_idx can specify the resolution AmvrShift of the motion vector difference, which can be defined as a non - restrictive example shown in Table 7 below. The array indices x0, y0 can indicate: the position (x0, y0) of the top - left luminance sample of the considered coding block relative to the top - left luminance sample of the picture; when amvr_precision_idx[x0][y0] does not exist, it can be inferred to be equal to 0. When inter_affine_flag[x0][y0] is equal to 0, the variables MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdL1[x0][y0][0], MvdL1[x0][y0][1] represent the modulation vector differences corresponding to the considered block, and these values can be modified by shifting them by AmvrShift, for example, using: MvdL0[x0][y0][0]=MvdL0[x0][y0][0]<<AmvrShift; MvdL0[x0][y0][1]=MvdL0[x0][y0][1]<<AmvrShift; MvdL1[x0][y0][0]=MvdL1[x0][y0][0]<<AmvrShift; and MvdL1[x0][y0][1]=MvdL1[x0][y0][1]<<AmvrShift. When inter_affine_flag[x0][y0] is equal to 1, the variables MvdCpL0[x0][y0][0][0], MvdCpL0[x0][y0][0][1], MvdCpL0[x0][y0][1][0], MvdCpL0[x0][y0][1][1], MvdCpL0[x0][y0][2][0] and MvdCpL0[x0][y0][2][1] can be modified by shifting, for example, as follows: MvdCpL0[x0][y0][0][0]= MvdCpL0[x0][y0][0][0]<<AmvrShift; MvdCpL1[x0][y0][0][1]= MvdCpL1[x0][y0][0][1]<<AmvrShift; MvdCpL0[x0][y0][1][0]= MvdCpL0[x0][y0][1][0]<<AmvrShift; MvdCpL1[x0][y0][1][1]= MvdCpL1[x0][y0][1][1]<<AmvrShift; MvdCpL0[x0][y0][2][0]= MvdCpL0[x0][y0][2][0]<<AmvrShift; and, MvdCpL1[x0][y0][2][1]= MvdCpL1[x0][y0][2][1]<<AmvrShift.
[0030]
[0031] Further referring to Figure 2 ,global motion can be relative to a previously encoded frame. When there is only one set of global motion parameters, the motion may be relative to the frame presented immediately before the current frame.
[0032] Continuing to refer to Figure 2 ,global motion can represent the main motion in a frame. Many blocks in the frame may have motion very similar to the global motion. The exception may be blocks with local motion. Keeping the block motion compensation compatible with the global motion can reduce the complexity of the encoder and the decoder and improve the compression efficiency.
[0033] In some embodiments, still referring to Figure 2 ,if global motion is identified in a frame header such as PPS or SPS, the motion model in the SPS can be applied to all blocks in the picture. For example, if the global motion uses translational motion (e.g., motion model = 0), then all prediction units (PUs) in the frame may also be limited to translational motion (e.g., motion model = 0). In this case, the adaptive motion model may not be used. This can also be identified in the SPS using the use_gm_constrained_motion_models flag. When this flag is set to 1, the adaptive motion model may not be used in the decoder, and instead, a single motion model can be used for all PUs.
[0034] Still referring to Figure 2In some implementations of the present topic, motion signaling may not change across multiple PUs. Instead, a fixed motion model can be used by identifying the motion model once in the SPS. This approach can replace global motion. The use of a fixed motion model can be specified at the encoder to reduce complexity; for example, the encoder might be limited to a translation model, which may be advantageous for low-power devices, such as low-computation-power devices. For example, an affine motion model might not be used; this can be specified in the encoder profile. Such examples may be useful for real-time applications such as video conferencing, Internet of Things (IoT) infrastructure, security cameras, etc. By using a fixed motion model, it may be unnecessary to include redundant signaling in the bitstream.
[0035] Continue to refer to Figure 2 The current topic is not limited to coding techniques that utilize global motion, but can be applied to a wide range of coding techniques.
[0036] As mentioned above, and still referring to Figure 2 Global motion can represent the main motion in a frame. Many blocks in a frame may have motions very similar to the global motion, except for blocks with local motion. Maintaining block motion compensation compatible with global motion can reduce encoder and decoder complexity and improve compression efficiency.
[0037] Still referencing Figure 3 Instead of restricting the motion of each block to the same motion model (e.g., but not limited to the global motion model) and identifying it in the frame header (e.g., PPS or SPS), the motion model applied to each block in the frame can be restricted to a similar motion model. Similar motion models may include those with the same or lower complexity. For example, in the first column of Table 5, the following three models are shown in ascending order of complexity.
[0038]
[0039] Still refer to Figure 3 The second column of Table 5 shows the motion models allowed for blocks used in inter-frame coding. For example, some implementations of the current topic may allow the PU to employ a motion model whose index is less than or equal to the index of the motion model for the global motion.
[0040] Continue to refer to Figure 4 Maintaining compatibility between the motion model and global motion allows the use of global motion control points as candidates for motion vector prediction. Global motion CPMV can represent motions similar to those of PU and serves as a good candidate for MV prediction.
[0041] Figure 4This is a flowchart illustrating an exemplary embodiment of the process of applying a similar motion model to the global motion model identified in the frame header.
[0042] In step 305, the decoder receives a bitstream. The current block may be included within the bitstream received by the decoder. The bitstream may include, for example, data found in the bitstream that serves as input to the decoder when data compression is used. The bitstream may include information required for decoding the video. Receiving the bitstream may include extracting and / or parsing blocks and associated signaling information from the bitstream. In some implementations, the current block may include: a coding tree unit (CTU), a coding unit (CU), and / or a prediction unit (PU).
[0043] In step 310, still refer to Figure 5 The frame header associated with the current frame can be extracted. This header includes signals indicating that global motion has been enabled and signals further characterizing the parameters of the motion model. In step 315, the current frame may be decoded. Decoding may include, for each current block, using a motion model with a complexity less than or equal to that of the global motion model.
[0044] Figure 5 This is a system block diagram illustrating an exemplary embodiment of a decoder 400 capable of decoding bitstreams, including motion models applied to a global motion model similar to a frame header identifier. The decoder 400 may include: an entropy decoder processor 404, an inverse quantization and inverse transform processor 408, a deblocking filter 412, a frame buffer 416, a motion compensation processor 420, and / or an intra-frame prediction processor 424.
[0045] During operation, still refer to Figure 6 The bitstream 428 may be received by decoder 400 and input to entropy decoder processor 404; entropy decoder processor 404 may decode a portion of the bitstream's entropy into quantization coefficients. The quantization coefficients may be provided to inverse quantization and inverse transform processor 408, which may perform inverse quantization and inverse transform to create a residual signal, which may be added to the output of motion compensation processor 420 or intra-frame prediction processor 424, depending on the processing mode. The output of motion compensation processor 420 and / or intra-frame prediction processor 424 may include block prediction based on previously decoded blocks. The sum of prediction and residual may be processed by deblocking filter 412 and stored in frame buffer 416.
[0046] Figure 6This is a flowchart illustrating an exemplary process 500 for video encoding based on aspects and applications of the current topic, and a motion model similar to a global motion model identified by a frame header. This process can reduce encoding complexity while improving compression efficiency. In step 505, the video frame may undergo initial block segmentation, for example, using a tree-structured macroblock segmentation scheme, which may include segmenting the image frame into CTUs and CUs.
[0047] In step 510, still refer to Figure 6 The global motion of the current block or frame may be determined. In step 515, the block may be encoded and included in the bitstream. Encoding may include setting flags in the frame header of all blocks of the frame to apply a model similar to the global motion model. For example, encoding may include utilizing inter-frame prediction and intra-frame prediction modes.
[0048] Figure 6 This is a system block diagram illustrating an exemplary embodiment of a decoder 600 capable of applying a motion model similar to a frame header identifier to a global motion model. The example video encoder 600 may receive input video 604, which may undergo initial segmentation and / or partitioning according to a processing scheme such as a tree-structured macroblock partitioning scheme (e.g., quadtree plus binary tree). Examples of a tree-structured macroblock partitioning scheme may include partitioning picture frames into large blocks called coding tree units (CTUs). In some implementations, each CTU may be further partitioned once or multiple times into several sub-blocks called coding units (CUs). The final result of this partitioning may include a set of sub-blocks that may be called prediction units (PUs). Transform units (TUs) may also be used.
[0049] Still refer to Figure 6 The example video encoder 600 may include: an intra-frame prediction processor 612, a motion estimation / compensation processor 612 (also known as an inter-frame prediction processor, which is capable of constructing a list of motion vector candidates, including adding individual global motion vector candidates to the list), a transform / quantization processor 616, an inverse quantization / inverse transform processor 620, a loop filter 624, a decoded image buffer 628, and / or an entropy coding processor 632. Bitstream parameters may be input to the entropy coding processor 632 to be included in the output bitstream 636.
[0050] In operation, still refer to Figure 6For each block of the input video 604, it may be determined whether to process the block using intra-frame prediction or motion estimation / compensation. The block may be provided to either the intra-frame prediction processor 608 or the motion estimation / compensation processor 612. If the block is to be processed via intra-frame prediction, the intra-frame prediction processor 608 may perform processing to output predicted values. If the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 612 may perform processing, if applicable, including constructing a list of motion vector candidates (including adding individual global motion vector candidates to the list of motion vector candidates).
[0051] Still refer to Figure 7 The residual may be formed by subtracting the predicted value from the input video. The residual may be received by a transform / quantization processor 616, which performs transform processing, such as a discrete cosine transform (DCT), to produce quantizable coefficients. The quantization coefficients and any associated identification information may be provided to an entropy coding processor 632 for entropy coding and included in the output bitstream 636. The entropy coding processor 632 may support the encoding of identification information related to the current block. Furthermore, the quantization coefficients may be provided to an inverse quantization / inverse transform processor 620, which may reproduce pixels. These quantization coefficients may be combined with the predicted value and processed by a loop filter 624. The output of the loop filter 624 may be stored in a decoded image buffer 628 for use by a motion estimation / compensation processor 612, which is capable of constructing a motion vector candidate list, including adding individual global motion vector candidates to the motion vector candidate list.
[0052] Continue to refer to Although some variations have been described in detail above, other modifications or additions are possible. For example, in some implementations, the current block may include any symmetrical block (8x8, 16x16, 32x32, 64x64, 128x128, etc.) as well as any asymmetrical block (8x4, 16x8, etc.).
[0053] In some implementations, reference is still made to This could potentially implement a quadtree plus binary decision tree (QTBT). In QTBT, at the level of the encoding tree unit, the partitioning parameters of the QTBT can be dynamically derived to adapt to local characteristics without any propagation overhead. Subsequently, at the encoding unit level, the joint classifier decision tree structure can eliminate unnecessary iterations and control the risk of mispredictions. In some implementations, the LTR frame block update mode may be available as an additional option at each leaf node of the QTBT.
[0054] In some implementations, reference continues. Additional syntax elements may be identified at different levels of the bitstream. For example, an enable flag can be used for the entire sequence by including an enable flag encoded in the Sequence Parameter Set (SPS). Furthermore, CTU flags may be encoded at the Code Tree Unit (CTU) level.
[0055] It should be noted that any one or more aspects and embodiments described herein can be readily implemented using digital electronic circuits, integrated circuits, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof, such as on one or more machines programmed according to the teachings of this specification (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices such as document servers), as will be apparent to those skilled in the art of computers. These different aspects or features may include implementations in one or more computer programs and / or software that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transfer data and instructions to the storage system, at least one input device, and at least one output device. Based on the teachings of this disclosure, a skilled programmer can readily prepare appropriate software code, as will be apparent to those skilled in the art of software. The aspects and implementations of employing software and / or software modules discussed above may also include: appropriate hardware for assisting in the implementation of machine-executable instructions for the software and / or software modules.
[0056] Such software can be a computer program product employing a machine-readable storage medium. A machine-readable storage medium can be any medium capable of storing and / or encoding a sequence of instructions executable by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to: magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state storage devices, EPROMs, EEPROMs, programmable logic devices (PLDs), and / or any combination thereof. As used herein, machine-readable media are intended to include: single media as well as collections of physically separate media, such as a set of optical disks combined with one or more hard disk drives and computer memory. As used herein, machine-readable storage media do not include temporary forms of signal transmission.
[0057] Such software may also include: information (e.g., data) carried as data signals on a data carrier (e.g., a carrier wave). For example, machine-executable information may be contained in the data carrier as data-carrying signals; in the data carrier, signals encode a sequence of instructions or a portion thereof for execution by a machine (e.g., a computing device) and encode any associated information (e.g., data structures and data) to enable the machine to perform any of the methods and / or embodiments described herein.
[0058] Examples of computing devices include, but are not limited to: e-book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablets, smartphones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions (which specify the actions the machine should take), and any combination thereof. In one example, a computing device may be included in and / or incorporated into an information kiosk.
[0059] The illustration shows an embodiment of a computing device in an example form of a computer system 700, in which an instruction set may be executed to cause the control system to perform any or more aspects and / or methods of this disclosure. It is also contemplated that multiple computing devices may be used to implement a specially configured set of instructions to cause one or more devices to perform any or more aspects and / or methods of this disclosure. The computer system 700 includes a processor 704 and a memory 708, which communicate with each other and with other components via a bus 712. The bus 712 may include any of several types of bus architectures (using any of a variety of bus architectures), including but not limited to: a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof.
[0060] Memory 708 may include various components (e.g., machine-readable media), including but not limited to random access memory components, read-only components, and any combination thereof. In one example, the basic input / output system 716 (BIOS) includes basic routines that facilitate the transfer of information between elements within the computer system 700, such as during startup; these basic routines may be stored in memory 708. Memory 708 may also include instructions (e.g., software) 720 stored in one or more machine-readable media; these instructions implement any or more aspects and / or methods of this disclosure. In another example, memory 708 may also include any number of program modules, including but not limited to: an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0061] Computer system 700 may also include storage device 724. Examples of storage devices (e.g., storage device 724) include, but are not limited to, hard disk drives, disk drives, optical disc drives combined with optical media, solid-state storage devices, and any combination thereof. Storage device 724 may be connected to bus 712 via a suitable interface (not shown). Example interfaces include, but are not limited to, SCSI, advanced technology attachment (ATA), serial ATA, universal serial bus (USB), IEEE 1394 (FIREWIRE), and any combination thereof. In one example, storage device 724 (or one or more components thereof) may removably interact with computer system 700 (e.g., via an external port connector (not shown)). In particular, storage device 724 and associated machine-readable medium 728 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 700. In one example, software 720 may reside wholly or partially within machine-readable medium 728. In another example, software 720 may reside entirely or partially within processor 704.
[0062] Computer system 700 may also include input device 732. In one example, a user of computer system 700 may input commands and / or other information into computer system 700 via input device 732. Examples of input device 732 include, but are not limited to: alphanumeric input devices (e.g., keyboard), pointing devices, joysticks, game controllers, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touchpads, optical scanners, video capture devices (e.g., still cameras, camcorders), touchscreens, and any combination thereof. Input device 732 may be connected to bus 712 via any of a variety of interfaces (not shown), including but not limited to: serial interfaces, parallel interfaces, game ports, USB interfaces, firewire interfaces, direct interfaces connected to bus 712, and any combination thereof. Input device 732 may include a touchscreen interface, which may be part of or separate from display 736, as will be discussed further below. Input device 732 may be used as a user selection device for selecting one or more graphical representations within the graphical interface described above.
[0063] Users can also input commands and / or other information to computer system 700 via storage device 724 (e.g., removable disk drive, flash drive, etc.) and / or network interface device 740. Network interface devices, such as network interface device 740, can be used to connect computer system 700 to one or more networks (e.g., network 744) and one or more remote devices 748 connected to that network. Examples of network interface devices include, but are not limited to: network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to: wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical areas), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof. Networks, such as network 744, may employ wired and / or wireless communication modes. Typically, any network topology can be used. Information (e.g., data, software 720, etc.) can be transmitted to computer system 200 and / or transmitted out of computer system 700 via network interface device 740.
[0064] Computer system 700 may further include a video display adapter 752 for transmitting displayable images to a display device, such as display device 736. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tube displays (CRTs), plasma displays, light-emitting diode displays (LEDs), and any combination thereof. Display adapter 752 and display device 736 may be used in conjunction with processor 704 to provide a graphical representation of aspects of this disclosure. In addition to display devices, computer system 700 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. These peripheral output devices may be connected to bus 712 via peripheral interface 756. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, firewire connections, parallel connections, and any combination thereof.
[0065] The foregoing is a detailed description of illustrative embodiments of the present invention. Various modifications and additions can be made without departing from the spirit and scope of the invention. Features of each of the various embodiments described above can be suitably combined with features of other described embodiments to provide multiple feature combinations in associated new embodiments. Furthermore, while several individual embodiments have been described above, the description herein is merely an illustration of the application of the principles of the invention. Moreover, although specific methods herein may be illustrated and / or described as being performed in a particular order, the order is highly variable within the ordinary technical scope of implementing the embodiments disclosed herein. Therefore, this description is intended only as an example and does not otherwise limit the scope of the invention.
[0066] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear after a list of combinations of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless implied or explicitly contradicted by the context in which it is used, such phrases are intended to mean any element or feature listed individually, or any listed element or feature combined with any other listed element or feature. For example, the phrases “at least one of A and B;”, “one or more of A and B;”, and “A and / or B” respectively mean “A alone, B alone, or A and B together.” A similar interpretation applies to lists containing three or more items. For example, the phrases “at least one of A, B, and C;”, “one or more of A, B, and C;”, and “A, B, and / or C” respectively mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.” Furthermore, the term “based on” as used in the foregoing and claims is intended to mean “at least partially based on,” allowing for the inclusion of features or elements not listed.
[0067] Depending on the desired configuration, the subject matter described herein can be embodied in systems, devices, methods, and / or articles. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those set forth herein. For example, the above embodiments may involve various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of several further features described above. Furthermore, the logical flows depicted in the drawings and / or described herein do not necessarily need to be in the specific order or sequence shown to achieve the desired results. Other implementations may be within the scope of the following claims.
Claims
1. A decoder configured to: The receiver receives a bitstream comprising an encoded image having a first region and a second region. The first region has global motion and includes a first consecutive plurality of encoded blocks, and the second region has local motion and includes a second consecutive plurality of encoded blocks. The first region is encoded by a motion model that encodes all blocks of the first region in increasing order of complexity. The motion model is one of translational motion, four-parameter affine motion, or six-parameter affine motion. The complexity of the motion model is a function of the number of motion vectors required to inter-frame encode the blocks using the motion model. A motion model is used to decode the encoded image for each block of the encoded image, which includes the first consecutive plurality of encoded blocks and the second consecutive plurality of encoded blocks, wherein the complexity of the motion model for each block is no greater than the complexity of the first region of the global motion; The decoded image is stored in a buffer.
2. A method for transmitting an encoded bitstream, the encoded bitstream being decoded by a decoder receiving the bitstream, the method comprising: Receive input video signals; Generate an encoded bitstream, the encoded bitstream including an encoded image, the encoded image having a first region and a second region, the first region having global motion and including a first consecutive plurality of encoded blocks, the second region having local motion and including a second consecutive plurality of encoded blocks, the first region being encoded by a motion model in ascending order of complexity, the motion model being one of translational motion, four-parameter affine motion, or six-parameter affine motion, the complexity of the motion model being a function of the number of multiple motion vectors required to inter-frame encode the blocks using the motion model; The encoded bitstream is transmitted through a channel to a decoder, wherein the decoder has instructions and is configured to: A motion model is used to decode the encoded image for each block of the encoded image, which includes the first consecutive plurality of encoded blocks and the second consecutive plurality of encoded blocks, wherein the complexity of the motion model for each block is no greater than the complexity of the first region of the global motion; The decoded image is stored in a buffer.
3. A video encoder, comprising circuitry configured to: Receive input video signals; A encoded bitstream is generated, comprising an encoded image having a first region and a second region. The first region has global motion and includes a first series of consecutive encoded blocks, and the second region has local motion and includes a second series of consecutive encoded blocks. The first region is encoded by a motion model that encodes all blocks of the first region in increasing order of complexity. The motion model is one of translational motion, four-parameter affine motion, or six-parameter affine motion. The complexity of the motion model is a function of the number of motion vectors required to inter-frame encode the blocks using the motion model. The encoded bitstream is decoded by a method comprising: A motion model is used to decode the encoded image for each block of the encoded image, which includes the first consecutive plurality of encoded blocks and the second consecutive plurality of encoded blocks, wherein the complexity of the motion model for each block is no greater than the complexity of the first region of the global motion; The decoded image is stored in a buffer.