Selective motion vector prediction candidates in frames with global motion
By using a global motion model to select a single motion vector candidate for frames with both global and local motion, the patent addresses inefficiencies in video compression, improving encoding and decoding efficiency and reducing bitrate.
Patent Information
- Application Number
- JP2024179735
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2024-10-15
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2040-04-24
AI Technical Summary
Existing video compression technologies face inefficiencies in encoding and decoding processes due to the complex relationship between video quality, data usage, encoding complexity, sensitivity to errors, and the need for improved motion vector prediction, particularly in frames with both global and local motion.
Incorporating a global motion model to select a single global motion vector candidate for motion vector prediction, which is signaled in a header, reducing the need for additional motion vector information and improving compression efficiency by leveraging common motion across blocks in a frame.
This approach enhances compression efficiency by reducing the bitrate required for motion vector signaling and simplifying the encoding and decoding processes, while maintaining video quality.
Smart Images

Figure 0007716140000027 
Figure 0007716140000028 
Figure 0007716140000029
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 838,594, filed on April 25, 2019, entitled "SELECTIVE MOTION VECTOR PREDICTION CANDIDATES IN FRAMES WITH GLOBAL MOTION", which is incorporated herein by reference in its entirety.
[0002] The present invention generally relates to the field of video compression. In particular, the present invention is directed to selective motion vector prediction candidates in frames with global motion.
Background Art
[0003] A video codec can include an electronic circuit or software that compresses or decompresses digital video. This can convert uncompressed video into a compressed format and vice versa. In the context of video compression, a device that compresses video (and / or performs certain functions thereof) can typically be called an encoder, and a device that decompresses video (and / or performs certain functions thereof) can be called a decoder.
[0004] The format of the compressed data can conform to standard video compression specifications. Compression can be lossy in that the compressed video lacks certain information that was present in the original video. This can include the result that the decompressed video may have lower quality than the original uncompressed video because there is insufficient information to accurately reconstruct the original video.
[0005] There can be a complex relationship between video quality, the amount of data used to represent the video (e.g., determined by the bit rate), the complexity of the encoding and decoding algorithms, the sensitivity to data loss and errors, ease of editing, random access, end-to-end latency (e.g., waiting time), and equivalents.
[0006] Motion compensation can include an approach for predicting a video frame or a portion thereof, on the premise of a reference frame such as a past and / or future frame, by taking into account the motion of a camera and / or an object in the video. This can be adopted in the encoding and decoding of video data for video compression, for example, in the encoding and decoding using the Motion Picture Experts Group (MPEG)-2 standard (also referred to as Advanced Video Coding (AVC) and H.264). Motion compensation can describe a picture from the perspective of the transformation of a reference picture to the current picture. The reference picture can be temporally past when compared with the current picture and can be from the future when compared with the current picture. When an image can be accurately synthesized from images transmitted and / or stored in the past, the compression efficiency can be improved. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0007] In one aspect, a decoder includes a circuit network configured to perform steps of receiving a bitstream, constructing a motion vector candidate list for a current block, the step of constructing the motion vector candidate list further including adding a single global motion vector candidate to the motion vector candidate list, the single global motion vector candidate being selected based on a global motion model utilized by the current block, and reconstructing pixel data of the current block using the motion vector candidate list.
[0008] In another aspect, the method comprises receiving, by a decoder, a bitstream, and constructing, for a current block, a motion vector candidate list, the constructing of the motion vector candidate list further including adding a single global motion vector candidate to the motion vector candidate list, the single global motion vector candidate being selected based on a global motion model utilized by the current block, and reconstructing pixel data of the current block using the motion vector candidate list.
[0009] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. This specification also provides, for example, the following items. (Item 1) A decoder, wherein the decoder A circuit network, wherein the circuit network Receives a bitstream, and Constructs, for a current block, a motion vector candidate list, the constructing of the motion vector candidate list further including adding a single global motion vector candidate to the motion vector candidate list, the single global motion vector candidate being selected based on a global motion model utilized by the current block, and Reconstructs pixel data of the current block using the motion vector candidate list And is configured to perform A circuit network (Item 2) The decoder according to Item 1, wherein the single global motion vector candidate is selected based on a predefined mapping of global motion model types to candidates. (Item 3) The decoder according to item 1, wherein the single global motion vector candidate includes a control point motion vector. (Item 4) The decoder according to item 3, wherein the control point motion vector is a translational motion vector. (Item 5) The decoder according to item 3, wherein the control point motion vector is a vector of a 4-parameter affine motion model. (Item 6) The decoder according to item 3, wherein the control point motion vector is a vector of a 6-parameter affine motion model. (Item 7) An entropy decoder processor, wherein the entropy decoder processor is configured to receive the bitstream and decode the bitstream into quantized coefficients, an entropy decoder processor; An inverse quantization and inverse transformation processor, wherein the inverse quantization and inverse transformation processor is configured to process the quantized coefficients including performing an inverse discrete cosine, an inverse quantization and inverse transformation processor; A deblocking filter; A frame buffer; An intra prediction processor The decoder according to item 1, further comprising: (Item 8) The decoder according to item 1, wherein the current block forms part of a quadtree + binary decision tree. (Item 9) The decoder according to item 1, wherein the current block is a coding tree unit. (Item 10) The decoder according to item 1, wherein the current block is a coding unit. (Item 11) The decoder according to item 1, wherein the current block is a prediction unit. (Item 12) The decoder according to item 1, wherein the global motion model includes translational motion. (Item 13) The global motion model is the decoder according to item 1, including affine motion. (Item 14) The global motion model is the decoder according to item 1, characterized by the header of the bitstream, and the header includes a picture parameter set. (Item 15) The global motion model is the decoder according to item 1, characterized by the header of the bitstream, and the header includes a sequence parameter set (SPS). (Item 16) A method comprising: receiving, by a decoder, a bitstream; constructing, for a current block, a motion vector candidate list, where constructing the motion vector candidate list further includes adding a single global motion vector candidate to the motion vector candidate list, and the single global motion vector candidate is selected based on a global motion model used by the current block; reconstructing pixel data of the current block and using the motion vector candidate list. A method. (Item 17) The method according to item 16, wherein the single global motion vector candidate is selected based on a predefined mapping of global motion model types to candidates. (Item 18) The method according to item 16, wherein the single global motion vector candidate includes a control point motion vector. (Item 19) The method according to item 18, wherein the control point motion vector is a translational motion vector. (Item 20) The method according to item 18, wherein the control point motion vector is a vector of a 4-parameter affine motion model. (Item 21) The method according to item 18, wherein the control point motion vector is a vector of a 6-parameter affine motion model. (Item 22) The decoder further is an entropy decoder processor configured to receive the bitstream and decode the bitstream into quantized coefficients, an entropy decoder processor; is an inverse quantization and inverse transform processor configured to process the quantized coefficients including performing an inverse discrete cosine transform, an inverse quantization and inverse transform processor; a deblocking filter; a frame buffer; and an intra prediction processor The method according to item 16, comprising: (Item 23) The method according to item 16, wherein the current block forms part of a quadtree + binary decision tree. (Item 24) The method according to item 16, wherein the current block is a coding tree unit. (Item 25) The method according to item 16, wherein the current block is a coding unit. (Item 26) The method according to item 16, wherein the current block is a prediction unit. (Item 27) The method according to item 16, wherein the global motion model includes translational motion. (Item 28) The method according to item 16, wherein the global motion model includes affine motion. (Item 29) The method according to item 16, wherein the global motion model is characterized by a header of the bitstream, and the header includes a picture parameter set. (Item 30) The method according to item 16, wherein the global motion model is characterized by a header of the bitstream, and the header includes a sequence parameter set (SPS).
Brief Description of the Drawings
[0010] For purposes of illustrating the present invention, the drawings show aspects of one or more embodiments of the present invention. However, it should be understood that the present invention is not limited to the precise arrangements and means shown in the drawings.
[0011]
Figure 1
[0012]
Figure 2
[0013]
Figure 3
[0014]
Figure 4
[0015]
Figure 5
[0016]
Figure 6
[0017]
Figure 7
[0018] The drawings are not necessarily to scale and may be illustrated by phantom lines, diagrammatic representations, and fragmentary views. In some instances, details that are not necessary for an understanding of the embodiments or that render other details difficult to perceive may be omitted. Like reference numerals in the various drawings indicate like elements.
Best Mode for Carrying Out the Invention
[0019] Global motion in video refers to motion that occurs across an entire frame. Global motion can be caused by camera motion; for example, camera pans and zooms can generate motion that typically affects the entire frame within the video. Motion that exists within a portion of the video may be referred to as local motion. Local motion can be caused by moving objects within the scene, such as an object moving from left to right within the scene. A video may contain a combination of local motion and global motion. Some implementations of the present subject matter may provide an efficient approach for communicating global motion to a decoder and the use of global motion vectors in improving compression efficiency.
[0020] FIG. 1 is a schematic diagram illustrating the motion vectors of an exemplary frame 100 with global and local motion. Frame 100 may include some blocks of pixels, illustrated as a square, and their associated motion vectors, illustrated as arrows. A square (e.g., a block of pixels) with an arrow pointing to the upper left may indicate a block with motion that can be considered global motion, and a square with an arrow pointing in another direction (indicated by 104) may indicate a block with local motion. In the illustrated embodiment of FIG. 1, many blocks have the same global motion. Signaling the global motion in a header, such as a Picture Parameter Set (PPS) or a Sequence Parameter Set (SPS), and using the signal global motion can reduce the motion vector information required by the blocks and can result in improved prediction. For illustrative purposes, the embodiments described below refer to the determination and / or application of global or local motion vectors at the block level, but the global motion vector can be determined and / or applied with respect to any region of a frame and / or picture (a region composed of multiple blocks, including but not limited to a region bounded by one or more lines and / or curves that can be angled and / or curved within it, and / or a region defined by geometric and / or exponential coding, and / or a region bounded by any geometric form, and / or including the entire frame and / or picture). Signaling is described herein as being performed at the frame level and / or within a parameter set of a header and / or frame, but signaling can alternatively or additionally be performed at the sub-picture level, where the sub-picture can include any region of a frame and / or picture as described above.
[0021] As an example, still referring to FIG. 1, a simple translational motion is described by two component MVs that explain the displacement amount of blocks and / or pixels within the current frame x , MV yThe motion vector (MV) with can be described. More complex motions such as rotation, scaling, and warping can be described using affine motion vectors, and an "affine motion vector" as used in the present disclosure is a vector that describes a uniform displacement of a set of pixels or points, such as a set of pixels that illustrate an object moving across the field of view in a video without apparently changing its shape during the motion, represented within a video picture and / or within a picture. Some approaches to video encoding and / or decoding may use a 4-parameter or 6-parameter affine model for motion compensation in inter-picture coding. For example, a 6-parameter affine motion can be described as follows. x’ = ax + by + c y’ = dx + ey + f A 4-parameter affine motion can be described as follows. x’ = ax + by + c y’ = -bx + ay + f
[0022] Where (x, y) and (x’, y’) are pixel locations in the current picture and the reference picture, respectively, and a, b, c, d, e, and f are parameters of the affine motion model.
[0023] Continuing to refer to FIG. 1, the parameters used to explain the affine motion can be signaled to the decoder for applying affine motion compensation in the decoder. In some approaches, the motion parameters may be signaled explicitly or by signaling the translational control point motion vector (CPMV) and then deriving the affine motion parameters from the translational motion vector. Two control point motion vectors (CPMV) may be utilized to derive the affine motion parameters for a four-parameter affine motion model, and three control point translational motion vectors (CPMV) may be utilized to obtain the parameters for a six-parameter motion model. The step of signaling the affine motion parameters using the control point motion vectors may enable the use of an efficient motion vector coding method for signaling the affine motion parameters. In some implementations, but not limited to, the translational motion model may be indexed by 0, the four-parameter affine model may be indexed by 1, and the six-parameter affine model may be indexed by 2.
[0024] In one embodiment, still referring to FIG. 1, the sps_affine_enabled_flag in the PPS and / or SPS may define whether affine model-based motion compensation can be used for inter prediction. If the sps_affine_enabled_flag is equal to 0, the syntax may be constrained such that affine model-based motion compensation is not used within the coded video sequence (CLVS), and the inter_affine_flag and cu_affine_type_flag may not be present within the coding unit syntax of the CLVS. Otherwise (the sps_affine_enabled_flag is equal to 1), affine model-based motion compensation can be used within the CLVS.
[0025] Referring further to FIG. 1, the sps_affine_type_flag in PPS and / or SPS may define whether 6-parameter affine model-based motion compensation can be used for inter prediction. If the sps_affine_type_flag is equal to 0, the syntax may be constrained such that 6-parameter affine model-based motion compensation is not used within the CLVS, and the cu_affine_type_flag may not be present within the coding unit syntax of the CLVS. Otherwise (the sps_affine_type_flag is equal to 1), 6-parameter affine model-based motion compensation may be used within the CLVS. When not present, the value of the sps_affine_type_flag may be presumed to be equal to 0.
[0026] Continuing to refer to FIG. 1, the step of generating a list of MV prediction candidates can be a step that is performed in some compression approaches that utilize motion compensation in the decoder. Some past approaches may define the use of spatial and temporal motion vector candidates. Global motion signaled in a header such as SPS or PPS may indicate the presence of global motion in the video. Such global motion is expected to be common to most blocks within a frame. By using the global motion as a prediction candidate, motion vector coding can be improved and the bitrate can be reduced. The candidate MVs added to the MV prediction list may be selected according to the motion model used to represent the global motion and the motion model used in intercoding.
[0027] Still referring to FIG. 1, some implementations of global motion can be described using one or more control point motion vectors (CPMVs), depending on the motion model used. Thus, depending on the motion model used, 1 to 3 control point motion vectors may be available and may be used as candidates for prediction. In some implementations, all available CPMVs may be added to a list as prediction candidates. The common case of adding all available CPMVs can increase the likelihood of finding good motion vector predictions and improve compression efficiency. [Table 1]
[0028] Continuing to refer to FIG. 1, the step of generating a list of motion vector (MV) prediction candidates is a step in performing motion compensation in the decoder. Some existing approaches (e.g., past compression standards) define the use of spatial and temporal motion vector candidates.
[0029] Continuing to refer to FIG. 1, global motion can be signaled in a header such as the SPS and can indicate the presence of global motion in the video. Such global motion is likely to be present in many blocks within a frame. Thus, a given block is likely to have motion similar to the global motion. By using the global motion as a prediction candidate, motion vector coding can be improved and the bitrate can be reduced. The candidate MV may be added to the MV prediction list, and the candidate MV may be selected depending on the motion model used to represent the global motion and the motion model used in intercoding.
[0030] Still referring to FIG. 1, the global motion can be described using one or more control point motion vectors (CPMVs), depending on the motion model used. Thus, depending on the motion model used, 1 to 3 control point motion vectors may be available and may be used as candidates for prediction. A list of MV prediction candidates can be reduced by selectively adding one CPMV to the prediction candidate list. Reducing the list size can reduce the computational complexity and improve the compression efficiency. In some implementations, still referring to FIG. 1, the CPMV selected as a candidate may follow a predefined mapping, such as those illustrated in Table 1 below.
Table 2
[0031] The use of selective prediction candidates from global motion is signaled within the picture parameter set of the sequence parameter set, thereby reducing the encoding and decoding complexity.
[0032] Still referring to FIG. 1, since the block is likely to have motion similar to the global motion, adding the global motion vector as the first candidate in the prediction list can reduce the bits required to signal the prediction candidates used and / or reduce the bits required to encode the motion vector difference. As a further non-limiting example, FIG. 2 illustrates three exemplary motion models 200 that can be used for global motion, including their index values (0, 1, or 2).
[0033] Still referring to FIG. 2, the PPS may be used to signal parameters that may change between pictures of the sequence. Parameters that remain the same for a sequence of pictures are signaled within the sequence parameter set, reducing the size of the PPS and reducing the video bit rate. An exemplary picture parameter set (PPS) is shown in Table 2.
Table 3-1
Table 3-2
Table 3-3
Table 3-4
Table 3-5
[0034] Additional fields may be added to the PPS to signal global motion. In the case of global motion, the presence of global motion parameters within the picture sequence may be signaled within the SPS, and the PPS may reference the SPS by the SPS ID. The SPS may be modified to add a field for signaling the presence of global motion parameters within the SPS in some approaches to decoding. For example, a 1-bit field may be added to the SPS. If the global_motion_present bit is 1, global motion related parameters may be expected to be within the PPS. If the global_motion_present bit is 0, the global motion parameter related field may not need to be present within the PPS. For example, the PPS of Table 2 may be extended to include a global_motion_present field as shown in Table 3, for example.
Table 4-1
Table 4-2
[0035] Similarly, the PPS can include, for example, a pps_global_motion_parameters field for a frame as shown in Table 4.
Table 5
[0036] More specifically, the PPS may include, for example, a field for characterizing global motion parameters that uses control point motion vectors as shown in Table 5.
Table 6-1
Table 6-2
[0037] As a further non-limiting example, Table 6 below may represent an exemplary SPS.
Table 7-1
Table 7-2
Table 7-3
Table 7-4
Table 7-5
Table 7-6
Table 7-7
Table 7-8
Table 7-9
[0038] The SPS table as described above may be extended as described above in order to incorporate a global motion presence indicator as shown in Table 7.
Table 8
[0039] Additional fields may be incorporated within the SPS to reflect further indicators as described within this disclosure.
[0040] In one embodiment, still referring to FIG. 2, the sps_affine_enabled_flag within the PPS and / or SPS may define whether affine model-based motion compensation can be used for inter prediction. If the sps_affine_enabled_flag is equal to 0, the syntax may be constrained such that affine model-based motion compensation is not used within the coded video sequence (CLVS), and the inter_affine_flag and cu_affine_type_flag may not be present within the coding unit syntax of the CLVS. Otherwise (the sps_affine_enabled_flag is equal to 1), affine model-based motion compensation can be used within the CLVS.
[0041] Continuing to refer to FIG. 2, the sps_affine_type_flag in PPS and / or SPS may define whether 6-parameter affine model-based motion compensation can be used for inter prediction. If the sps_affine_type_flag is equal to 0, the syntax may be constrained such that 6-parameter affine model-based motion compensation is not used within the CLVS, and the cu_affine_type_flag may not be present within the coding unit syntax of the CLVS. Otherwise (the sps_affine_type_flag is equal to 1), 6-parameter affine model-based motion compensation may be used within the CLVS. When not present, the value of the sps_affine_type_flag may be assumed to be equal to 0.
[0042] Still referring to FIG. 2, the translational CPMV may be signaled within the PPS. The control points may be predefined. For example, control point MV 0 may be for the upper left corner of the picture, MV 1 may be for the upper right corner, and MV 3 may be for the lower left corner of the picture. Table 5 illustrates an exemplary approach for signaling CPMV data according to the motion model used.
[0043] In an exemplary embodiment, still referring to FIG. 2, an array amvr_precision_idx, which can be signaled within a coding unit, coding tree, or the like, may define a resolution AmvrShift of a motion vector difference, which can be defined as a non-limiting example as shown in Table 8 below. Array indices x0, y0 may define the location (x0, y0) of the top-left luminance sample of the coding block under consideration with respect to the top-left luminance sample of the picture, and when amvr_precision_idx[x0][y0] does not exist, this can be presumed to be equal to 0. When inter_affine_flag[x0][y0] is equal to 0, variables MvdL0[x0][y0][0], MvdL0[x0][y0][1], MvdL1[x0][y0][0], MvdL1[x0][y0][1], which represent the motion vector difference values corresponding to the block under consideration, may be modified, for example,
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Chemical formula
Table 9
[0044] FIG. 3 is a process flow diagram illustrating an exemplary embodiment of process 300 for constructing a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list.
[0045] In step 305, the current block is received by the decoder. The current block may be contained within the bitstream received by the decoder. The bitstream may include data found within a stream of bits, for example, an input to the decoder when using data compression. The bitstream may include information necessary to decode the video. The receiving step may include extracting and / or parsing the block and associated signaling information from the bitstream. In some implementations, the current block may include a coding tree unit (CTU), a coding unit (CU), or a prediction unit (PU).
[0046] In step 310, a motion vector candidate list may be constructed for the current block, and the constructing step may include adding a single global motion vector candidate to the motion vector candidate list. The single global motion vector candidate may be selected based on the global motion model utilized by the current block. In step 315, the pixel data of the current block may be reconstructed using the motion vector candidate list.
[0047] FIG. 4 is a system block diagram illustrating an exemplary decoder 400 capable of decoding a bitstream 428 by constructing a motion vector candidate list, including at least adding a single global motion vector candidate to the motion vector candidate list. The decoder 400 may include an entropy decoder processor 404, an inverse quantization and inverse transform processor 408, a deblocking filter 412, a frame buffer 416, a motion compensation processor 420, and / or an intra prediction processor 424.
[0048] During operation, still referring to FIG. 4, a bitstream 428 may be received by decoder 400 and input to entropy decoder processor 404, which may entropy decode a portion of the bitstream into quantized coefficients. The quantized coefficients may be provided to inverse quantization and inverse transform processor 408, which may perform inverse quantization and inverse transform to generate a residual signal, which may be added to the output of motion compensation processor 420 or intra prediction processor 424 according to the processing mode. The outputs of motion compensation processor 420 and intra prediction processor 424 may include block predictions based on previously decoded blocks. The sum of the prediction and the residual may be processed by deblocking filter 412 and stored in frame buffer 416.
[0049] FIG. 5 is a process flow diagram illustrating an exemplary embodiment of a process 500 for encoding video by constructing a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list according to some aspects of the present subject matter that can improve compression efficiency while reducing the complexity of the encoding steps. At step 505, a video frame may undergo an initial block segmentation, for example, by using a tree-structured macroblock partitioning scheme that may include partitioning a picture frame into CTUs and CUs. At step 510, global motion may be determined. At step 515, a block may be encoded and included in the bitstream. Encoding may include constructing a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list. Encoding may include steps such as utilizing an inter prediction mode and an intra prediction mode.
[0050] FIG. 6 is a system block diagram illustrating an exemplary embodiment of a video encoder 600 that can construct a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list. The exemplary video encoder 600 may receive an input video 604, which may first be segmented and / or partitioned according to a processing scheme such as a tree-structured macroblock partitioning scheme (e.g., quadtree + binary tree). An example of a tree-structured macroblock partitioning scheme may include the step of partitioning a picture frame into large block elements called coding tree units (CTUs). In some implementations, each CTU may be further partitioned one or more times into several sub-blocks called coding units (CUs). The final result of this partitioning may include a group of sub-blocks that may be called prediction units (PUs). Transform units (TUs) may also be utilized.
[0051] Still referring to FIG. 6, the exemplary video encoder 600 may include an intra prediction processor 612, a motion estimation / compensation processor 612 (also referred to as an inter prediction processor) that can construct a motion vector candidate list, including the step of adding a single global motion vector candidate to the motion vector candidate list, a transform / quantization processor 616, an inverse quantization / inverse transform processor 620, an in-loop filter 624, a decoded picture buffer 628, and / or an entropy coding processor 632. Bitstream parameters may be input to the entropy coding processor 632 for inclusion in the output bitstream 636.
[0052] During operation, referring continuously to FIG. 6, for each block of frames of the input video 604, it may be determined whether to process the block via intra-picture prediction or using motion estimation / compensation. The block may be provided to the intra prediction processor 608 or the motion estimation / compensation processor 612. If the block is to be processed via intra prediction, the intra prediction processor 608 may perform the processing and output a predictor. If the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 612 may perform the processing, including steps of constructing a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list if applicable.
[0053] Still referring to FIG. 6, the residual may be formed by subtracting the predictor from the input video. The residual may be received by the transform / quantization processor 616, which may perform a transform process (e.g., discrete cosine transform (DCT)) and produce coefficients that can be quantized. The quantized coefficients and any associated signaling information may be provided to the entropy coding processor 632 for entropy coding and inclusion within the output bitstream 636. The entropy coding processor 632 may assist in encoding the signaling information associated with the step of encoding the current block. Additionally, the quantized coefficients may be provided to the inverse quantization / inverse transform processor 620, which may reproduce pixels that can be processed by the in-loop filter 624 in combination with the predictor, and its output may be stored in the decoded picture buffer 628 for use by the motion estimation / compensation processor 612, which is capable of constructing a motion vector candidate list, including adding a single global motion vector candidate to the motion vector candidate list.
[0054] Referring further to FIG. 6, while several variations have been described in detail above, other modifications or additions are also conceivable. For example, in some implementations, the current block may include any symmetric block (8×8, 16×16, 32×32, 64×64, 128×128, and equivalents) and any asymmetric block (8×4, 16×8, and equivalents).
[0055] In some implementations, still referring to FIG. 6, a quadtree + binary decision tree (QTBT) may be implemented. In QTBT, at the coding tree unit level, the QTBT partition parameter may be dynamically derived to conform to local characteristics without transmitting any overhead. Subsequently, at the coding unit level, a joint classifier decision tree structure may eliminate unnecessary iterations and control the risk of incorrect predictions. In some implementations, an LTR frame block update mode may be available as an additional option for each leaf node of the QTBT.
[0056] In some implementations, still referring to FIG. 6, additional syntax elements may be signaled at different hierarchical levels of the bitstream. For example, a flag may be enabled over the entire sequence by including an enabled flag coded within the sequence parameter set (SPS). Further, a CTU flag may be coded at the coding tree unit (CTU) level.
[0057] One or more of the aspects and embodiments described herein may be readily implemented and / or realized within one or more machines (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices such as document servers, etc.) programmed in accordance with the teachings herein, such that they would be apparent to one of ordinary skill in the computer art. It should be noted that these various aspects or features may be readily implemented using digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include an implementation within one or more computer programs and / or software that are executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive and transmit data and instructions from and to a storage system, at least one input device, and at least one output device, which may be of special purpose or general purpose. Appropriate software coding can be readily prepared by a skilled programmer based on the teachings of the present disclosure, as would be apparent to one of ordinary skill in the software art. Aspects and implementations discussed above that employ software and / or software modules may also include appropriate hardware to assist in the implementation of machine-executable instructions of the software and / or software modules.
[0058] Such software may be a computer program product that employs a machine-readable storage medium. The machine-readable storage medium can store and / or encode a sequence of instructions for execution by a machine (e.g., a computing device), and can be any medium that causes the machine to perform any one of the methodologies and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CD, CD-R, DVD, DVD-R, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid state memory devices, EPROM, EEPROM, programmable logic devices (PLD), and / or any combination thereof. As used herein, a machine-readable medium includes a single medium and a collection of physically distinct media, such as a collection of compact discs or one or more hard disk drives in combination with computer memory. As used herein, a machine-readable storage medium does not include transitory forms of signal transmissions.
[0059] Such software may also include information (e.g., data) carried as a data signal on a data carrier such as a carrier wave. For example, machine-executable information can include a data carrier signal embodied in a data carrier that encodes a sequence of instructions or a portion thereof for execution by a machine (e.g., a computing device), and any associated information (e.g., data structures and data) that causes the machine to perform any one of the methodologies and / or embodiments described herein.
[0060] Examples of computing devices include, but are not limited to, e - book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smartphones, etc.), web devices, network routers, network switches, network bridges, and any machine capable of executing a sequence of instructions that specify actions to be taken by that machine and any combination thereof. In one embodiment, the computing device may include and / or be included within a kiosk.
[0061] FIG. 7 shows a graphical representation of one embodiment of a computing device in an exemplary form of a computer system 700 in which a set of instructions for causing the control system to implement any one or more of the aspects and / or methodologies of the present disclosure can be executed. It is also envisioned that multiple computing devices may be utilized to implement a specially configured set of instructions for causing any one or more of the devices to implement any one or more of the aspects and / or methodologies of the present disclosure. The computer system 700 includes a processor 704 and a memory 708 that communicate with each other and with other components via a bus 712. The bus 712 may include any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures.
[0062] Memory 708 may include various components (e.g., machine-readable media), including but not limited to, random access memory components, read-only components, and any combination thereof. In one embodiment, a basic input / output system 716 (BIOS), including basic routines that help transfer information between elements within computer system 700 during startup and the like, may be stored in memory 708. Memory 708 may also include instructions (e.g., software) 720 that embody any one or more of the aspects and / or methodologies of the present disclosure (e.g., stored on one or more machine-readable media). In another embodiment, memory 708 may further include any number of program modules, including but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0063] Computer system 700 may also include a memory device 724. Examples of memory devices (e.g., memory device 724) include, but are not limited to, hard disk drives, magnetic disk drives, optical disk drives in combination with optical media, solid state memory devices, and any combination thereof. Memory device 724 may be connected to bus 712 by an appropriate interface (not shown). Exemplary interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE (registered trademark)), and any combination thereof. In one embodiment, memory device 724 (or one or more of its components) may be removably interfaced with computer system 700 (e.g., via an external port connector (not shown)). In particular, memory device 724 and associated machine-readable medium 728 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 700. In one embodiment, software 720 may reside, in whole or in part, within machine-readable medium 728. In another embodiment, software 720 may reside, in whole or in part, within processor 704.
[0064] Computer system 700 may also include an input device 732. In one embodiment, a user of computer system 700 may enter commands and / or other information into computer system 700 via input device 732. Examples of input device 732 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, game pads, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touch pads, optical scanners, video capture devices (e.g., still cameras, video cameras), touch screens, and any combination thereof. Input device 732 may interface with bus 712 via any of a variety of interfaces (not shown), including, but not limited to, serial interfaces, parallel interfaces, game ports, USB interfaces, FIREWIRE (registered trademark) interfaces, direct interfaces to bus 712, and any combination thereof. Input device 732 may further include a touch screen interface, which may be part of, or separate from, display 736, discussed further below. Input device 732 may be utilized as a user selection device for selecting one or more graphical representations within a graphical interface as described above.
[0065] The user may also input commands and / or other information into computer system 700 via a memory device 724 (e.g., removable disk drive, flash drive, etc.) and / or a network interface device 740. A network interface device, such as network interface device 740, may be utilized to connect computer system 700 to one or more of various networks, such as network 744, and one or more remote devices 748 connected thereto. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, data networks associated with telephone / voice providers (e.g., data and / or voice networks of mobile communication providers), direct connections between two computing devices, and any combination thereof. Networks, such as network 744, may employ wired and / or wireless modes of communication. Generally, any network topology may be used. Information (e.g., data, software 720, etc.) may be communicated to and / or from computer system 700 via network interface device 740.
[0066] The computer system 700 may further include a video display adapter 752 for communicating an image viewable on a display device, such as the display device 736. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. The display adapter 752 and the display device 736 may be utilized in combination with the processor 704 to provide a graphical representation of aspects of the present disclosure. In addition to the display device, the computer system 700 may include one or more other peripheral output devices including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to the bus 712 via a peripheral interface 756. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, FIREWIRE® connections, parallel connections, and any combination thereof.
[0067] The foregoing is a detailed description of illustrative embodiments of the invention. Various modifications and additions can be made without departing from the spirit and scope of the invention. Each feature of the various embodiments described above can be combined, as appropriate, with features of other described embodiments to provide combinations of multiple features in associated new embodiments. Further, while the foregoing describes several separate embodiments, what is described herein is merely illustrative of the application of the principles of the invention. In addition, the specific methods described herein may be illustrated and / or described as being performed in a particular order, but the order is highly variable among those of ordinary skill in the art to achieve embodiments as disclosed herein. Accordingly, this description is intended to be regarded only as an example and is not intended to limit the scope of the invention in any other way.
[0068] In the foregoing description, and in the claims, phrases such as "at least one of" or "one or more of" may occur, followed by a conjunctive listing of elements or features. The term "and / or" may also occur within a listing of two or more elements or features. Unless otherwise implied or explicitly contradicted by the context in which such phrases are used, this is intended to mean any of the individually recited elements or features, or any combination of any of the recited elements or features with any other recited element or feature. For example, the phrases "at least one of A and B", "one or more of A and B", and "A and / or B" are each intended to mean "A alone, B alone, or A and B together". Similar interpretations are also intended with respect to listings that include three or more items. For example, the phrases "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together". Additionally, the use of the term "based on" in the foregoing and in the claims is intended to mean "at least based on" such that features or elements not recited are also allowable.
[0069] The subject matter described in this specification can be embodied in a system, apparatus, method, and / or article, depending on the desired configuration. The implementations described in the foregoing description do not represent all implementations consistent with the subject matter described in this specification. Instead, they are only some examples consistent with aspects related to the described subject matter. Some variations have been described in detail above, but other modifications or additions are also possible. In particular, further features and / or variations can be provided in addition to those described herein. For example, the implementations described above can be directed to various combinations and sub-combinations of the disclosed features and / or combinations and sub-combinations of some additional features disclosed above. Additionally, the logical flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order or sequential order shown to achieve the desired result. Other implementations may be within the scope of the following claims.
Claims
1. A video encoder for encoding a bitstream for decoding by a video decoder, the bitstream including a sequence parameter set and a coded picture, the coded picture including a first continuous region including a first plurality of coding blocks and a second continuous region including a second plurality of coding blocks, the first continuous region including global motion, the second continuous region including local motion, the video decoder a circuit network, the circuit network receiving the bitstream, for each coding block of the first plurality of coding blocks in the first continuous region, determining a global motion model, the global motion model being common to all of the first plurality of coding blocks in the first continuous region, the global motion model being one of a translational motion, a four-parameter affine motion, or a six-parameter affine motion, the global motion model being enabled in the sequence parameter set, when the global motion model is a translational motion, for each of the first plurality of coding blocks in the first continuous region, constructing a prediction candidate list including a first candidate, the first candidate being a spatial motion vector candidate in the coded picture, decoding each of the first plurality of coding blocks in the first continuous region using the first candidate for translational motion compensation to perform, when the global motion model is a four-parameter affine motion, for each of the first plurality of coding blocks in the first continuous region, constructing a prediction candidate list including a second candidate including two control point motion vectors, each of the two control point motion vectors being a spatial motion vector candidate in the coded picture, decoding each of the first plurality of coding blocks in the first continuous region using the second candidate for four-parameter affine motion compensation to perform, and, when the global motion model is a six-parameter affine motion, Constructing a prediction candidate list including a third candidate containing three control point motion vectors for each of the first plurality of coding blocks within the first continuous region, wherein each of the three control point motion vectors is a spatial motion vector candidate within the coded picture, and Decoding each of the first plurality of coding blocks within the first continuous region using the third candidate for six-parameter affine motion compensation Performing By decoding the first continuous region of the coded picture, reconstructing the global motion, and By decoding each of the second plurality of coding blocks within the second continuous region using the individual motion information included in the bitstream for each of the second plurality of coding blocks, decoding the second continuous region of the coded picture to reconstruct the local motion A circuit network configured to perform A video encoder comprising
2. The video encoder according to claim 1, wherein all of the first plurality of coding blocks within the first continuous region are 64×64 or all are 128×128.
3. The video encoder according to claim 1, wherein the global motion within the first continuous region is caused by camera motion.
4. The video encoder according to claim 1, wherein the local motion within the second continuous region is caused by the motion of an object within the scene.
5. The video encoder according to claim 1, wherein the first continuous region and the second continuous region constitute the entire coded picture.
6. The video encoder according to claim 1, wherein the first continuous region has more coding blocks than the second continuous region.
7. A video encoder for encoding a bitstream for decoding by a video decoder, wherein the encoded bitstream includes a sequence parameter set and a coded picture, the coded picture including a first continuous region including a first plurality of coding blocks and a second continuous region including a second plurality of coding blocks, the first continuous region including a global motion caused by the motion of a camera, the second continuous region including a local motion caused by the motion of an object, and the video decoder is a network of circuits, and the network of circuits receives the bitstream, for each coding block of the first plurality of coding blocks in the first continuous region, determining a global motion model, the global motion model being common to all of the first plurality of coding blocks in the first continuous region, the global motion model being one of a translational motion, a four-parameter affine motion, or a six-parameter affine motion, and the global motion model being enabled in the sequence parameter set. When the global motion model is a translational motion, for each of the first plurality of coding blocks in the first continuous region, constructing a prediction candidate list including a first candidate, the first candidate being a spatial motion vector candidate in the coded picture. Decoding each of the first plurality of coding blocks in the first continuous region using the first candidate for translational motion compensation. Performing When the global motion model is a four-parameter affine motion, for each of the first plurality of coding blocks in the first continuous region, constructing a prediction candidate list including a second candidate including two control point motion vectors, each of the two control point motion vectors being a spatial motion vector candidate in the coded picture. Decoding each of the first plurality of coding blocks in the first continuous region using the second candidate for four-parameter affine motion compensation. performing, and when the global motion model is a six-parameter affine motion, constructing a prediction candidate list including a third candidate including three control point motion vectors for each of the first plurality of coding blocks in the first continuous region, each of the three control point motion vectors being a spatial motion vector candidate in the coded picture, decoding each of the first plurality of coding blocks in the first continuous region using the third candidate for six-parameter affine motion compensation performing thereby reconstructing the global motion by decoding the first continuous region of the coded picture, reconstructing the local motion by decoding the second continuous region of the coded picture using the individual motion information included in the bitstream for each of the second plurality of coding blocks in the second continuous region a circuit network configured to perform A video encoder comprising.
Citation Information
Patent Citations
Method and apparatus for encoding and decoding moving image
JP2010011075A
Motion vector prediction
US20180359483A1
Image processing device and method
WO2014013880A1
Method and apparatus for global motion compensation in video coding system
WO2017087751A1
Image encoding / decoding method and device, and recording medium having bitstream stored thereon
WO2018097589A1