Implicit derivation of MVD sign and parity for video coding

By implicitly deriving the sign of the last non-zero MVD component using parity calculations based on the magnitudes of other components, the method reduces the bitstream size and enhances coding efficiency in video coding techniques.

WO2025106217A1PCT designated stage expired Publication Date: 2025-05-22GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/052083
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2024-10-18
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing video coding techniques explicitly encode the sign of each motion vector difference (MVD) component, which increases the bitstream size and reduces coding efficiency.

Method used

The method involves implicitly deriving the sign of the last non-zero MVD component based on the magnitudes of the other MVD components, using parity associated with the sum of their absolute values, thereby reducing the number of bits needed to represent the MVD components.

Benefits of technology

This approach reduces the size of the compressed bitstream by eliminating the need to explicitly encode the sign of the last non-zero MVD component, while maintaining decoding accuracy through implicit derivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024052083_22052025_PF_FP_ABST
    Figure US2024052083_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Decoding motion vector difference (MVD) components for a current block is disclosed. Respective magnitudes of the MVD components are decoded. Respective signs of each of the non-zero MVD components, excluding the last non-zero component in the sequence, are decoded. The sign of the last non-zero MVD component is derived based on the magnitudes of the decoded MVD components.
Need to check novelty before this filing date? Find Prior Art

Description

IMPLICIT DERIVATION OF MVD SIGN AND PARITY FOR VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 548,446, filed November 14, 2023, the entire disclosure of which is incorporated herein by reference.BACKGROUND

[0002] Digital video streams may represent video using a sequence of frames or still images. Digital video can be used for various applications including, for example, video conferencing, high-definition video entertainment, video advertisements, or sharing of usergenerated videos. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various approaches have been proposed to reduce the amount of data in video streams, including compression and other coding techniques. These techniques may include both lossy and lossless coding techniques.SUMMARY

[0003] This disclosure relates generally to encoding and decoding video data and more particularly relates to deriving, instead of coding, certain data relating to motion vectors.

[0004] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.

[0005] One general aspect includes a method for decoding a sequence of motion vector difference (MVD) components for a current block. The method includes decoding respective magnitudes of the MVD components; decoding respective signs of each non-zero MVD component except for a last non-zero MVD component in the sequence of MVD components; and deriving a sign of the last non-zero MVD component based on the respective magnitudes of the MVD components.

[0006] Another general aspect includes a method for encoding motion vector difference (MVD) components for a current block. The method includes determining whether a condition is met, the condition including that a sign of a last non-zero MVD component in a sequence is equivalent to a parity of a sum obtained from the MVD components; in response to determining that the condition is not met, modifying at least one of the MVD components to meet the condition; encoding respective magnitudes of the MVD components; and encoding respective signs of all of the MVD components except for the sign of the last nonzero MVD component.

[0007] Another general aspect includes a method for decoding motion vector difference (MVD) components for a current block. The method includes decoding respective magnitudes and signs for each of the MVD components except for a last MVD component of the MVD components; decoding a half value of the last MVD component; deriving an absolute value of the last MVD component based on a parity derived from a sum obtained based on the magnitudes of the each of the MVD components except the last MVD component; and decoding the sign of the last MVD component.

[0008] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects may be implemented by appropriate computer programs which may be carried on appropriate carrier media which may be tangible carrier media (e.g. disks) or intangible carrier media (e.g. communications signals). Aspects may also be implemented using suitable apparatus which may take the form of programmable computers running computer programs arranged to implement the methods and / or techniques disclosed herein. Aspects can be combined such that features described in the context of one aspect may be implemented in another aspect.

[0009] These and other aspects of the present disclosure are disclosed in the following detailed description of the embodiments, the appended claims, and the accompanying figures.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The description herein refers to the accompanying drawings described below wherein like reference numerals refer to like parts throughout the several views.

[0011] FIG. 1 is a schematic of a video encoding and decoding system.

[0012] FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.

[0013] FIG. 3 is a diagram of an example of a video stream to be encoded and subsequently decoded.

[0014] FIG. 4 is a block diagram of an encoder.

[0015] FIG. 5 is a block diagram of a decoder.

[0016] FIG. 6 is a diagram of motion vectors representing full and sub-pixel motion.

[0017] FIG. 7 is a diagram of a sub-pixel prediction block.

[0018] FIG. 8 is a diagram of full and sub-pixel positions.

[0019] FIG. 9 is an example of a flowchart of a technique for inferring a sign of an MVD component.

[0020] FIG. 10 is a flowchart of an example of a technique for deriving a sign for a last non-zero component of MVD components based on their respective magnitudes.

[0021] FIG. 11 is a flowchart of an example of a technique for deriving a sign for a last non-zero component of MVD components based on the respective magnitudes.

[0022] FIG. 12 is an example of a flowchart of a technique for efficiently decoding a magnitude of an MVD component.

[0023] FIG. 13 is a flowchart of an example of a technique for implicit derivation of motion vector signs.

[0024] FIG. 14 is a flowchart of an example of a technique for encoding signs of motion vectors.

[0025] FIG. 15 is a flowchart of an example of a technique for decoding magnitudes of MVDs.DETAILED DESCRIPTION

[0026] As mentioned, compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream (i.e., an encoded bitstream) using one or more techniques to limit the information included in the output bitstream. A received bitstream can be decoded to re-create the blocks and the source images from the limited information. Encoding a video stream, or a portion thereof, such as a frame or a block, can include using temporal or spatial similarities in the video stream to improve coding efficiency. For example, a current block of a video stream may be encoded based on identifying a difference (residual) between the previously coded pixel values, or between a combination of previously coded pixel values, and those in the current block.

[0027] Encoding using temporal similarities is known as inter prediction or motion- compensated prediction (MCP). A prediction block of a current block (i.e., a block being coded) is generated by finding a corresponding block in a reference frame following a motion vector (MV). That is, inter prediction attempts to predict the pixel values of a block using apossibly displaced block or blocks from a temporally nearby frame (i.e., a reference frame) or frames. A temporally nearby frame is a frame that appears earlier or later in time in the video stream than the frame (i.e., the current frame) of the block being encoded (i.e., the current block). A motion vector used to generate a prediction block refers to (e.g., points to or is used in conjunction with) a frame (i.e., a reference frame) other than the current frame. A motion vector may be defined to represent a block or pixel offset between the reference frame and the corresponding block or pixels of the current frame.

[0028] The motion vector(s) for a current block in motion-compensated prediction may be encoded into, and decoded from, a compressed bitstream. A motion vector for a current block (i.e., a block being encoded) is described with respect to a co-located block in a reference frame. The motion vector describes an offset (i.e., a displacement) in the horizontal direction (i.e., MVX) and a displacement in the vertical direction (i.e., MVy) from the colocated block in the reference frame. As such, an MV can be characterized as a 3-tuple (f, MVX, MVy) where f is indicative of (e.g., is an index of) a reference frame, MVXis the offset in the horizontal direction from a collocated position of the reference frame, and MVyis the offset in the vertical direction from the collocated position of the reference frame. As such, at least the offsets MVXand MVy(referred to herein as MV components) are written (i.e., encoded) into the compressed bitstream and read (i.e., decoded) from the encoded bitstream.

[0029] Some blocks may be coded using one MV (in the case of a single prediction mode) and using two MVs in certain cases of compound prediction. A compound predictor may be formed by combining two prediction blocks, using coding modes such as inter and / or intra prediction. This includes combinations such as intra+intra, intra+inter, or inter+inter compound predictors. The inter+inter compound prediction mode involves using two motion vectors from different reference frames, which could be from the past, future, or a mix thereof, to obtain two prediction blocks that are then combined. The two prediction blocks can be combined as defined by the semantics of the prediction mode.

[0030] To lower the rate cost of encoding the motion vectors, a motion vector may be encoded differentially. Namely, a predicted motion vector (PMV) is selected as a reference motion vector, and only a difference between the motion vector (MV) and the reference motion vector (also called the motion vector difference (MVD)) is encoded into the compressed bitstream. The reference (or predicted) motion vector may be a motion vector of one of the neighboring blocks, for example. Thus, MVD=MV-PMV. The neighboring blocks can include spatial neighboring blocks (i.e., blocks in the same current frame as the current block). The neighboring blocks can include temporal neighboring blocks (i.e., blocks inframes other than the current frame). An encoder codes the MVD in the compressed bitstream; and a decoder decodes the MVD from the compressed bitstream and adds it to prediction motion vector (PMV) to obtain the motion vector (MV) of a current block. That is, at the decoder, MV=MVD+PMV. Even more specifically, MVX=MVDX+PMVXand MVy=MVDy+PMVy.

[0031] Coding an MV as used herein includes coding an MVD for the MV. As such, coding an MV includes coding the horizontal offset (i.e., MVDX) and coding the vertical offset (i.e., MVDy) of the motion vector difference. When implemented by an encoder, “coding” means encoding in a compressed bitstream. When implemented by a decoder, “coding” means decoding from a compressed bitstream.

[0032] Coding an MVD may include or involve separately coding respective magnitudes and signs for the MVD components (i.e., MVDXand MVDy). Conventional techniques explicitly code a respective sign bit for each of the MVDs of a block.

[0033] To illustrate, in one known approach, coding MVD(s) for a block includes signaling a 4-ary symbol, “mv Joint ” which specifies which components of the motion vector difference are non-zero. If both MVDx=MVDy=0, then mv Joint is set to 0; if M VDs0 and MVDy=0, then mv Joint is set to 1 ; if MVDx=0 and M VD 0, then mv Joint is set to 2; and if M VDs0 and M VD 0, then mv Joint is 3. Then, for each non-zero MVD component (either MVDXor MVDy), several additional parameters are signaled. First, the sign of the MVDXor MVDyis indicated, denoted mv_sign. Then, the class of the MVD component, denoted mv_class, is signaled indicating the magnitude of the MVD component. Finally, an offset value may be signaled indicating the difference between the integer pixel MVD component and the base value of the associated mv_class.

[0034] Implementations according to this disclosure can implicitly derive certain data related to an MVD component based on data related to other MVD components. The derived data can be inferred based on data that is said to be hidden in the data of the other MVD components. An encoder performs the data hiding. More specifically, parity associated with certain data of the MVD components is used to derive the other data.

[0035] “Parity” associated with the sum of absolute values of some (e.g. all) MVD components can be used to determine the sign (positive or negative) of the last non-zero MVD component. As such, the sign of the last non-zero MVD component need not be explicitly transmitted in a compressed bitstream therewith reducing the number of bits in the compressed bitstream. In another example, “parity” associated with the sum of absolute values of all but the last of the MVD components can be used to infer the parity of the lastMVD component. As such, only a rounded down half of the value of the last MVD component need be transmitted in the compressed bitstream and, based on the parity, the decoder can determine whether the MVD component should be set to (value*2) or (value*2+l). By only transmitting half of the value of the last MVD component reduces the size of the compressed bitstream.

[0036] While the disclosure herein is mainly described with respect to MVD components, the teachings herein also apply to coding MV components that are not differentially coded. Additionally, the teachings herein can be adapted to use data associated with MVD components to derive at the decoder (and omit encoding at the encoder) other data related to MVD components. Further details of implicit derivation of data related to motion vectors, such as an MVD sign and a parity, are described herein with initial reference to a system in which it can be implemented.

[0037] FIG. 1 is a schematic of a video encoding and decoding system 100. A transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.

[0038] A network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream. Specifically, the video stream can be encoded in the transmitting station 102 and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone network or any other means of transferring the video stream from the transmitting station 102 to, in this example, the receiving station 106.

[0039] The receiving station 106, in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.

[0040] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and / or some communication pathway) the encoded video stream and stores the video stream for laterdecoding. In an example implementation, a real-time transport protocol (RTP) is used for transmission of the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used, e.g., a video streaming protocol based on the Hypertext Transfer Protocol (HTTP).

[0041] When used in a video conferencing system, for example, the transmitting station 102 and / or the receiving station 106 may include the ability to both encode and decode a video stream as described below. For example, the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference server (e.g., the transmitting station 102) to decode and view and further encodes and transmits its own video bitstream to the video conference server for decoding and viewing by other participants.

[0042] FIG. 2 is a block diagram of an example of a computing device 200 (e.g., an apparatus) that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.

[0043] A CPU 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the CPU 202 can be any other type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed. Although the disclosed implementations can be practiced with one processor as shown, e.g., the CPU 202, advantages in speed and efficiency can be achieved using more than one processor.

[0044] A memory 204 in computing device 200 can be a read only memory (ROM) device or a random-access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that is accessed by the CPU 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the CPU 202 to perform the methods described here. For example, the application programs 210 can include applications 1 through N, which further include a video coding application that performs the methods described here.Computing device 200 can also include a secondary storage 214, which can, for example, be a memory card used with a mobile computing device. Because the video communicationsessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memory 204 as needed for processing.

[0045] The computing device 200 can also include one or more output devices, such as a display 218. The display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid crystal display (LCD), a cathode-ray tube (CRT) display or light emitting diode (LED) display, such as an organic LED (OLED) display.

[0046] The computing device 200 can also include or be in communication with an image-sensing device 220, for example a camera, or any other image-sensing device 220 now existing or hereafter developed that can sense an image such as the image of a user operating the computing device 200. The image-sensing device 220 can be positioned such that it is directed toward the user operating the computing device 200. In an example, the position and optical axis of the image-sensing device 220 can be configured such that the field of vision includes an area that is directly adjacent to the display 218 and from which the display 218 is visible.

[0047] The computing device 200 can also include or be in communication with a soundsensing device 222, for example a microphone, or any other sound-sensing device now existing or hereafter developed that can sense sounds near the computing device 200. The sound-sensing device 222 can be positioned such that it is directed toward the user operating the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.

[0048] Although FIG. 2 depicts the CPU 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized. The operations of the CPU 202 can be distributed across multiple machines (wherein individual machines can have one or more processors) that can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines such as a network-based memory or memory in multiple machines performing the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 can be composed of multiple buses. Further, the secondary storage 214 can be directly coupled to the other components of the computing device 200 or can be accessed viaa network and can comprise an integrated unit such as a memory card or multiple units such as multiple memory cards. The computing device 200 can thus be implemented in a wide variety of configurations.

[0049] FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, e.g., a frame 306. At the next level, the frame 306 can be divided into a series of planes or segments 308. The segments 308 can be subsets of frames that permit parallel processing, for example. The segments 308 can also be subsets of frames that can separate the video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.

[0050] Whether or not the frame 306 is divided into segments 308, the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to, for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of pixel data. The blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macro-block are used interchangeably herein.

[0051] FIG. 4 is a block diagram of an encoder 400. The encoder 400 can be implemented, as described above, in the transmitting station 102 such as by providing a computer software program stored in memory, for example, the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4. The encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.

[0052] The encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to reconstruct a frame for encoding of future blocks. In FIG. 4, the encoder 400 has thefollowing stages to perform the various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.

[0053] When the video stream 300 is presented for encoding, respective frames 304, such as the frame 306, can be processed in units of blocks. At the intra / inter prediction stage 402, respective blocks can be encoded using intra-frame prediction (also called intra-prediction) or inter-frame prediction (also called inter-prediction). In any case, a prediction block can be formed. In the case of intra-prediction, a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed. In the case of interprediction, a prediction block may be formed from samples in one or more previously constructed reference frames.

[0054] Next, still referring to FIG. 4, the prediction block can be subtracted from the current block at the intra / inter prediction stage 402 to produce a residual block (also called a residual). The transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms. The quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated. The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy-encoded coefficients, together with other information used to decode the block, which may include for example the type of prediction used, transform type, motion vectors and quantizer value, are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding. The compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.

[0055] The reconstruction path in FIG. 4 (shown by the dotted connection lines) can be used to ensure that the encoder 400 and a decoder 500 (described below) use the same reference frames to decode the compressed bitstream 420. The reconstruction path performs functions that are similar to functions that take place during the decoding process that are discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (alsocalled a derivative residual). At the reconstruction stage 414, the prediction block that was predicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.

[0056] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames. In another implementation, an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.

[0057] FIG. 5 is a block diagram of a decoder 500. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.

[0058] The decoder 500, similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512 and a post loop filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.

[0059] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400. Using header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter prediction stage 508 to create the same prediction block as was created in the encoder 400, e.g., at the intra / inter prediction stage 402. At the reconstruction stage 510, the prediction block can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts.

[0060] Other filtering can be applied to the reconstructed block. In this example, the post loop filtering stage 514 is applied to the reconstructed block to reduce blocking distortion, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein. Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the post loop filtering stage 514.

[0061] FIG. 6 is a diagram of motion vectors representing full and sub-pixel motion. In FIG. 6, several blocks 602, 604, 606, 608 of a current frame 600 are inter predicted using pixels from a reference frame 630. In this example, the reference frame 630 is a reference frame, also called the temporally adjacent frame, in a video sequence including the current frame 600, such as the video stream 300. The reference frame 630 is a reconstructed frame (i.e., one that has been encoded and decoded such as by the reconstruction path of FIG. 4) that has been stored in a so-called last reference frame buffer and is available for coding blocks of the current frame 600. Other (e.g., reconstructed) frames, or portions of such frames may also be available for inter prediction. Other available reference frames may include a golden frame, which is another frame of the video sequence that may be selected (e.g., periodically) according to any number of techniques, and a constructed reference frame, which is a frame that is constructed from one or more other frames of the video sequence but is not shown as part of the decoded output, such as the output video stream 516 of FIG. 5.

[0062] A prediction block 632 for encoding the block 602 corresponds to a motion vector 612. A prediction block 634 for encoding the block 604 corresponds to a motion vector 614. A prediction block 636 for encoding the block 606 corresponds to a motion vector 616. Finally, a prediction block 638 for encoding the block 608 corresponds to a motion vector 618. Each of the blocks 602, 604, 606, 608 is inter predicted using a single motion vector and hence a single reference frame in this example, but the teachings herein also apply to inter prediction using more than one motion vector (such as bi-prediction and / or compound prediction using two different reference frames), where pixels from each prediction are combined in some manner to form a prediction block.

[0063] FIG. 7 is a diagram of a sub-pixel prediction block. In some situations, the prediction block that results in the best (e.g., smallest) residual may not correspond to (e.g., align with) pixels in the reference frame. That is, the best motion vector may point to a location that is between pixels of blocks in the reference frame. In this case, motion compensated prediction at the sub-pixel level is useful. As such, MCP may involve the use ofa sub-pixel interpolation filter that generates filtered sub-pixel values at defined locations between the full pixels (also called integer pixels) along rows, columns, or both. The interpolation filter may be one of a number of interpolation filters available for use in motion compensated prediction.

[0064] FIG. 7 includes the block 632 and neighboring pixels of the block 632 of the reference frame 630 of FIG. 6. Integer pixels within the reference frame 630 are shown as unfilled circles. The integer pixels, in this example, represent reconstructed pixel values of the reference frame 630. The integer pixels are arranged in an array along x- and y-axes. Pixels forming the prediction block 632 are shown as filled circles. The prediction block 632 results from sub-pixel motion along two axes.

[0065] Generating the prediction block 632 can require two interpolation operations. In some cases, generating a prediction block can require only one interpolation operation along one of x and y axes. A first interpolation operation to generate intermediate pixels followed by a second interpolation operation to generate the pixels of the prediction block from the intermediate pixels. The first and the second interpolation operations can be along the horizontal direction (i.e., along the x axis) and the vertical direction (i.e., along the y axis), respectively. Alternatively, the first and the second interpolation operations can be along the vertical direction (i.e., along the y axis) and the horizontal direction (i.e., along the x axis), respectively. The first and second interpolation operations can use a same interpolation filter type. Alternatively, the first and second interpolation operations can use different interpolation filter types.

[0066] In order to produce pixel values for the sub-pixels of the prediction block 632, an interpolation process may be used. In one example, the interpolation process is performed using interpolation filters such as finite impulse response (FIR) filters. An interpolation filter may comprise a 6-tap filter, an 8-tap filter, or other number of taps. The taps of an interpolation filter weight spatially neighboring pixels (integer or sub-pel pixels) with coefficient values to generate a sub-pixel value. In general, the interpolation filters used to generate each sub-pixel value at different sub-pixel positions (e.g., 1 / 2, 1 / 4, 1 / 8, 1 / 16 or other sub-pixel positions) between two pixels are different (i.e., have different coefficient values).

[0067] FIG. 8 is a diagram of full and sub-pixel positions. Different interpolation filters may be available. Each of the interpolation filters may be designed to provide a different frequency response. In an example, the available interpolation filters may include a smooth filter, a normal filter, a sharp filter, and a bilinear filter. The interpolation filter to be used by a decoder to generate a prediction block may be signaled in the header of the framecontaining the block to be predicted. As such, the same interpolation filter is used to generate sub-pixel prediction blocks for all blocks of the frame. The interpolation filter may be signaled at a coding unit level. As such, the same interpolation filter is used for every subblock (e.g., every prediction block) of the coding unit to generate sub-pixel prediction blocks for the sub-blocks of the coding unit.

[0068] An encoder may generate a prediction block based on each of the available interpolation filters. The encoder then selects (i.e., to signal to the decoder) the filter that results in, e.g., the best rate-distortion ratio. A rate-distortion ratio refers to a ratio that balances an amount of distortion (i.e., loss in video quality) with rate (i.e., the number of bits) required for encoding. A coding unit (sometimes called a super-block or a macro-block) may have a size of 128x128 pixels, 64x64 pixels, or some other size, and can be recursively decomposed all the way down to blocks having sizes as small as, in an example, 4x4 pixels.

[0069] In the example of FIG. 8, a 6-tap filter is used. This means that values for the subpixels or pixel positions 820, 822, 824 can be interpolated by applying an interpolation filter to the pixels 800-810. Only sub-pixel positions between the two pixels 804 and 806 are shown in FIG. 8. However, sub-pixel values between the other full pixels of the line of pixels can be determined in a like manner. For example, a sub-pixel value between the two pixels 806 and 808 may be determined or generated by applying an interpolation filter to the pixels 802, 804, 806, 808, 810, and an integer pixel adjacent to the pixel 810, if available.

[0070] Using different coefficient values in an interpolation filter, regardless of its size, results in different characteristics of filtering and hence different compression performance. In some implementations, the set of interpolation filters may be designed for 1 / 16-pixel precision and include at least two of a Bi-linear filter, an 8-tap filter (EIGHTTAP), a sharp 8- tap filter (EIGHTTAP_SHARP), or a smooth 8-tap filter (EIGHTTAP_SMOOTH). Each interpolation filter has a different frequency response.

[0071] FIG. 9 is an example of a flowchart of a technique 900 for inferring a sign of an MVD component. The technique 900 can be used to encode into a compressed bitstream, such as the compressed bitstream 420 of FIG. 4, or decode from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5, respective magnitudes and signs of MVD components for a current block.

[0072] The technique 900 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by aprocessor, such as CPU 202, may cause the computing device to perform the technique 900. The technique 900 may be implemented in whole or in part in the intra / inter prediction stage 508 of the decoder 500 of FIG. 5 or in the intra / inter prediction stage 402 of the encoder of FIG. 4 . The technique 900 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.

[0073] At 902, respective magnitudes of MVD components of one or more MVs of the current block are coded. As already mentioned, the current block may be coded using one or more than one (e.g., two) MVs. For example, if the current block is coded using one MV, then MVD components would include the components MVDXand MVDy; and if the current block is coded using two MVs, then MVD components would include the components MVDi,x, MVDi,y, MVD2,x, and MVD2,y; and so on. Again, when the technique 900 is implemented in an encoder, coding means encoding into the compressed bitstream; and when the technique 900 is implemented in a decoder, coding means decoding from the compressed bitstream.

[0074] Any number of techniques can be used to code the MVs. In one example, for each of the MVs, the technique 900 may code the symbol mv Joint, which is described above, indicating which of the components of the MVD are non-zero. The technique 900 may then code each of the non-zero MVD components. Coding a non-zero MVD component may include coding a class (mv_class) of the MVD component and an offset within the class. The syntax element mv_class categorizes the magnitude of the MVD component into different classes. Each class represents a range of magnitudes. The offset codes the difference between the actual MVD component value and the base value of its class, as determined by the mv_class syntax element. Additional information to represent the sub-pixel precision of the MVD component may also be coded, if necessary.

[0075] At 904, the technique 900 may determine whether sign derivation is appropriate for the current block. If so, then the technique 900 proceeds to 906; otherwise (not shown), the technique 900 proceeds to code the sign of the last non-zero MVD component using conventional techniques (not shown). Sign derivation may not be suitable for all blocks. For example, if the MVDs of the Advanced Motion Vector Difference (AMVD) modes (described below) are very high (e.g., are greater than 1), then sign derivation may not be applied. In such cases, the sign is coded similarly to conventional techniques. If the MVD vector (e.g., the set of MVD components) does not contain at least some MV classes 0 or 1, then the sign of the last non-zero MVD component cannot be derived as described herein.

[0076] As further described blow, at least one MVD component magnitude may be changed by the encoder so that the sum of the magnitudes conforms to (e.g., is consistent with) the sign of the last non-zero MVD component. However, the change cannot be so drastic as to result in significant modification to the MVD component, therewith potentially resulting in significant distortion. To illustrate, assume that the class of an MVD component is assumed to be 4. That is, the MVD component represents at least a 24=16 pixel offset. If the class were to be changed to 3, resulting in an 8-pixel offset, the prediction block and residuals could change significantly.

[0077] The last MVD component of the MVD components that is non-zero is identified. Then, at 906, respective signs of all the MVD components, except for a sign of the last nonzero component, are coded. Stated another way, the signs (positive or negative) for each component of all of the MVD components are coded, except for the last non-zero component, whose sign should not be coded. The sign of an MVD component may be coded using one bit.

[0078] At 908, the sign for the last non-zero MVD component is derived (inferred) based on the respective magnitudes. When implemented at the encoder, deriving the sign of the last non-zero component can mean not encoding the sign in the compressed bitstream; and when implemented at the decoder, deriving the sign of the last non-zero component mean determining the sign bit without explicitly decoding the sign from the compressed bitstream. When implemented at the encoder, deriving the sign of the last non-zero MVD component, at 908, includes determining whether to modify (and actually modifying, if necessary) at least one MVD component.

[0079] As a first example, assume that the current block is coded using two MVs and that the MVs are given by a first MV (-2, 0) and a second MV (1, 0). As such, the MVD components may be [-2, 0, 1, 0]. In this first example, the last non-zero MVD component is 1. The sign (which is positive) of this last non-zero MVD component is not coded. That is, the encoder does not encode a bit (a flag) in the compressed bitstream indicating that the sign of this MVD component is positive; and the decoder derives the sign based on some predefined rules (described below). As a second example, assume that the current block is coded using one MV given by (-5, 0). In this second example, the last non-zero MVD component is -5. The sign (which is negative) of this last non-zero MVD component is not coded. As a third example, assume that the current block is coded using one MV given by (2, 8). In this third example, the last non-zero MVD component is 8. The sign (which is positive) of this last non-zero MVD component is not coded.

[0080] FIG. 10 is a flowchart of an example of a technique 1000 for deriving a sign for a last non-zero component of MVD components based on their respective magnitudes. The technique 1000 is an example of a decoder-based implementation of the step 906 of FIG. 9. The technique 1000 may be performed before or after decoding the sign bits of all MVD components (i.e., all non-zero MVD components) except for the sign of the last non-zero MVD components. As such, the technique 1000 may receive the magnitudes of MVD components for a current block or the signed values of all but the last non-zero MVD component.

[0081] At 1002, a variable sum is initialized to zero. At 1004, a determination is made as to whether more MVD components remain to be processed. If there are no more MVD components to process, the technique 1000 proceeds to 1018; otherwise the technique 1000 proceeds to 1006. At 1006, a variable abs_mvd_comp is set to the absolute value of the next MVD component to be processed.

[0082] The magnitude of an MVD component may be coded in any number of ways, which may be identified based on a prediction mode associated with the current block. Two modes of coding MV and MVD magnitudes are described herein: an enum-based mode and an AMVD mode. However, others are possible. The enum-based and the AMVD modes are now briefly described.

[0083] In the enum-based mode, the encoder and the decoder may support a same set of allowed MV (and therefore MVD) precisions. The set of allowed MV precisions can be an ordered set. The set of allowed MV (and therefore MVD) precisions can be implemented (or represented) using any suitable data structure. In an example, the data structure can be an enumeration given by enum {MV_PRECISION_4_PEL = 0, MV_PRECISION_2_PEL = 1, MV_PRECISION_1_PEL = 2, MV_PRECISION_HALF_PEL = 3, MV_PRECISION_QTR_PEL = 4, MV_PRECISION_EIGHTH_PEL = 5, NUM_MV_PRECISIONS, };

[0084] In the enumeration, MV_PRECISION_4_PEL, MV_PRECISION_2_PEL, MV_PRECISION_1_PEL, MV_PRECISION_HALF_PEL, MV_PRECISION_QTR_PEL,and MV_PRECISION_EIGHTH_PEL define constants, each having the corresponding value (e.g., the constant MV_PRECISION_1_PEL has the integer value 2). The constant NUM_MV_PRECISIONS is the number of allowed MV precisions in the set. While described herein a set of allowed MV precisions that includes six (6) MV precisions, more or fewer MV precisions are possible. For example, the set of allowed MV precisions may include an MV precision MV_PRECISION_SIXTEENTH_PEL corresponding to 1716thfractional pixel precision. In another example, the set of allowed MV precisions may also include MV precision MV_PRECISI0N_8_PEL corresponding to 8-pel precision. For example, the set of allowed MV precisions may not include the MV precision MV_PRECISION_EIGHTH_PEL.

[0085] That the set allowed MV precisions is ordered means that two MV precisions of the set can be compared to determine which is smaller or larger or whether the MV precisions are equal. References to an “MV precision” should be understood from the context to mean one of the MV precision values from the set of allowed MV precisions.

[0086] The MV precision MV_PRECISION_EIGHTH_PEL indicates 1 / 8 MV precision. The MV precision MV_PRECISION_QTR_PEL indicates 1 / 4 MV precision. The MV precision MV_PRECISION_HALF_PEL indicates 1 / 2 MV precision. The MV precisions MV_PRECISION_1_PEL = 2, MV_PRECISION_2_PEL = 1, and MV_PRECISION_4_PEL indicate magnitudes of motion vectors in integer precisions. MV magnitudes are multiples of 2 when the MV precision is MV_PRECISION_2_PEL, and multiples of 4 when the MV precision is MV_PRECISION_4_PEL. For example, if the MV precision of a block is MV_PRECISION_2_PEL, then the supported motion vectors are 0, 2, 4, 6, 8, and so on. Similarly, if the MV precision is equal to MV_PRECISION_4_PEL, then the value of motion vectors are multiple of 4, for example, 0, 4, 8, 12, 16, and so on.

[0087] Conventionally, an MV may be encoded using first bits that indicate an integer (or magnitude) portion of the motion vector and second bits that indicate a fractional portion (i.e., the sub-pixel precision). In an example, a motion vector may be coded using 12 bits, where the 9 most significant bits (MSBs) indicate the integer portion and the 3 least significant bits (LSBs) indicate the fractional portion up to 1 / 8 MV precision. The first LSB indicates whether the precision is either 1 / 8 or not. If the value of first LSB is equal to 1, the precision is 1 / 8; otherwise the precision is not equal to 1 / 8. When the precision is not equal to 1 / 8 ( i.e., the first LSB is equal to 0), then the second LSB indicates whether the precision is 1 / 4 or not. If the first LSB is equal to 0 and the second LSB is equal to 1, then the precision is 1 / 4. If the first LSB is 0 and the second LSB is 0, then the precision can be either integer or1 / 2 . If both of the first and second LSBs are 0, then the third LSB indicates whether the precision is 1 / 2 or integer. Thus, 000, 001, 010, Oi l, 100, 101, 110, 111 may indicate, respectively, integer, 1 / 8, 1 / 4, 1 / 8, 1 / 2, 1 / 8, 1 / 4, and 1 / 8 precision for a particular motion vector. Encoding an MV (e.g., the MV component or the MVD components) using enumbased mode can reduce the number of bits requires to code the MV at least because only the minimum number of bits based on the mode are transmitted. However, when an MV is coded using the enum-based mode, certain bits need not be included in the bitstream. For example, if the MV precision is MV_PRECISION_2_PEL, then there is no need to transmit bits relating to sub-pixel precision as they can only be zero.

[0088] In the AMVD mode, a logarithmic scale is used to encode MVD components. In AMVD, a combination of direct or scaled encoding are used for smaller motion vector differences and logarithmic (specifically log2) encoding for larger differences. The logarithmic scale is beneficial because it allows larger numbers to be represented with fewer bits compared to linear scaling. Essentially, instead of coding the MVD magnitude (e.g., 24) in a bitstream, the log base 2 (e.g., 4) is encoded. More specifically, the floor = [log2(M7£) component magnitude^ of the magnitude (i.e., only the integer part) is coded. An offset is then coded to indicate the difference between the magnitude and the 2floor.

[0089] Referring to FIG. 10 again, at 1008, if the MVD component is coded using the AMVD mode, then the technique 1000 proceeds to 1010; otherwise (i.e., the MVD component is coded using the enum-based mode), the technique 1000 proceeds to 1016.

[0090] At 1010, if the value of the variable abs_mvd_comp is less than a threshold (e.g., 8 in here), half of the abs_mvd_comp is added to the variable sum, at 1012; otherwise, (1 + ,QG2(gbs_mvd_compy) is added to the variable sum, at 1014, which represents a scaled value of the MV component. As mentioned above, in the AMVD mode, the magnitudes can be 0, 2, 4, 6, 8, 16, 32, 64, 128, 256, 512, 1024, and the like. As such, the coded actual MVD component values (e.g., the exponents) are converted to the index vector abs_mvd_comp with values such as 0, 1, 2, 3, 4, 5, 6, 7, 8,9, 10, and 11.

[0091] If the MVD component is determined, at 1008, to be coded using the enum-based mode, then (gbs_mvd_comp » precision_shift) is added to the variable sum, at 1016. The purpose of this shifting is now described.

[0092] The abs_mvd_comp is shifted by precision_shift because some of the right-most bits of the MVD component may have been inferred and are thus zero based on the MV precision. If the block is coded with an MV precision that is lower than l / 8th precision (MV_PRECISION_EIGHTH_PEL), then less than all of the possible bits for coding an MVDare available. To illustrate, if an MVD component is coded as MV_PRECISION_HALF_PEL, then the two least significant bits may not be conveyed in the compressed bitstream and are assumed to be zero. Accordingly, instead of using direct sum of MVD components, parity of the sum of MVD component right shifted by the precision_shift is used to derive the sign. That is, the parity hiding, as performed by the encoder, is based on the available MVD data (e.g., based on data transmitted by the encoder). The precision_shift can be calculated using equation (1), where pb_mv -precision is the MV precision of the current block. precision_shift = MV_PRECISION_EIGHTH_PEL - pb_mv -precision (1)

[0093] To illustrate, if the precision of an MVD component is !4 pel (e.g., pb_mv _precision= MV_PRECISION_QTR_PEL=4), then precision_shift = 5-4 = 1. As such, the value added to sum will be the magnitude of the MVD component shifted by 1. Again, the shifting operation is for the purpose of identifying the information transmitted by the encoder or where the transmitted information begins, which the encoder may have modified. It is noted that precision_shift is equal to 1 for the AMVD mode.

[0094] From 1012, 1014, and 1016, the technique 1000 proceeds back to 1004 to determine whether there are additional MVD components to process. At 1018, the technique 1018 determines whether the value of the accumulated sum is even or odd. If the sum is odd, then the last non-zero MVD component is determined (e.g., inferred) to be negative, at 1020. For example, the value of the decoded last non-zero MVD component may be multiplied by minus one (-1) or a sign bit associated with the last non-zero MVD component may be set to a value (e.g., 1) indicating that it is negative. If the sum is even, the last non-zero MVD component is determined (e.g., inferred) to be positive, at 1022. For example, the sign bit associated with the last non-zero MVD component may be set to a value (e.g., 0) indicating that it is positive.

[0095] To summarize, the technique 1000 is configured based on the assumption that the sign of the last non-zero MVD component is equal to (e.g., is set based on) the parity of the sum of the absolute value of all MVD components, in the case of the enum-based mode, or the sum of log2(MVD components), in the case of the AMVD mode. If the sum is odd, then the sign is assumed to be negative; and if the sum is even, then the sign is assumed to be positive. However, the technique 1000 can be easily adapted to operate on the opposite assumption (e.g., that if the sum is odd, then the sign is positive; otherwise, the sign is negative).

[0096] FIG. 11 is a flowchart of an example of a technique 1100 for deriving a sign for a last non-zero component of MVD components based on the respective magnitudes. Thetechnique 1100 is an example of an encoder-based implementation of the step 906 of FIG. 9. The technique 1100 may modify MVs (e.g., MVD components) so that a condition is satisfied. The condition can be as described above. Namely, the condition can be that the sign of the last non-zero MVD component is equal to the parity of the sum of the absolute value of all MVD components (in the case of the enum-based mode), or the sum of log2(MVD components) (in the case of the AMVD mode). If the sum is odd, then the sign is assumed to be negative; and if the sum is even, then the sign is assumed to be positive.

[0097] At 1102, the technique 1100 determines whether magnitudes of the MVD components conform to) to the sign of the last-non zero MVD component. That is, the technique 1100 determines whether the magnitudes are such that the condition is satisfied. Whether the magnitudes conform to the sign can be determined by first calculating the sum (as described with respect to the variable sum of FIG. 10). If the magnitudes do not conform, then the technique proceeds to 1106; otherwise the technique 1100 proceeds to 1108 to encode the signs of all but the last non-zero MVD component.

[0098] To illustrate, assume that the motion search process of the encoder produces the MVD components [-2, 0, 3, 2]. Here, the sign of the last non zero MVD component (e.g., 2) is positive and the sum of absolute values of the MVD components is ( 2 + 0 + 3 + 2) = 7. The parity bit of 7 is 1 (i.e., the sum is odd). As such, the condition is not satisfied and the technique 1100 determines, at 1104, that the magnitudes do not conform with the sign of the last non-zero MVD component. Therefore, signaling the MVD components [-2, 0, 3, 2], as they are, is not valid.

[0099] As such, at 1106, the technique 1100 modifies at least one of the MVD components based on rate-distortion (RD) costs so that the condition is satisfied. To illustrate, the technique 1100 may modify the MVD components to [-2, 0, 4, 2]. In the updated MVD components, the sign of the last non-zero MVD component is positive and the parity of the sum of absolute values is (2 +0+4+2) == 8, which is even. As such, the sum is even and the sign is positive and the condition is now satisfied.

[0100] In the case of the enum-based mode, modifying at least one of the MVD components can include identifying rate-distortion costs associated with varying at least some (e.g., each) of the MV components by +1 and by -1 and taking the variation that corresponds to the best rate-distortion rate. In another example, the MVD components can be changed by +1 and +3. Other ways of modifying the at least one of the MVD components are possible.

[0101] In the case of the AMVD mode, less information is encoded for MVD components than in the case of enum-based mode. Specifically, the information transmitted isthe exponent of the MVD component (e.g., the mv_class), as described above. As such, there are fewer options for hiding the sign of the last non- zero MVD component. However, as mentioned above, even minor deviations (adjustments) to high mv_class (e.g., to the high exponents or exponents that are greater than 1) can lead to bigger rate-distortion losses. As such, if modification of the MVD components is necessary, then only those MVD components associated with low exponents (0 or 1) are modified. As such, a 0 exponent may be changed to 1 and a 1 exponent may be changed to a 0. If it is not possible to change the exponents (i.e., classes) of at least one of the MVD components, then sign derivation is not appropriate for the block.

[0102] To reiterate, in the case of the AMVD mode, if the sum does not conform to the sign of the last non-zero MVD component then 1) only those MVD components associated with classes 0 and 1 are considered for varying; 2) for each MVD component in the sub-set, the class is changed to 0 if it was 1, and to 1 if it was 0; 3) respective rate-distortion costs of the changes are calculated; and 4) the final MVD components will contain one MVD component whose class was flipped, based on the smallest rate-distortion cost. However, if the MVD components don’t include classes 0 or 1, then sign derivation is skipped for the block.

[0103] FIG. 12 is an example of a flowchart of a technique 1200 for efficiently decoding a magnitude of an MVD component. The technique 1200 can be used to decode from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5, respective magnitudes and signs of MVD components for a current block.

[0104] The technique 1200 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1200. The technique 1200 may be implemented in whole or in part in the intra / inter prediction stage 508 of the decoder 500 of FIG. 5 or in the intra / inter prediction stage 402 of the encoder of FIG. 4 . The technique 1200 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.

[0105] The technique 1200 enables more efficient encoding of the MVD components by reducing the amount of data needed to represent the last MVD component of a set of MVD components. By deriving a parity of the last MVD component based on the sum of the absolute values of the previous components, the encoder can omit explicit coding of this partand only a rounded down half of the MVD component is encoded therewith achieving better compression.

[0106] To illustrate, assume that a current block is encoded using two motion vectors. As such, the MVD components include MVDCO, MVDC1, MVDC2, and MVDC3. The MVD components MVDCO, MVDC1, and MVDC2 can be coded using any conventional technique. Additionally, half of the last MVD component MVDC3 is also coded. The absolute value of the last MVD component (e.g., MVDC3) can be derived based on the sum of the absolute values of the other MVD components. The sum can be calculated as sum = abs(MVDCO) + abs(MVDCl) + abs(MVDC2). The parity bit (i.e., the LSB) of the sum value is extracted. That is, it is determined whether sum is even or odd. The parity bit, denoted p, is 1 if sum is odd and 0 if sum is even. The last MVD component can now be computed. The absolute value of the last MVD component (MVDC3) can be computed as MVDC3 = 2 * half_MVDC3 + p. This formula reconstructs the full value of the last MVD component by doubling half_MVDC3 and adding the derived parity bit. The sign of the last MVD component can be coded using conventional techniques.

[0107] At 1201_1, the respective magnitudes and signs of all MVD components except those of the last MVD component are decoded using any conventional technique. At 1201_2, half of the magnitude of the last MVD component is decoded and stored in the variable half_last_mvd. The half value half_last_mvd can also be decoded using any conventional technique.

[0108] Each of 1202 and 1206-1218 can be similar to 1002 and 1006-1018 of FIG. 10, respectively and descriptions therefor are omitted. Whereas the test at 1004 of FIG. 10 acts as an iterator over all MVD components, the test at 1204 acts as an iterator over all MVD components except the last MVD component. That is, the steps 1206-1216 are not performed with respect to the last MVD component.

[0109] At 1218, if the sim is odd, then at 1220, the parity value p is set to 1; otherwise, at 1222, the parity value p is set to 0. At 1224, the last MVD component can be calculated as last_MVD = 2*half_last_mvd + p. At 1226, the sign of the last MVD component can be decoded.

[0110] At an encoder, and consistent with the description of FIG. 11, if the sum does not conform to the value of the last MVD component, then at least one of the MVD components other than the last MVD component can be changed as described herein. As such, the encoder can, amongst other steps, obtain (e.g., select, determine, etc.) the MVD components (using a conventional technique), modify (if necessary) at least one of the MVD components so thatthe condition is satisfied, encode signs of the MVD components, encode all but the last MVD component, and encode half of the last MVD component. The encoder additionally encodes signs of the MVD components.

[0111] In other implementations according to this disclosure, other syntax elements can be derived instead of, or in addition to, the sign and parity symbols described above. For example, an index of an interpolation filter may be derived. To illustrate, assume that N interpolation filters are possible. The value of a selected filter can be signaled by the encoder using modulo sum of the MVD components. For example, the decoder may derive the interpolation filter by 1) obtaining a sum of MVD components as described above and 2) deriving the interpolation filter as sum % N. Consistent with the above description, the encoder can modify at least some of MVD components so that the sum % N conforms to the filter selected by the encoder.

[0112] To further describe some implementations in greater detail, reference is next made to examples of techniques which may be performed for implicit derivation of data of motion vectors. FIG. 13 is a flowchart of an example of a technique 1300 for implicit derivation of motion vector signs. The technique 1300 can be executed using computing devices, such as the systems, hardware, and software described with respect to FIGS. 1-12. The technique 1300 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1300. The technique 1300 may be implemented in whole or in part in the intra / inter prediction stage 508 of the decoder 500 of FIG. 5. The technique 1300 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.

[0113] The technique 1300 can be used for decoding a sequence of MVD components for a current block. At 1302, respective magnitudes of the MVD components are decoded. In an example, for each non-zero MVD component, a class (e.g., mv_class) and an offset value may be decoded. At 1304, respective signs of each non-zero MVD component except for a last non-zero MVD component in the sequence of MVD components are decoded. At 1306, a sign of the last non-zero MVD component is derived based on the respective magnitudes of the MVD components. The sign of the last non-zero MVD component can be derived based on a parity of a sum of absolute values of all MVD components, such as described with respect to the enum-based mode. As such, the parity can be determined based on a precisionshift related to a motion vector precision of the current block. In another example, the sign of the last non-zero MVD component can be derived based on a parity of a sum of respective classes of the MVD components, as described with respect to the AMVD mode.

[0114] FIG. 14 is a flowchart of an example of a technique 1400 for encoding signs of motion vectors. The technique 1400 can be executed using computing devices, such as the systems, hardware, and software described with respect to FIGS. 1-12. The technique 1400 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1400. The technique 1400 may be implemented in whole or in part in the intra / inter prediction stage 402 of the encoder 400 of FIG. 4. The technique 1400 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.

[0115] Given a sequence of MVD components, regardless of how determined, the technique 1400 encodes all but the sign of the last non-zero MVD component is the sequence in a compressed bitstream. At 1402, the technique 1400 determines whether a condition relating to the MVD components is met. As described above, the condition can be or include that a sign of a last non-zero MVD component in the sequence is equivalent to a parity of a sum obtained from the MVD components. The sum can be a sum of absolute values of all MVD components, as described above with respect to the enum-based mode. As such, the sum of the absolute values can be adjusted by a precision shift related to a motion vector precision of the current block. In another example, the sum can be a sum of logarithmic values in the case of the Advanced Motion Vector Difference (AMVD) mode. As such, the sum can be based on log2 values of the MVD components.

[0116] At 1404, in response to determining that the condition is not met, at least one of the MVD components is modified so that the condition is met. In the case of the AMVD mode, at least one of the MVD components that is to be modified corresponds to an exponent of 0 or 1. Modifying at least one of the MVD components can include altering individual MVD values so that the condition is met. At 1406, respective magnitudes of the MVD components are encoded in a compressed bitstream. At 1408, respective signs of all of the MVD components except for the sign of the last non-zero MVD component are encoded in the compressed bitstream.

[0117] FIG. 15 is a flowchart of an example of a technique 1500 for decoding magnitudes of MVDs. The technique 1500 can be executed using computing devices, such as the systems, hardware, and software described with respect to FIGS. 1-12. The technique 1500 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1500. The technique 1500 may be implemented in whole or in part in the intra / inter prediction stage 508 of the decoder 500 of FIG. 5. The technique 1500 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.

[0118] At 1502, respective magnitudes and signs for each of the MVD components except for a last MVD component of the MVD components are decoded from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5. At 1504, a half value of the last MVD component is decoded from the compressed bitstream. That is, a value that is decoded from the compressed bitstream is equal to half of the value of the last MVD component.

[0119] At 1506, an absolute value of the last MVD component is derived based on a parity derived from a sum obtained based on the magnitudes of the each of the MVD components except the last MVD component. The absolute value of the last MVD can then be computed as twice the half value of the last MVD plus the derived parity. At 1508, the sign of the last MVD component can be decoded.

[0120] The parity can be derived by computing a sum of absolute values of all except the last MVD, and extracting a least significant bit (LSB) of the sum as the parity bit. In an example, and as described above, if the prediction mode is Adaptive Motion Vector Difference (AMVD), the sum can be adjusted based on a modified calculation for MVD components with absolute values less than 8. In an example, and as described with respect to the enum-based mode, the sum can include a precision shift adjustment related to a motion vector (MV) precision of the current block.

[0121] For simplicity of explanation, the techniques described herein, such as the techniques 900-1500 of FIGS. 9-15, respectively, are depicted and described as respective series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustratedsteps or operations may be required to implement a method in accordance with the disclosed subject matter.

[0122] Some implementations are described below as numbered examples (Example 1, 2, 3, etc.). These examples are provided as examples only and do not limit the other implementations disclosed herein.

[0123] Example 1 : A method for decoding a sequence of motion vector difference (MVD) components for a current block. The method includes decoding respective magnitudes of the MVD components; decoding respective signs of each non-zero MVD component except for a last non-zero MVD component in the sequence of MVD components; and deriving a sign of the last non-zero MVD component based on the respective magnitudes of the MVD components.

[0124] Example 2: The method of Example 1, where the sign of the last non-zero MVD component is derived based on a parity of a sum of absolute values of all MVD components.

[0125] Example 3: The method of Example 2, where the parity is determined based on a precision shift related to a motion vector precision of the current block.

[0126] Example 4: The method of Example 1, where the sign of the last non-zero MVD component is derived based on a parity of a sum of respective classes of the MVD components.

[0127] Example 5: The method of Example 1, where for the each non-zero MVD component, the method includes decoding a class and an offset value.

[0128] Example 6: A method for encoding motion vector difference (MVD) components for a current block. The method includes determining whether a condition is met, where the condition comprises that a sign of a last non-zero MVD component in a sequence is equivalent to a parity of a sum obtained from the MVD components; in response to determining that the condition is not met, modifying at least one of the MVD components to meet the condition; encoding respective magnitudes of the MVD components; and encoding respective signs of all of the MVD components except for the sign of the last non-zero MVD component.

[0129] Example 7: The method of Example 6, where the sum is a sum of absolute values of all MVD components.

[0130] Example 8: The method of Example 7, where the sum of the absolute values is adjusted by a precision shift related to a motion vector precision of the current block.

[0131] Example 9: The method of Example 6, where the sum is a sum of logarithmic values in an Advanced Motion Vector Difference (AMVD) mode.- l-

[0132] Example 10: The method of Example 9, where the sum is based on log2 values of the MVD components, and where the at least one of the MVD components that is modified corresponds to an exponent of 0 or 1.

[0133] Example 11: The method of Example 6, where modifying at least one of the MVD components involves altering individual MVD values so that the condition is met.

[0134] Example 12: A method for decoding motion vector difference (MVD) components for a current block. The method includes decoding respective magnitudes and signs for each of the MVD components except for a last MVD component of the MVD components; decoding a half value of the last MVD component; deriving an absolute value of the last MVD component based on a parity derived from a sum obtained based on the magnitudes of each of the MVD components except the last MVD component; and decoding the sign of the last MVD component.

[0135] Example 13: The method of Example 12, where the parity is derived by computing a sum of absolute values of all except the last MVD component, and extracting a least significant bit (LSB) of the sum as a parity bit.

[0136] Example 14: The method of Example 12, where if a prediction mode of the current block is Adaptive Motion Vector Difference (AMVD), the sum is adjusted based on a modified calculation for MVD components with absolute values less than 8.

[0137] Example 15: The method of Example 12, where the sum includes a precision shift adjustment related to a motion vector (MV) precision of the current block.

[0138] Example 16: The method of Example 12, where the absolute value of the last MVD is computed as twice the half value of the last MVD plus the derived parity.

[0139] Example 17: A device that includes a processor that is configured to perform the method of any of Examples 1-5 or 12-16.

[0140] Example 18: A device that includes a memory and a processor. The processor is configured to execute instructions stored in the memory to perform the method of any of Examples 1-5 or 12-16.

[0141] Example 19: A non-transitory computer-readable storage medium that includes executable instructions that, when executed by a processor, facilitate performance of operations that perform the method of any of Examples 1-5 or 12-16.

[0142] Example 20: A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, where the encoded bitstream is configured for decoding by the method of any of Examples 1-5 or 12-16.

[0143] Example 21: A device that includes a processor that is configured to perform the method of any of Examples 6-11.

[0144] Example 22: A device that includes a memory and a processor. The processor is configured to execute instructions stored in the memory to perform the method of any of Examples 6-11.

[0145] Example 23: A non-transitory computer-readable storage medium that includes executable instructions that, when executed by a processor, facilitate performance of operations that perform the method of any of Examples 6-11.

[0146] Example 24: A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, where the encoded bitstream is generated by an encoder performing the method of any of Examples 6-11.

[0147] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it is to be understood that encoding and decoding, as those terms are used in the claims, could mean compression, decompression, transformation, or any other processing or change of data.

[0148] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X includes A or B” is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form. Moreover, use of the term “an implementation” or “one implementation” throughout is not intended to mean the same embodiment or implementation unless described as such.

[0149] Implementations of the transmitting station 102 and / or the receiving station 106 (and the algorithms, methods, instructions, etc., stored thereon and / or executed thereby, including by the encoder 400 and the decoder 500) can be realized in hardware, software, or any combination thereof. The hardware can include, for example, computers, intellectual property (IP) cores, application- specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers,servers, microprocessors, digital signal processors or any other suitable circuit. In the claims, the term “processor” should be understood as encompassing any of the foregoing hardware, either singly or in combination. The terms “signal” and “data” are used interchangeably. Further, portions of the transmitting station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.

[0150] Further, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a general-purpose computer or general-purpose processor with a computer program that, when executed, carries out any of the respective methods, algorithms and / or instructions described herein. In addition, or alternatively, for example, a special purpose computer / processor can be utilized which can contain other hardware for carrying out any of the methods, algorithms, or instructions described herein.

[0151] The transmitting station 102 and the receiving station 106 can, for example, be implemented on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server and the receiving station 106 can be implemented on a device separate from the server, such as a hand-held communications device. In this instance, the transmitting station 102 can encode content using an encoder 400 into an encoded video signal and transmit the encoded video signal to the communications device. In turn, the communications device can then decode the encoded video signal using a decoder 500. Alternatively, the communications device can decode content stored locally on the communications device, for example, content that was not transmitted by the transmitting station 102. Other suitable transmitting and receiving implementation schemes are available. For example, the receiving station 106 can be a generally stationary personal computer rather than a portable communications device and / or a device including an encoder 400 may also include a decoder 500.

[0152] Further, all or a portion of implementations of the present disclosure can take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transport the program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or a semiconductor device. Other suitable mediums are also available.

[0153] The above-described embodiments, implementations and aspects have been described to allow easy understanding of the present invention and do not limit the present invention. On the contrary, the invention is intended to cover various modifications andequivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation to encompass all such modifications and equivalent structure as is permitted under the law.

Claims

What is claimed is:

1. A method for decoding a sequence of motion vector difference (MVD) components for a current block, comprising: decoding respective magnitudes of the MVD components; decoding respective signs of each non-zero MVD component except for a last nonzero MVD component in the sequence of MVD components; and deriving a sign of the last non-zero MVD component based on the respective magnitudes of the MVD components.

2. The method of claim 1, wherein the sign of the last non-zero MVD component is derived based on a parity of a sum of absolute values of all MVD components.

3. The method of claim 2, wherein the parity is determined based on a precision shift related to a motion vector precision of the current block.

4. The method of any one of claims 1 to 3, wherein the sign of the last non-zero MVD component is derived based on a parity of a sum of respective classes of the MVD components.

5. The method of any one of claims 1 to 4, wherein for the each non-zero MVD component, the method comprises: decoding a class and an offset value.

6. A method for encoding motion vector difference (MVD) components for a current block, comprising: determining whether a condition is met, wherein the condition comprises that a sign of a last non-zero MVD component in a sequence is equivalent to a parity of a sum obtained from the MVD components; in response to determining that the condition is not met, modifying at least one of the MVD components to meet the condition; encoding respective magnitudes of the MVD components; and encoding respective signs of all of the MVD components except for the sign of the last non-zero MVD component.

7. The method of claim 6, wherein the sum is a sum of absolute values of allMVD components.

8. The method of claim 7, wherein the sum of the absolute values is adjusted by a precision shift related to a motion vector precision of the current block.

9. The method of any one of claims 6 to 8, wherein the sum is a sum of logarithmic values in an Advanced Motion Vector Difference (AMVD) mode.

10. The method of claim 9, wherein the sum is based on log2 values of the MVD components and wherein the at least one of the MVD components that is modified corresponds to an exponent of 0 or 1.

11. The method of any one of claims 6 to 10, wherein modifying the at least one of the MVD components comprises: altering individual MVD values so that the condition is met.

12. A method for decoding motion vector difference (MVD) components for a current block, comprising: decoding respective magnitudes and signs for each of the MVD components except for a last MVD component of the MVD components; decoding a half value of the last MVD component; deriving an absolute value of the last MVD component based on a parity derived from a sum obtained based on the magnitudes of the each of the MVD components except the last MVD component; and decoding the sign of the last MVD component.

13. The method of claim 12, wherein the parity is derived by computing a sum of absolute values of all except the last MVD, and extracting a least significant bit (LSB) of the sum as a parity bit.

14. The method of any one of claims 12 to 13, wherein if a prediction mode of the current block is Adaptive Motion Vector Difference (AMVD), the sum is adjusted based on a modified calculation for MVD components with absolute values less than 8.

15. The method of any one of claims 12 to 14, wherein the sum includes a precision shift adjustment related to a motion vector (MV) precision of the current block.

16. The method of any one of claims 12 to 15, wherein the absolute value of the last MVD is computed as twice the half value of the last MVD plus the derived parity.

17. A device, comprising: a processor that is configured to perform the method of any of claims 1-5 or 12-16.

18. A device, comprising: a memory; and a processor, the processor configured to execute instructions stored in the memory to perform the method of any of claims 1-5 or 12-16.

19. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any of claims 1-5 or 12-16.

20. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is configured for decoding by the method of any of claims 1-5 or 12-16.

21. A device, comprising: a processor that is configured to perform the method of any of claims 6-11.

22. A device, comprising: a memory; and a processor, the processor configured to execute instructions stored in the memory to perform the method of any of claims 6-11.

23. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any of claims 6-11.

24. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is generated by an encoder performing the method of any of claims 6-11.

Citation Information

Patent Citations

  • Motion vector sign bit hiding

    US20130235936A1

  • Motion Vector Difference Coding And Decoding

    US20190141346A1