Adaptive motion vector candidates list
By adaptively determining the number of motion vector candidates based on block-specific criteria, the method addresses the inefficiencies of traditional video coding techniques, achieving improved bit rate efficiency and prediction quality.
Patent Information
- Application Number
- PCT/US2024/057544
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Traditional video coding techniques face inefficiencies due to a fixed and inflexible size of the list of candidate motion vectors, which does not adapt to varying block sizes, reference frame qualities, or prediction modes, leading to suboptimal prediction quality and bit rate usage.
The method adaptively determines the number of motion vector candidates based on various criteria related to the current block, such as reference frame quality, block size, prediction mode, and quantization parameter, allowing for dynamic adjustment of the list size to optimize compression efficiency.
This adaptive approach enhances both bit rate efficiency and prediction quality by optimizing the number of motion vector candidates for each block, aligning with the specific characteristics and requirements of the current block.
Smart Images

Figure US2024057544_05062025_PF_FP_ABST
Abstract
Description
ADAPTIVE MOTION VECTOR CANDIDATES LISTCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 604,500, filed November 30, 2023, the entire disclosure of which is incorporated herein by reference.BACKGROUND
[0002] Digital video streams may represent video using a sequence of frames or still images. Digital video can be used for various applications including, for example, video conferencing, high-definition video entertainment, video advertisements, or sharing of usergenerated videos. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various approaches have been proposed to reduce the amount of data in video streams, including compression and other coding techniques. These techniques may include both lossy and lossless coding techniques.SUMMARY
[0003] This disclosure relates generally to encoding and decoding video data and more particularly relates to motion vector coding candidate signaling.
[0004] An aspect of the disclosed implementations is a method for coding a current block of a current frame. The method includes determining a number of motion vector candidates based on one or more criteria related to the current block; generating the number of the motion vector candidates; adding the motion vector candidates to a list of candidate motion vectors; and coding an index of one of the motion vector candidates, where the index is coded using a minimum number of bits needed for coding the number of the motion vector candidates.
[0005] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. Another embodiment includes a non- transitory computer-readable storage medium having stored thereon an encoded bitstream that is configured for decoding by the operations of the method. Another embodimentincludes a non-transitory computer-readable storage medium having stored thereon an encoded bitstream that is generated by an encoder performing the operations of the method.
[0006] Implementations may include one or more of the following features.
[0007] The method where the number of the motion vector candidates can be based on whether a reference frame of the current frame is a generated reference frame.
[0008] The number of the motion vector candidates can be based on a quality of a reference frame.
[0009] The quality of the reference frame can be based on at least one of a quantization parameter associated with the reference frame or a temporal distance between the reference frame and the current frame.
[0010] The number of the motion vector candidates can be based on a block size of the current block.
[0011] The number of the motion vector candidates can be based on whether the current block is predicted using a compound prediction mode.
[0012] The number of the motion vector candidates can be based on whether a prediction mode of the current block is a unidirectional prediction mode or a bidirectional prediction mode.
[0013] The number of the motion vector candidates can be based on a quantization parameter associated with the current block.
[0014] The number of the motion vector candidates can be limited to a maximum number, where the maximum number is coded in a header of the current frame.
[0015] The number of the motion vector candidates can be based on a diversity metric calculated with respect to possible motion vector candidates.
[0016] The number of the motion vector candidates can be based on a size of a list of candidate motion vectors associated with a neighboring block.
[0017] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects may be implemented by appropriate computer programs which may be carried on appropriate carrier media which may be tangible carrier media (e.g. disks) or intangible carrier media (e.g. communications signals). Aspects may also be implemented using suitable apparatus which may take the form of programmable computers running computer programs arranged to implement the methods and / or techniques disclosed herein. For example, a non-transitory computer-readable storage medium may include executable instructions that, when executed by a processor, facilitate performance of operations operable to cause the processor to carry out any of the methods described herein. Aspects can becombined such that features described in the context of one aspect may be implemented in another aspect.
[0018] These and other aspects of the present disclosure are disclosed in the following detailed description of the embodiments, the appended claims, and the accompanying figures.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The description herein refers to the accompanying drawings described below wherein like reference numerals refer to like parts throughout the several views.
[0020] FIG. 1 is a schematic of a video encoding and decoding system.
[0021] FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0022] FIG. 3 is a diagram of an example of a video stream to be encoded and subsequently decoded.
[0023] FIG. 4 is a block diagram of an encoder.
[0024] FIG. 5 is a block diagram of a decoder.
[0025] FIG. 6 is a diagram of motion vectors representing full and sub-pixel motion.
[0026] FIG. 7 is a block diagram of an example of a reference frame buffer .
[0027] FIG. 8 is an illustration of compound inter prediction.
[0028] FIGS. 9A-9B illustrate examples of tools for generating motion vector candidates.
[0029] FIG. 10 is a flowchart diagram of a method or technique for coding a current block of a current frame.DETAILED DESCRIPTION
[0030] As mentioned, compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream (i.e., a compressed bitstream) using one or more techniques to limit the information included in the compressed bitstream. A received compressed bitstream can be decoded to re-create the blocks and the source images from the limited information. A current block of a video stream may be encoded based on identifying a difference (residual) between previously coded pixel values, or between a combination of previously coded pixel values, and those in the current block. Encoding a video stream, or a portion thereof, such as a frame or a block, can include using temporal similarities in the video stream to improve coding efficiency.
[0031] Encoding using temporal similarities is known as inter prediction or motion- compensated prediction (MCP). A prediction block of a current block (i.e., a block beingcoded) is generated by finding a corresponding block in a reference frame following a motion vector (MV). That is, inter prediction attempts to predict the pixel values of a block using a possibly displaced block or blocks from a temporally nearby frame (i.e., a reference frame) or multiple temporally nearby frames. A temporally nearby frame is a frame that appears earlier or later in time in the video stream than the current frame that includes the current block. An MV used to generate a prediction block refers to (e.g., points to or is used in conjunction with) a frame (i.e., a reference frame) other than the current frame. An MV may be defined to represent a block or pixel offset between the reference frame and the corresponding block or pixels of the current frame.
[0032] MCP can be performed either from a single reference frame or from two reference frames. Inter prediction modes that perform motion compensation from two reference frames may be referred to as compound inter prediction modes (or compound modes, for brevity). In compound modes, two MVs can be signaled to (or may be derived from a list of candidate MVs at) the decoder. For example, the motion vector(s) for a current block may be encoded into, and decoded from, a compressed bitstream. If both reference frames are, in display order, located on the same side of the current frame, the prediction mode may be referred to as a unidirectional prediction mode. If one of the reference frames is in the backward direction and another reference frame is in the forward direction, in the display order, the compound mode may be referred to as bidirectional prediction mode.
[0033] An MV for a current block is described with respect to a co-located block in a reference frame. The motion vector describes an offset (i.e., a displacement) in the horizontal direction (i.e., MVX) and a displacement in the vertical direction (i.e., MVy) from the colocated block in the reference frame. As such, an MV can be characterized as a 3-tuple (f, MVX, MVy) where f is indicative of (e.g., is an index of) a reference frame, MVXis the offset in the horizontal direction from a collocated position of the reference frame, and MVyis the offset in the vertical direction from the collocated position of the reference frame. As such, at least the offsets MVXand MVyare written (i.e., encoded) into the compressed bitstream and read (i.e., decoded) from the encoded bitstream.
[0034] To lower the rate cost of encoding the motion vectors, a motion vector may be encoded differentially. Namely, a predicted motion vector (PMV) may be selected as a reference motion vector, and only a difference (also called the motion vector difference (MVD)), if any, between the motion vector (MV) of a current block and the reference motion vector is encoded into the compressed bitstream. The reference (or predicted) motion vector may be an MV of one of the neighboring blocks, for example, and may be selected from thelist of candidate MVs. Thus, MVD-MV-PMV. The neighboring blocks can include spatial neighboring blocks (i.e., blocks in the same current frame as the current block) and / or can include temporal neighboring blocks (i.e., blocks in frames other than the current frame). An encoder codes the MV or MVD in the compressed bitstream; the encoder may code the PMV (e.g., an indication thereof, such as an index in the list of candidate MVs) in the compressed bitstream; and a decoder decodes the MVD from the compressed bitstream and adds it to the predicted (or reference) motion vector (PMV) to obtain the motion vector (MV) of a current block.
[0035] As mentioned above, coding an MV may include coding the horizontal offset (i.e., MVX) and coding the vertical offset (i.e., MVy) of the MV or coding the horizontal offset (i.e., MVDX) and coding the vertical offset (i.e., MVDy) of the MVD. When implemented by an encoder, “coding” means encoding in a compressed bitstream. When implemented by a decoder, “coding” means decoding from a compressed bitstream.
[0036] As is known and as already mentioned, there is generally a need to construct a list (e.g., at least one list) of candidate MVs and to code an index of a reference MV (i.e., a selected MV) in the list of candidate MVs. To construct the list of candidate MVs for a current block, a predetermined set of search rules are performed to identify the candidate MVs. For example, the candidate MVs can be identified from spatial neighboring blocks within the current frame or temporal neighboring blocks in a reference frame. The encoder encodes the index of a selected MV candidate in a compressed bitstream; and, at the decoder, the list of candidate MVs may be constructed (e.g., generated) according to the same predetermined rules and the index of the selected MV candidate may be decoded from the compressed bitstream.
[0037] Conventionally, the size of the list of candidate MVs is fixed to a predefined number. That is, the number (i.e., quantity) of candidate MVs that are added to the list of candidate MVs is fixed to the predefined number. In an example, the list may generally be fixed to include four candidate MVs. The candidate MVs can be identified from spatial neighboring blocks within the current frame or using temporal motion vector prediction, as described with respect to FIGS. 9A and 9B, respectively. In one example, if less than four candidates are identified, the list may be filled with zero MVs. The predefined number may also be referred to herein as a baseline list size.
[0038] However, such rigidity in the number of candidates is not optimal. That is, always having to generate the predefined number (e.g., 4) of candidates for motion vector prediction may not always be optimal. For example, for certain blocks, generating the predefinednumber (e.g., 4) candidates may be unnecessary, whereas for others, the predefined number may not suffice to achieve the best prediction quality - more candidates could enhance prediction accuracy in some cases.
[0039] Consider a scenario with a 4-candidate list: selecting one of these candidates would require coding a 2-bit index. However, if in practice, 90% of the blocks in a frame predominantly utilize only the first or second candidate from this list, each of these could be coded with just a single bit (e.g., a 0 may indicate the first candidate and a 1 may indicate the second candidate). In such cases, the use of an additional bit for encoding the index becomes superfluous, leading to less efficient compression. On the other hand, when utilizing a high- quality reference frame, it may be beneficial to consider more than four candidate MVs, which can potentially enhance prediction quality.
[0040] As such, a challenge in video coding lies in the static and inflexible nature in the size of the list of candidate MVs, particularly when dealing with varying block sizes, reference frame qualities, prediction modes (e.g., compound, unidirectional prediction or bidirectional), or other criteria. Traditional codecs lack the adaptability to efficiently select a size for the list of candidate motion vectors for different blocks, leading to inefficiencies in bit rate usage and suboptimal prediction quality.
[0041] Implementations according to this disclosure adaptively change the number of candidate MVs generated for a current block. That is, the size of the list of MV candidates can be adaptively set to be lower or higher than the predefined (e.g., baseline) number. The size of the list of MV candidates can be based on one or more criteria related to the current block. The compression rate and / or quality can be improved by adaptively modifying the number of candidate MVs at the block level depending on different criteria (e.g., conditions) related to the blocks themselves. Conditions related to a current block can be used to determine the number of candidates to be generated for the block. The criteria (e.g., conditions) related to a current block broadly include characteristics of the current block itself or characteristics related to coding the current block.
[0042] The criteria related to a current block may include one or more of whether the reference frame is a generated (e.g., synthesized) frame, reference frame quality (including the quantization parameter (QP) of the reference frame, reference frame index, and temporal distance from the current frame), the size of the current block, the prediction mode, QP of the current block, header information, or a combination thereof. The number of MV candidates (i.e., the size of the list of MV candidates) can be optimized at a block level therewith enhancing both the bit rate efficiency and prediction quality.
[0043] To illustrate, and without limitations, where a reference frame is of high quality, the number of candidate MVs can be increased, which can be conducive to leveraging the detailed information available in high-quality reference frames; conversely, for smaller block sizes, a reduced number of candidate motion vectors may be preferable. This reduction aligns with the lower complexity and reduced motion detail typically associated with smaller blocks, thereby optimizing computational efficiency and reducing bit rate while maintaining prediction accuracy.
[0044] Further details of template matching using available peripheral pixels are described herein with initial reference to a system in which it can be implemented. FIG. 1 is a schematic of a video encoding and decoding system 100. A transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.
[0045] A network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream. Specifically, the video stream can be encoded in the transmitting station 102 and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone network or any other means of transferring the video stream from the transmitting station 102 to, in this example, the receiving station 106.
[0046] The receiving station 106, in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0047] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and / or some communication pathway) the encoded video stream and stores the video stream for later decoding. In an example implementation, a real-time transport protocol (RTP) is used for transmission of the encoded video over the network 104. In another implementation, atransport protocol other than RTP may be used, e.g., a Hypertext Transfer Protocol (HTTP) video streaming protocol.
[0048] When used in a video conferencing system, for example, the transmitting station 102 and / or the receiving station 106 may include the ability to both encode and decode a video stream as described below. For example, the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference server (e.g., the transmitting station 102) to decode and view and further encodes and transmits its own video bitstream to the video conference server for decoding and viewing by other participants.
[0049] FIG. 2 is a block diagram of an example of a computing device 200 (e.g., an apparatus) that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
[0050] A CPU 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the CPU 202 can be any other type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed. Although the disclosed implementations can be practiced with one processor as shown, e.g., the CPU 202, advantages in speed and efficiency can be achieved using more than one processor.
[0051] A memory 204 in computing device 200 can be a read only memory (ROM) device or a random-access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that is accessed by the CPU 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the CPU 202 to perform the methods described here. For example, the application programs 210 can include applications 1 through N, which further include a video coding application that performs the methods described here.Computing device 200 can also include a secondary storage 214, which can, for example, be a memory card used with a mobile computing device. Because the video communication sessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memory 204 as needed for processing.
[0052] The computing device 200 can also include one or more output devices, such as a display 218. The display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid crystal display (LCD), a cathode-ray tube (CRT) display or light emitting diode (LED) display, such as an organic LED (OLED) display.
[0053] The computing device 200 can also include or be in communication with an image-sensing device 220, for example a camera, or any other image-sensing device 220 now existing or hereafter developed that can sense an image such as the image of a user operating the computing device 200. The image-sensing device 220 can be positioned such that it is directed toward the user operating the computing device 200. In an example, the position and optical axis of the image-sensing device 220 can be configured such that the field of vision includes an area that is directly adjacent to the display 218 and from which the display 218 is visible.
[0054] The computing device 200 can also include or be in communication with a soundsensing device 222, for example a microphone, or any other sound-sensing device now existing or hereafter developed that can sense sounds near the computing device 200. The sound- sensing device 222 can be positioned such that it is directed toward the user operating the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.
[0055] Although FIG. 2 depicts the CPU 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized. The operations of the CPU 202 can be distributed across multiple machines (wherein individual machines can have one or more processors) that can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines such as a network-based memory or memory in multiple machines performing the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 can be composed of multiple buses. Further, the secondary storage 214 can be directly coupled to the other components of the computing device 200 or can be accessed via a network and can comprise an integrated unit such as a memory card or multiple units suchas multiple memory cards. The computing device 200 can thus be implemented in a wide variety of configurations.
[0056] FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, e.g., a frame 306. At the next level, the frame 306 can be divided into a series of planes or segments 308. The segments 308 can be subsets of frames that permit parallel processing, for example. The segments 308 can also be subsets of frames that can separate the video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.
[0057] Whether or not the frame 306 is divided into segments 308, the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to, for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of pixel data. The blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macro-block are used interchangeably herein.
[0058] FIG. 4 is a block diagram of an encoder 400. The encoder 400 can be implemented, as described above, in the transmitting station 102 such as by providing a computer software program stored in memory, for example, the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4. The encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.
[0059] The encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to reconstruct a frame for encoding of future blocks. In FIG. 4, the encoder 400 has the following stages to perform the various functions in the reconstruction path: a dequantizationstage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0060] When the video stream 300 is presented for encoding, respective frames 304, such as the frame 306, can be processed in units of blocks. At the intra / inter prediction stage 402, respective blocks can be encoded using intra-frame prediction (also called intra-prediction) or inter-frame prediction (also called inter prediction). In any case, a prediction block can be formed. In the case of intra-prediction, a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed. In the case of inter prediction, a prediction block may be formed from samples in one or more previously constructed reference frames.
[0061] Next, still referring to FIG. 4, the prediction block can be subtracted from the current block at the intra / inter prediction stage 402 to produce a residual block (also called a residual). The transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms. The quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated. The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy-encoded coefficients, together with other information used to decode the block, which may include for example the type of prediction used, transform type, motion vectors and quantizer value, are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding. The compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.
[0062] The reconstruction path in FIG. 4 (shown by the dotted connection lines) can be used to ensure that the encoder 400 and a decoder 500 (described below) use the same reference frames to decode the compressed bitstream 420. The reconstruction path performs functions that are similar to functions that take place during the decoding process that are discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also called a derivative residual). At the reconstruction stage 414, the prediction block that was-lipredicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.
[0063] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames. In another implementation, an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.
[0064] FIG. 5 is a block diagram of a decoder 500. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.
[0065] The decoder 500, similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512 and a post-loop filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
[0066] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400. Using header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter prediction stage 508 to create the same prediction block as was created in the encoder 400, e.g., at the intra / inter prediction stage 402. At the reconstruction stage 510, the prediction block can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be applied to the reconstructed block to reduce blocking artifacts.
[0067] Other filtering can be applied to the reconstructed block. In this example, the postloop filtering stage 514 is applied to the reconstructed block to reduce blocking distortion, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein. Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the post- loop filtering stage 514.
[0068] FIG. 6 is a diagram of motion vectors representing full and sub-pixel motion. In FIG. 6, several blocks 602, 604, 606, 608 of a current frame 600 are inter predicted using pixels from a reference frame 630. In this example, the reference frame 630 is a reference frame, also called the temporally adjacent frame, in a video sequence including the current frame 600, such as the video stream 300. The reference frame 630 is a reconstructed frame (i.e., one that has been encoded and decoded such as by the reconstruction path of FIG. 4) that has been stored in a so-called last reference frame buffer and is available for coding blocks of the current frame 600. Other (e.g., reconstructed) frames, or portions of such frames may also be available for inter prediction. Other available reference frames may include a golden frame, which is another frame of the video sequence that may be selected (e.g., periodically) according to any number of techniques, and a constructed reference frame, which is a frame that is constructed from one or more other frames of the video sequence but is not shown as part of the decoded output, such as the output video stream 516 of FIG. 5.
[0069] A prediction block 632 for encoding the block 602 corresponds to a motion vector 612. A prediction block 634 for encoding the block 604 corresponds to a motion vector 614. A prediction block 636 for encoding the block 606 corresponds to a motion vector 616. Finally, a prediction block 638 for encoding the block 608 corresponds to a motion vector 618. Each of the blocks 602, 604, 606, 608 is inter predicted using a single motion vector and hence a single reference frame in this example, but the teachings herein also apply to inter prediction using more than one motion vector (such as bi-prediction and / or compound prediction using two different reference frames), where pixels from each prediction are combined in some manner to form a prediction block.
[0070] As mentioned above, a list of candidate MV s may be generated according to predetermined rules. The predetermined rules for generating (e.g., deriving, or constructing and ordering) the list of candidate MVs and the number of candidates in the list may vary by codec. For example, in High Efficiency Video Coding (H.265), the list of candidate MVs can include up to 5 candidate MVs.
[0071] Codecs may populate the list of candidate MVs using different algorithms, techniques, or tools (collectively, tools). Each of the tools may produce a group of MVs that are added to the list of candidate MVs. For example, in Versatile Video Coding (H.266), the list of candidate MVs may be constructed using several modes, including intra-block copy (IBC) merge, block level merge, and sub-block level merge. The details of these modes are not necessary for the understanding of this disclosure. H.266 limits the number of candidate MVs obtained using IBC merge, block- level merge, and sub-block level merge, to 6 candidates, 6 candidates, and 5 candidates, respectively. In the Alliance for Open Media (AOMedia) Video 1 (AV 1) codec, the list of candidate MVs is limited to 4 candidates. Different codecs may use different techniques for generating lists of candidate MVs.However, such nuances are not necessary for the understanding of this disclosure. As such, the disclosure merely assumes a use of a list of candidate MVs.
[0072] FIG. 7 is a block diagram of an example of a reference frame buffer 700. The reference frame buffer 700 stores reference frames used to encode or decode blocks of frames of a video sequence. The reference frame buffer 700 is shown as including eight reference frames. However, the reference frame buffer 700 may include more or ewer frames.
[0073] At least some of the frames of the reference frame buffer 700 can be available for coding the blocks of the current frame. In an example, up to two reference frames can be selected from the reference frame buffer 700 for use as reference frames for inter prediction. Only one of the frames may be used as the reference frame in the case of single reference inter prediction; and two reference frames may be used in the case of compound inter prediction, where a prediction is generated by combining two prediction blocks based on the two reference frames.
[0074] Type (e.g., roles), with respect to the current frame, may be associated with the reference frames stored in the reference frame buffer 700. The reference frame buffer 700 is shown as including at least a last frame LAST_FRAME 702, a golden frame GOLDEN_FRAME 704, an alternative reference frame ALTREF_FRAME 706, and a backward frame BWDREF_FRAME 708.
[0075] The last frame LAST_FRAME 702 can be, for example, the adjacent frame immediately before the current frame in the video sequence, which is a forward reference frame. The LAST_FRAME 702 can be a frame that was displayed in the near past. The golden frame GOLDEN_FRAME 704 can be, for example, a reconstructed video frame for use as a reference frame that may or may not be adjacent to the current frame. The alternative reference frame ALTREF_FRAME 706 can be, for example, a video frame in the non-adjacent future, which is a backward reference frame. The backward frame BWDREF_FRAME 708 can be a frame that will be displayed in the future and may be stored as an additional backward prediction reference frame in addition to ALTREF_FRAME 706. The backward frame BWDREF_FRAME 708 can be closer in relative distance to the current frame than the ALTREF_FRAME 706, for example. The alternative reference frame ALTREF_FRAME 706 can be a frame from either the past or the future. The reference frame buffer 700 can also include additional alternative reference frames (ALTREF_FRAMEs). At least some of the alternative reference frames can also have the option of not being displayed. At least some of the alternative reference frames can be synthesized by the encoder, for example, by temporal filtering along motion trajectories of multiple frames.
[0076] The frame header of a reference frame includes a virtual index 710 to a location within the reference frame buffer 700 at which the reference frame is stored. A reference frame mapping 714 maps the virtual index 710 of a reference frame to a physical index 716 of memory at which the reference frame is stored. Where two reference frames are the same frame, those reference frames will have the same physical index even if they have different virtual indexes. One or more refresh flags 712 can be used to remove one or more of the stored reference frames from the reference frame buffer 700, for example, to clear space in the reference frame buffer 700 for new reference frames, where there are no further blocks to encode or decode using the stored reference frames, or where a new golden frame is encoded or decoded.
[0077] The reference frames stored in the reference frame buffer 700 can be used to identify motion vectors for predicting blocks of frames to be encoded or decoded. Different reference frames may be used depending on the type of prediction (e.g., the prediction mode) used to predict a current block of a current frame. For example, in bidirectional inter prediction, blocks of the current frame can be forward predicted using either the LAST_FRAME 702 or the GOLDEN_FRAME 704, or backward predicted using the ALTREF_FRAME 706. When compound prediction is used, multiple frames, such as one for forward prediction (e.g., LAST_FRAME 702 or GOLDEN_FRAME 704) and one for backward prediction (e.g., ALTREF_FRAME 706) can be used for predicting the current block.
[0078] There may be a finite number of reference frames that can be stored within the reference frame buffer 700. As shown in FIG. 7, the reference frame buffer 700 can store up to eight reference frames, wherein each stored reference frame may be associated with a different virtual index 702 of the reference frame buffer. Although four of the eight spaces inthe reference frame buffer 700 are used by the LAST_FRAME 702, the GOLDEN_FRAME 704, the ALTREF_FRAME 706, and the BWDREF_FRAME 708, four spaces remain available to store other reference frames. In particular, one or more available spaces in the reference frame buffer 700 may be used to store a second last frame LAST2_FRAME and / or a third last frame LAST3_FRAME as additional forward reference frames, in addition to the LAST_FRAME 702.
[0079] The list of candidate MVs may be constructed as described for each reference frame associated with a previously coded block or sub-block. For example, up to six reference frames may be available for each frame as described above — e.g., LAST_FRAME 702, LAST2_FRAME, LAST3_FRAME, GOLDEN_FRAME 704, ALTREF_FRAME 706, and BWDREF_FRAME 708. In this case, separate lists of candidate MVs may be constructed using those previously coded blocks or sub-blocks having motion vectors pointing each of the candidate reference frames.
[0080] One or more available spaces in the reference frame buffer 700 may be used to store additional alternative reference frames (e.g., ALTREF1_FRAME, ALTREF2_FRAME, etc., wherein the original alternative reference frame ALTREF_FRAME 706 could be referred to as ALTREF0_FRAME). The ALTREF_FRAME 706 is a frame of a video sequence that is distant from a current frame in a display order, but is encoded or decoded earlier than it is displayed. For example, the ALTREF_FRAME 706 may be ten, twelve, or more (or fewer) frames after the current frame in a display order.
[0081] It is possible for the BWDREF_FRAME 708 or the additional alternative reference frames to be frames located nearer to the current frame in the display order. In one example, the BWDREF_FRAME 708 can be one frame after the current frame in the display order. In another example, a first additional alternative reference frame, ALTREF2_FRAME, can be five or six frames after the current frame in the display order, whereas a second additional alternative reference frame, ALTREF3_FRAME, can be three or four frames after the current frame in the display order. Being closer to the current frame in display order increases the likelihood of the features of a reference frame being more similar to those of the current frame. As such, one or more of the BWDREF_FRAME 708 or additional alternative reference frames can be stored in the reference frame buffer 700 as additional options usable for backward prediction.
[0082] FIG. 8 is an illustration 800 of compound inter prediction. The illustration 800 includes a current frame 802 that includes a current block 804 to be coded (i.e., encoded or decoded) using a first MV 806 (i.e., MVo) that refers (i.e., points) to a first reference frame808 (i.e., Ro) and a second MV 810 (i.e., MVi) that refers to a second reference frame 812 (i.e., Ri). A line 814 illustrates the display order, in time, of the frames. As such, the illustration 800 is an example of the bidirectional inter prediction mode since the current frame 802 is between the first reference frame 808 and the second reference frame 812 in the display order. However the disclosure herein is not limited to the bidirectional inter prediction mode and the techniques described herein can also be used with (e.g., adapted to) unidirectional inter prediction.
[0083] The distance, in display order, between the first reference frame 808 and the current frame 802 is denoted do; and the distance, in display order, between the current frame 802 and the second reference frame 812 is denoted di. While not specifically shown in FIG.8, each of the first MV 806 and the second MV 810 includes a horizontal and vertical offset. Thus, MVo.x and MVo.y can denote, respectively, the horizontal and the vertical components of the first MV 806; and MVi,xand MVi,ycan denote, respectively, the horizontal and the vertical components of the second MV 810. The first MV 806 and the first reference frame 808 can be used to obtain a first prediction block 816 (denoted Po) for the current block 804; and the second MV 810 and the second reference frame 812 can be used to obtain a second prediction block 818 (denoted Pi) for the current block 804. A final prediction block for the current block 804 can be obtained as a combination (e.g., a pixel-wise weighted average) of the first prediction block 816 and the second prediction block 818.
[0084] FIGS. 9A-9B illustrate examples of tools for generating motion vector candidates. The processes described with respect to FIGS. 9A-9B are those implemented in the AV 1 codec. However, adaptively generating motion vector candidates list, as described herein, is not limited to or by any particular implementation of MV candidate list generation.
[0085] As mentioned above, a list of candidate MV s may be obtained using different tools. An encoder, such as the encoder 400 of FIG. 4, and a decoder, such as the decoder 500 of FIG. 5, may use the same tools for obtaining (e.g., populating, constructing, etc.) the list of candidate MVs. The candidate MVs obtained using a tool may be referred to herein as a group of candidate MVs. At least some of the tools described herein may be known or may be similar to or used by other codecs. However, the disclosure is not limited to or by any particular tools that can generate groups of MV candidates. The groups of motion vectors may be or may be combined to form a list of candidate MVs.
[0086] As mentioned above, merge candidates or candidate MVs may be derived using different tools. Some such tools are now described. Depending on the inter prediction mode,diff erent motion information may be coded in a compressed bitstream, such as the compressed bitstream 420 of FIGS. 4 or 5.
[0087] If a block is coded using the MERGE mode, a reference frame index and a motion vector of the list of candidate MVs are set as the reference frame index and motion vector of the block. A merge candidate corresponding to a merge index (e.g., the index of the candidate in the list of candidate MVs) is selected from the merge candidate list and the motion information of the merge candidate is set as the motion information of the block. The merge index (e.g., the index of the candidate in the list of candidate MVs) may be coded in the compressed bitstream. In the MERGE mode, a current block is merged with its neighboring block(s) to form a region therewith sharing the same motion parameters. Thus, there is no need to code and transmit motion parameters for the current block. Instead, for a region, only one set of motion parameters is coded and transmitted in the compressed bitstream.
[0088] If a motion vector is coded differentially, an MVP is selected from list of candidate MVs. The index of the MVP in the list of candidate MVs may be included in the compressed bitstream. The MVD may also be included (i.e., coded) in the compressed bitstream. Additionally, a reference frame index may also be included (i.e., coded) in the compressed bitstream.
[0089] For the derivation of MV predictors, all the spatial MV candidates (which may be generated as described with respect to FIG. 9A) and the temporal MV candidates (which may be generated as described with respect to FIG. 9B) can be pooled, and each predictor can be assigned a weighting that is determined during the scanning of the spatial and temporal neighboring blocks. Based on the associated weightings, the candidates can be sorted and ranked. Conventionally, a predefined number (e.g., 4) of the candidate MVs are added to the list of MV candidates, which may also be referred to a dynamic reference list (DRL). However, as described herein, the number of generated MV candidates can be adaptively based on one or more criteria related to the current block. Said another way, the size of the list of candidate MV s can be adaptively set to be higher or lower than the predefined number.
[0090] FIG. 9A illustrates an example 900 of generating a group of motion vector candidates for a current block from spatial neighboring blocks of the current block.
[0091] Obtaining candidate MVs from spatial neighbors is first described. Given a current block 902, spatial MV predictors can be identified by utilizing spatial neighboring blocks, including adjacent spatial neighboring blocks, which are direct neighbors of the current block to the top and left sides (such as the blocks in regions 904A and 904B), as well as non-adjacent spatial neighboring blocks, which are close but not directly adjacent to thecurrent block (such the other blocks shown in FIG. 9A). An example of a set of spatial neighboring blocks for a luma block is illustrated in FIG. 9A, wherein each spatial neighboring block (such as each of blocks 906A-906E) is an 8x8 block.
[0092] The spatial neighboring blocks are examined to find one or multiple MVs that are associated with the same reference frame index as the current block 902. For the current block 902, the search order of spatial neighboring 8x8 luma blocks can be as indicated by the numbers 1-8 in FIG. 9A: the top adjacent row (e.g., the blocks of the region 904A) is checked from left to right, then the left adjacent column (e.g., the blocks of the region 904B) is checked from top to bottom, then the top-right neighboring block (e.g., the block 906B) is checked, then the top-left block neighboring block (e.g., the block 906C) is checked, and so.
[0093] More specifically, in the first step, the bottom two 4x4 blocks of each 8x8 neighboring block (e.g., each of the blocks of the region 904A) are checked. In the second step, the right two 4x4 blocks of each 8x8 neighboring block (e.g., each of the blocks of the region 904B) are checked. In the third step, the bottom-left 4x4 block of the top-right 8x8 neighboring block (e.g., the block 906B) is checked. Thereafter, in the steps 4- 8, the bottomright 4x4 block of each 8x8 neighboring block is checked.
[0094] In the case of single reference inter prediction, the spatial MV predictors are generated by identifying the spatial neighboring blocks that are predicted using the same single reference frame as the current block 902, and their associated MVs are used as the spatial MV predictors. In compound prediction, when a spatial neighboring block is predicted by the compound prediction mode using the same reference frames as the current block 902, the associated MVs can be used as the spatial MV predictor. In the case that compound prediction utilizes two reference frames, non-adjacent spatial neighbors may not be checked when deriving the MV predictor.
[0095] FIG. 9B illustrates an example 940 of temporal motion vector prediction. In addition to spatial neighboring blocks, as described with respect to FIG. 9A, MV predictors (also referred to as temporal MV predictors) can also be derived based on a motion field.
[0096] A motion field can be created for each reference frame ahead of processing the current frame. Motion trajectories can be built between the current frame and the previously coded frames by exploiting motion vectors from previously coded frames through either linear interpolation or extrapolation. The motion trajectories can be associated with 8x8 blocks in the current frame. The motion field between the current frame and a given reference frame can then be formed by extending the motion trajectories from the current frame towards the reference frame.
[0097] Given the coordinates of a current block, the associated MVs stored in the temporal MV buffer are identified and projected to derive a temporal MV predictor that points from the current block to its reference frame. In an example, up to seven blocks may be checked to find valid temporal MV predictors.
[0098] FIG. 9B illustrates using a motion trajectory for a current frame 950 to predict motion of a current block 952. The motion trajectory shows motion between the current frame 950 and three reference frames 954, 956, and 958. The motion trajectory is determined based on a reference motion vector 960, which indicates motion between the reference frame 0 956 and the reference frame 1 954. For example, the reference frame 0 956 may be a reference frame used to predict motion of one or more blocks of the reference frame 954. After the reference motion vector 960 is determined, the motion trajectory is determined as the trajectory corresponding to the direction of the reference motion vector 960.
[0099] The motion trajectory identifies the current block 952 as the location of the current frame 950 intersected by the motion trajectory. A first temporal motion vector candidate 962 may then be determined as indicating motion between the reference frame 1 954 and the current frame 950. A second temporal motion vector candidate 964 may be determined as indicating motion between the reference frame 2 958 and the current frame 950. One or more of the reference motion vector 960, the first temporal motion vector candidate 962, or the second temporal motion vector candidate 964 may be included in a motion vector candidate list from which a motion vector is selected for predicting motion of the current block 952.
[0100] The reference frame 0 956 and the reference frame 2958 are shown as past frames with respect to the current frame 950. The reference frame 1 954 is shown as a future frame with respect to the current frame 950. However, other numbers of past or future reference frames may be used. For example, a motion trajectory can be determined where there is one past reference frame and one future reference frame. In another example, a motion trajectory can be determined where there is one past frame and two future reference frames. In another example, a motion trajectory can be determined where there are two or more past reference frames and two or more future reference frames.
[0101] In an implementation, motion information associated with blocks of a reference frame can be used to perform linear projections of blocks of a current frame in order to estimate the motion field of the current frame. That is, the motion information associated with blocks of the reference frame can be used to estimate the motion information (i.e., the motion field) of some blocks of the current frame.
[0102] As described above, a block can be predicted using inter prediction. When a frame is reconstructed (as described above with respect to the reconstruction stage 414 of FIG. 4 and the reconstruction stage 510 of FIG. 5), a type and the motion vector associated with each coding block can be retained (i.e., saved). The type can be associated with a frame rather than with each block of the frame. The type can indicate which reference frame(s) is (are) used to code blocks of the frame. As such, the type can indicate which reference frame a motion vector points to, and the motion vector indicates the offset within the reference frame. The type and motion vector information (collectively, the motion information) can be retained for blocks of size 4x4, 8x8, or some other same -dimensioned block sizes.
[0103] For example, when the reference frame 1 954 is reconstructed, the prediction type and motion vector information associated with blocks of the reference frame 1 954 can be retained. For example, the reference motion vector 960 and a type indicating the reference frame 956 are retained for a block 966 of the reference frame 1 954.
[0104] As mentioned above, the reference frame 1 954 can be used as a reference frame for the current frame 950. As such, the retained information of the blocks of the reference frame 1 954 can be used to perform linear projections to estimate the motion field of the current frame 950. The motion field of the current frame means the collective motions fields of blocks of the current frame.
[0105] The spatial MV candidates and the temporal MV candidates may be categorized based on their location with respect to the current block. Two categories may be used: nearest spatial neighbors to the current block and all others. MVs from immediate neighbors (e.g., above, left, top-right blocks) may be considered to have higher correlation with the current block and are prioritized over the MVs of other blocks. The candidate MVs may then be ranked based on their appearance frequency, thereby forming a ranked list. Conventionally, and as already mentioned, the top predefined number (e.g., 4) of the ranked list are used as candidate MV predictors. That is, the top predefined number of MVs constitutes the list of MV candidates. The encoder selects the closest match in the list to the actual MV of the current block and transmits the index of the closest in the compressed bitstream. However, as described herein the number of candidates added to the list of candidate MV s can be adaptively set based on criteria related to a current block. Said another way, the size of the list of candidate MV s can be adaptively set.
[0106] In the case of single reference inter prediction, separate lists of MVs may be maintained for each reference frame. A syntax element, drl_idx, can be used (e.g., signaled in the compressed bitstream) to specify the motion vector prediction mode. The predictionmodes may include: NEARESTMV, NEARMV, NEWMV, and GLOBALMV. The inter prediction mode NEARESTMV indicates that the 0-indexed entry from the list of MV candidates is to be used, the inter prediction mode NEARMV indicates that one of the 1, 2, or 3-indexed entries, as indicated by another index, is to be used. The inter prediction mode NEWMV indicates that one of the motion vector predictors in the list signaled by an index as reference is to be used and to apply a delta (e.g., a DMV) to the MVP. The inter prediction mode GLOBALMV indicates that an MV that is based on frame-level global motion parameters is to be used.
[0107] In the case of compound inter prediction modes, the syntax element drl_idx can be used (e.g., signaled in the compressed bitstream) to specify which of the following modes to use: NEAREST_NEARESTMV, NEAR_NEARMV, NEAREST_NEWMV, NEW_NEARESTMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV.
[0108] The inter prediction mode NEAREST_NEARESTMV indicates that the 0-indexed MV pair from the lists of MV candidates are used. The inter prediction mode NEAREST_NEWMV (NEW_NEARESTMV) indicates that the 0-indexed MV pair from the lists of candidate MVs are used and that an MVD is signaled for the second (first) MV. The inter prediction mode NEAR_NEARMV indicates that one of the motion vector predictors (MVP) in the list signaled by a DRL index is to be used. The inter prediction mode NEAR_NEWMV indicates that one of the motion vector predictors (MVP) in the list signaled by a DRL index as reference is used and that a delta MV for a second MV is included in the compressed bitstream. The inter prediction mode NEW_NEARMV indicates that one of the motion vector predictors (MVP) in the list signaled by a DRL index is used as reference and that a delta MV for the first MV is included in the compressed bitstream. The inter prediction mode NEW_NEWMV indicates that one of the motion vector predictors (MVP) in the list signaled by a DRL index is used as reference and that delta MV s are included in the compressed bitstream for each of MVs. The inter prediction mode GLOBAL_GLOBALMV indicates to use MVs from each reference based on their framelevel global motion parameters.
[0109] FIG. 10 is a flowchart diagram of a method or technique 1000 for coding a current block of a current frame. The technique 1000 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine -readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that,when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 1000. The technique 1000 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
[0110] The technique 1000 may be implemented in whole or in part in the intra / inter prediction stage 402 of the encoder 400 of FIG. 4 and / or the intra / inter prediction stage 508 of the decoder 500 of FIG. 5. When implemented by an encoder, “coding” means encoding in a compressed bitstream, such as the compressed bitstream 420 of FIG. 4; and when implemented by a decoder, “coding” means decoding from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5. The description herein refers to “higher” or “lower” number of candidate MVs. These relative terms can be with respect to a baseline (e.g., a predefined) number of candidate MVs. To illustrate, if the baseline number is 4, then a lower (or fewer) number of candidate MVs added to the list is 2, and a higher number of candidate MVs added to the list can be 6 or 8 or some other number. The size of the list of candidate MVs is ideally a power of two.
[0111] At 1002, a number of MV candidates is determined based on one or more criteria related to the current block. That is, the determined number sets the size of the list of candidate MVs. The MV candidates are then used to generate a list of candidate MVs. The one or more criteria related to the current block can be or include one or more of the following.
[0112] The number of motion vector candidates (i.e., the size of the list of candidate MVs) can be based, at least in part, on whether the reference frame of the current frame is a generated reference frame. A generated reference frame can be a synthesized frame generated to improve coding efficiency. If the reference frame is a generated frame, then it can be expected that the current block has high quality matches in the reference frame. As such, less MV candidates are required when the reference frame is a generated frame; and more MV candidates are required when the reference frame is not a generated frame.
[0113] When implemented at the encoder, the technique 1000 may generate a respective list of candidate MVs for each reference frame. Then, the particular reference frame used for encoding the current block is signaled in the compressed bitstream. When implemented at the decoder, the decoder generates a list of candidate MVs only for the signaled reference frame.
[0114] One example of a generated reference frame is a temporally interpolated picture (TIP) reference frame. A TIP reference frame can be used in bidirectional inter prediction. A TIP reference frame is a reference frame generated by interpolating reference blocks from a forward reference frame and a backward reference frame (e.g., as the nearest future and pastreference frames relative to the current frame). In particular, the coded motion vectors available in the forward and backward reference frames are used to generate a motion field for the current frame, and the motion field is the used to fetch the reference blocks which are used to generate the TIP reference frame. TIP video coding thus refers to an inter-prediction mode whereby a TIP reference frame is used to predict the motion of a current frame. TIP video coding typically involves a relatively small motion vector being applied against the TIP reference frame, which small motion vector is not only cheaper to encode, but also improves prediction detail and quality due to the TIP reference frame leveraging forward and backward reference data. The TIP reference frame is independently generated at each of the encoder and the decoder. In particular, the encoder generates the TIP reference frame using data determined as part of an encoder search process, and the decoder generates the TIP reference frame using bitstream data indicative of that encoder search process. The use of this TIP mode for video coding has shown remarkable coding gain achievements relative to video coding schemes which do not use the TIP mode.
[0115] The number of MV candidates can be based, at least in part, on a determined quality of the reference frame. The number of MV candidates can be directly related to the determined quality. That is, the higher the quality, the higher the number of MV candidates; and vice versa. A quality score can be associated with at least some (e.g., each) of the reference frames of a reference frame buffer. The reference frame with the highest quality score can be assigned the index 0, the reference frame with second highest quality score is assigned the index 1, and so on. The quality score can be calculated based on the QPs associated with the reference frames, their distances to the current frame, or a combination thereof.
[0116] In an example, the quality of the reference frame can be related to or based on the QP used to encode the reference frame. The number of MV candidates (i.e., the size of the list of candidate MV s) can be inversely related to the QP. That is, the higher the QP, the lower the number of MV candidates; and the lower the QP, the higher the number of MV candidates.
[0117] In an example, the quality of the reference frame can be related to the distance between the reference frame and the current frame. The closer the reference frame and the current frame, the more correlated they are likely to be. As such, the closer the reference frame is to the current frame, the higher the quality of the reference frame. In an example, the distance between the reference frame and the current frame can be determined based on the label (e.g., role) associated with the reference frame, as described with respect to FIG. 6.
[0118] The number of MV candidates can be based, at least in part, on the size of the current block. Smaller blocks (e.g., 8x8, 8x4, 4x8, or smaller) typically have associated therewith very little transform coefficient information. As such, the rate cost in signaling an index of a larger number of MV candidates would outweigh the transform coefficient information. As such, for smaller blocks, the number of MV candidates can be smaller than for larger blocks. For example, for smaller blocks, the number of MV candidates can be 2.
[0119] The number of MV candidates can be determined based, at least in part, on the prediction mode of the current block. Specifically, a higher number of MV candidates can be generated when the block utilizes a compound prediction mode, as opposed to a lower number if the prediction mode is a non-compound prediction mode.
[0120] The number of MV candidates can be based, at least in part, on whether the prediction mode of the current block is a unidirectional inter prediction mode or a bidirectional inter prediction mode. The determination of the number of MV candidates can be contingent upon the type of prediction mode utilized: unidirectional or bidirectional inter prediction. In a unidirectional inter prediction mode, a single list of MV candidates is generated, whereas in a bidirectional inter prediction mode, two lists are generated, one for each direction. To optimize bit usage in coding MV candidate indexes for a bidirectional inter prediction, a reduced number of MV candidates can be generated for each direction, compared to the number of MV candidates in unidirectional inter prediction. For instance, in a scenario where four MV candidates may be generated for a unidirectional inter prediction mode, two MV candidates per direction may be generated in the case of bidirectional inter prediction mode.
[0121] The number of MV candidates can be based, at least in part, on a QP used for (e.g. associated with) the current block. The QP used for the current block may be specified at the current frame level, at a super block level (i.e., the super block that include the current block), or at the current block level. A “super block” may also be referred to as a macroblock and is the largest possible coding block. If no block-level QP is set, then the frame-level QP may be used when coding the current block. In an example, the QP used for the current block may be a QP associated with a segment that includes the current block. As such, the QP used for the block may be obtained using a segment identifier associated with the current block.
[0122] In scenarios characterized by high QP, coding of transform coefficients may be omitted, leading to the non-transmission of transform coefficients to the decoder. This results in an increased relative proportion of bits allocated for MV candidates indexes. Consequently, in high QP contexts, it becomes advantageous to consider a reduction in thenumber of MV candidate indexes to maintain bit efficiency. Conversely, in low QP settings where a larger volume of transform coefficients is transmitted, the bit allocation for MV candidate indexes constitutes a relatively minor proportion of the overall bit usage. In such scenarios, the allocation of additional bits for MV candidate indexes can be beneficial. Increasing the bit allocation for MV candidate indexes enhances prediction quality by expanding the available pool of MV candidates. To illustrate, in the case of low QP, the number of MV candidates may be increased from four to eight to improve the likelihood of selecting a more optimal MV candidate, thereby enhancing the overall prediction efficacy.
[0123] In an example, the number of MV candidates can be limited to a maximum number (e.g., a ceiling) of motion vector candidates. The maximum number can be coded in a header of the current frame. That is, the encoder may encode the maximum number in the header of the current frame. The encoder can determine the maximum number in any number of ways.
[0124] The encoder may determine the ceiling based on statistical analysis of previously encoded frames, or as a result of multi-pass encoding. For example, an initial encoding pass of a frame might utilize a predetermined number of MV candidates (e.g., eight MV candidates). Subsequent analysis of this initial pass could reveal that more than a threshold number (e.g., more than 90%) of the blocks only utilize the top two of these MV candidates. Based on this insight, the encoder may adjust the ceiling to a lower value, such as two, in the final encoding pass. This approach optimizes the number of MV candidates in the list of candidate MVs, ensuring efficient encoding by aligning the number of MV candidate indexes with actual usage patterns observed in the frame data.
[0125] In an example, the number of MV candidates can be based, at least in part, on the variations in (e.g., heterogeneity of) the MV candidates generated. That is, the number of MV candidates can be based on the diversity amongst the MV candidates. The number of MV candidates can be directly proportional to the diversity observed among the MV candidates. A higher diversity among the MV candidates can result in an increased number of MV candidates being added to the list of MV candidates. In contrast, a lower diversity, indicated by the MV candidates being proximally similar (e.g., within a predetermined pixel range), reduces the number of MV candidates being added to the list of candidate MVs. That is, if the MV candidates are close to each other, then the additional bit cost for the additional indexes may not be justified.
[0126] The diversity of the MV candidates can be determined based on a metric. The metric, which can be the variance in the MV candidates, evaluates the degree of variationwithin their horizontal and vertical components. For instance, if the variances in both the horizontal and vertical components of the MV candidates fall below (or exceed) a predefined threshold, the candidates are deemed to lack sufficient diversity (or to be adequately diverse), consequently leading to a smaller (or larger) subset of these candidates being incorporated into the list of candidate MVs. Alternatively, the diversity metric may be defined as the maximum difference observed between the horizontal and vertical components of the MV candidates. As such, the number of motion vector candidates is based on a diversity metric calculated with respect to possible motion vector candidates.
[0127] In an example, the number of MV candidates can be based, at least in part, on the respective numbers of candidate MVs used by (e.g., for or with respect to) one or more neighboring inter-predicted blocks of the current block. This is based on the likelihood that the current block exhibits similar characteristics to the neighboring blocks. Simply stated, if a neighboring block uses more candidates in its list of candidate MVs, then it may be preferable to follow the same trend for the current block. The one or more neighboring interpredicted blocks can be one or more neighboring blocks that are coded prior to the current block. In an example, the one or more neighboring inter-predicted blocks can be the block directly to the left (e.g., the left-neighboring block), the block directly above (e.g., the aboveneighboring block), other neighboring inter-predicted blocks, or a combination thereof. To illustrate, if the left-neighboring block is used, then the number of MV candidates for the current block can be based on the number of MV candidates in the list of MV candidates used by that left-neighboring block. As such, the number of motion vector candidates can be based on a size of a list of MV candidates associated with a neighboring block.
[0128] In a specific, nonlimiting implementation, a smaller number of MV candidates than a predefined number of MV candidates can be generated if the width of the current block is less than 8, the height of the current block is less than 8, the current block is not coded using skip mode, the prediction mode of the current block is not a unidirectional inter prediction mode, and the reference frame is not the best quality reference frame. The smaller number of MV candidates can be generated if the reference frame is a TIP frame. In another specific implementation, a larger number of MV candidates than the predefined number of MV candidates can be generated if the width of the current block is less than 8, the height of the current block is less than 8, the current block is not coded using skip mode, the prediction mode of the current block is not a unidirectional inter prediction mode, and the reference frame is the best quality reference frame.
[0129] At 1004, the number of the motion vector candidates is generated. To be more specific, more than just the number of motion vector candidates may be generated. Any number of techniques can be used to generate and rank motion vectors. However, only the top number of the motion vector candidates are added to the list of candidate motion vectors. That is, at 1006, the motion vector candidates are added to the list of candidate motion vectors.
[0130] At 1008, an index of one of the motion vector candidates from the list of candidate motion vectors is coded. When the technique 1000 is implemented at the encoder, coding the index can mean encoding the index in the compressed bitstream; and when the technique 1000 is implemented at the decoder, coding the index can mean decoding the index from the compressed bitstream. The index is coded using a minimum number of bits needed for coding the number of the motion vector candidates.
[0131] For simplicity of explanation, the technique 1000 of FIG. 10 is depicted and described as a series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustrated steps or operations may be required to implement a method in accordance with the disclosed subject matter.
[0132] Some implementations are described below as numbered examples (Example 1 , 2, 3, etc.). These examples are provided as examples only and do not limit the other implementations disclosed herein.
[0133] Example 1 is a method that includes determining a number of motion vector candidates based on one or more criteria related to the current block; generating the number of the motion vector candidates; adding the motion vector candidates to a list of candidate motion vectors; and coding an index of one of the motion vector candidates, where the index is coded using a minimum number of bits needed for coding the number of the motion vector candidates.
[0134] Example 2 is the method of Example 1, where the number of the motion vector candidates is based on whether a reference frame of the current frame is a generated reference frame.
[0135] Example 3 is the method of Example 1, where the number of the motion vector candidates is based on a quality of a reference frame.
[0136] Example 4 is the method of Example 3, where the quality of the reference frame is based on at least one of a quantization parameter associated with the reference frame or a temporal distance between the reference frame and the current frame.
[0137] Example 5 is the method of Example 1, where the number of the motion vector candidates is based on a block size of the current block.
[0138] Example 6 is the method of Example 1, where the number of the motion vector candidates is based on whether the current block is predicted using a compound prediction mode.
[0139] Example 7 is the method of Example 1, where the number of the motion vector candidates is based on whether a prediction mode of the current block is a unidirectional prediction mode or a bidirectional prediction mode.
[0140] Example 8 is the method of Example 1, where the number of the motion vector candidates is based on a quantization parameter associated with the current block.
[0141] Example 9 is the method of Example 1, where the number of the motion vector candidates is limited to a maximum number, where the maximum number is coded in a header of the current frame.
[0142] Example 10 is the method of Example 1, where the number of the motion vector candidates is based on a diversity metric calculated with respect to possible motion vector candidates.
[0143] Example 11 is the method of Example 1 , where the number of the motion vector candidates is based on a size of a list of candidate motion vectors associated with a neighboring block.
[0144] Example 12 is a device that includes a processor configured to perform the method of any one of Examples 1-11.
[0145] Example 13 is a device that includes a memory and a processor, where the processor is configured to execute instructions stored in the memory to perform the method of any one of Examples 1-11.
[0146] Example 14 is a non-transitory computer-readable storage medium that includes executable instructions which, when executed by a processor, facilitate performance of operations that perform the method of any one of Examples 1-11.
[0147] Example 15 is a non-transitory computer-readable storage medium having stored thereon an encoded bitstream, where the encoded bitstream is configured for decoding by the method of any one of Examples 1-11.
[0148] Example 16 is a non-transitory computer-readable storage medium having stored thereon an encoded bitstream, where the encoded bitstream is generated by an encoder performing the method of any one of Examples 1-11.
[0149] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it is to be understood that encoding and decoding, as those terms are used in the claims, could mean compression, decompression, transformation, or any other processing or change of data.
[0150] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X includes A or B” is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form. Moreover, use of the term “an implementation” or “one implementation” throughout is not intended to mean the same embodiment or implementation unless described as such.
[0151] Implementations of the transmitting station 102 and / or the receiving station 106 (and the algorithms, methods, instructions, etc., stored thereon and / or executed thereby, including by the encoder 400 and the decoder 500) can be realized in hardware, software, or any combination thereof. The hardware can include, for example, computers, intellectual property (IP) cores, application-specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors or any other suitable circuit. In the claims, the term “processor” should be understood as encompassing any of the foregoing hardware, either singly or in combination. The terms “signal” and “data” are used interchangeably. Further, portions of the transmitting station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.
[0152] Further, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a general-purpose computer or general-purpose processor with a computer program that, when executed, carries out any of the respectivemethods, algorithms and / or instructions described herein. In addition, or alternatively, for example, a special purpose computer / processor can be utilized which can contain other hardware for carrying out any of the methods, algorithms, or instructions described herein.
[0153] The transmitting station 102 and the receiving station 106 can, for example, be implemented on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server and the receiving station 106 can be implemented on a device separate from the server, such as a hand-held communications device. In this instance, the transmitting station 102 can encode content using an encoder 400 into an encoded video signal and transmit the encoded video signal to the communications device. In turn, the communications device can then decode the encoded video signal using a decoder 500. Alternatively, the communications device can decode content stored locally on the communications device, for example, content that was not transmitted by the transmitting station 102. Other suitable transmitting and receiving implementation schemes are available. For example, the receiving station 106 can be a generally stationary personal computer rather than a portable communications device and / or a device including an encoder 400 may also include a decoder 500.
[0154] Further, all or a portion of implementations of the present disclosure can take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transport the program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or a semiconductor device. Other suitable mediums are also available.
[0155] The above-described embodiments, implementations and aspects have been described in order to allow easy understanding of the present invention and do not limit the present invention. On the contrary, the invention is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structure as is permitted under the law.
Claims
What is claimed is:
1. A method for coding a current block of a current frame, comprising: determining a number of motion vector candidates based on one or more criteria related to the current block; generating the number of the motion vector candidates; adding the motion vector candidates to a list of candidate motion vectors; and coding an index of one of the motion vector candidates, wherein the index is coded using a minimum number of bits needed for coding the number of the motion vector candidates.
2. The method of claim 1, wherein the number of the motion vector candidates is based on whether a reference frame of the current frame is a generated reference frame.
3. The method of claim 1, wherein the number of the motion vector candidates is based on a quality of a reference frame.
4. The method of claim 3, wherein the quality of the reference frame is based on at least one of a quantization parameter associated with the reference frame or a temporal distance between the reference frame and the current frame.
5. The method of claim 1, wherein the number of the motion vector candidates is based on a block size of the current block.
6. The method of claim 1, wherein the number of the motion vector candidates is based on whether the current block is predicted using a compound prediction mode.
7. The method of claim 1, wherein the number of the motion vector candidates is based on whether a prediction mode of the current block is a unidirectional prediction mode or a bidirectional prediction mode.
8. The method of claim 1, wherein the number of the motion vector candidates is based on a quantization parameter associated with the current block.
9. The method of claim 1, wherein the number of the motion vector candidates is limited to a maximum number, wherein the maximum number is coded in a header of the current frame.
10. The method of claim 1, wherein the number of the motion vector candidates is based on a diversity metric calculated with respect to possible motion vector candidates.
11. The method of claim 1 , wherein the number of the motion vector candidates is based on a size of a list of candidate motion vectors associated with a neighboring block.
12. A device, comprising: a processor that is configured to perform the method of any one of claims 1-11.
13. A device, comprising: a memory; and a processor, the processor configured to execute instructions stored in the memory to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any one of claims 1-11.
15. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is configured for decoding by the method of any one of claims 1-11.
16. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is generated by an encoder performing the method of one any of claims 1-11.
Citation Information
Patent Citations
Video signal encoding / decoding method and apparatus for said method
CN113382234B
Motion vector prediction
EP3639519B1
Method for encoding / decoding image signal, and device for same
EP3860123B1
Method for refining a motion vector derived under a merge mode using a difference vector
US11457235B2