Video coding in a merge motion vector difference mode utilizing dynamically generated probabilities
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITAL CE PATENT HOLDINGS SAS
- Filing Date
- 2024-06-24
- Publication Date
- 2026-05-13
AI Technical Summary
The computational cost of testing a large refinement vector set in the merge motion vector difference (MMVD) mode for finding the most accurate refined base motion vector is high, which hampers efficient video coding in versatile video coding (VVC) standards.
Adaptive modification of refinement vectors based on dynamically generated probability values, reducing the search space by modifying the distance and direction tables, and reordering or shifting elements to accelerate the encoding process and reduce bitrate.
This approach reduces computational complexity and bitrate while maintaining accurate motion refinement, enhancing video coding efficiency in MMVD mode by dynamically adjusting the refinement vector set based on probability metrics.
Smart Images

Figure EP2024067620_09012025_PF_FP_ABST
Abstract
Description
VIDEO CODING IN A MERGE MOTION VECTOR DIFFERENCE MODE UTILIZING DYNAMICALLY GENERATED PROBABILITIESCROSS REFERENCE TO RELATED APPLICATIONS[1] This application claims the benefit of European Application No. 23306121.7, filed on July 3, 2023, which is incorporated herein by reference in its entirety.BACKGROUND[2] In the recent versatile video coding (VVC) standard and enhanced compression model (ECM), several enhancements have been made to the coding of motion associated with a video block. Particularly, a merge motion vector difference (MMVD) mode has been introduced to enhance motion coding in a merge mode. Operating in an MMVD mode allows for the refinement of a motion vector (namely, a base motion vector) derived in a merge mode. The refined base motion vector can then be used for motion compensated prediction of the video block. The best refinement vector, selected from a refinement vector set, is used to refine a base motion vector, selected from motion vectors of respective merge video block candidates. The larger the refinement vector set is and the larger the number of base motion vectors being tested is, the more accurate the refined base motion vector that can be found will be. However, it is computationally costly to test a large refinement vector set for each of the base motion vectors to find the refinement vector that results in the most accurate refined base motion vector.SUMMARY[3] Aspects disclosed in the present disclosure describe methods for encoding video data. The methods include coding a video block of the video data into a bitstream. The coding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values and selecting a refinement vector from the modified set of refinement vectors, where the selected refinement vector is applied to refine a base motion vector. An index indicating the selected refinement vector is then coded into the bitstream. Aspects disclosed in the present disclosure also describe methods for decoding the video data. The methods include decoding a video block of the video data from the bitstream. The decoding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values. The decoding of the video block further includesdecoding from the bitstream an index indicating a refinement vector from the modified set of refinement vectors. Based on the index, the refinement vector is extracted from the modified set of refinement vectors.[4] Aspects disclosed in the present disclosure describe an apparatus for encoding video data. The apparatus comprises at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatus to code a video block of the video data into a bitstream. The coding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values and selecting a refinement vector from the modified set of refinement vectors, where the selected refinement vector is applied to refine a base motion vector. An index indicating the selected refinement vector is then coded into the bitstream. Aspects disclosed in the present disclosure also describe an apparatus for decoding video data. The apparatus comprises at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatus to decode a video block of the video data from the bitstream. The decoding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values. The decoding of the video block further includes decoding from the bitstream an index indicating a refinement vector from the modified set of refinement vectors. Based on the index, the refinement vector is extracted from the modified set of refinement vectors.[5] Further aspects disclosed in the present disclosure describe a non-transitory computer- readable medium comprising instructions executable by at least one processor to perform methods for encoding video data. The methods include coding a video block of the video data into a bitstream. The coding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values and selecting a refinement vector from the modified set of refinement vectors, where the selected refinement vector is applied to refine a base motion vector. An index indicating the selected refinement vector is then coded into the bitstream. Further aspects disclosed in the present disclosure also describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for encoding video data. The methods include decoding a video block of the video data from the bitstream. The decoding of the video block comprises adaptively modifying a set of refinement vectors based on respective probability values. The decoding of the video block further includes decoding from the bitstream an index indicating a refinement vector from the modified set of refinementvectors. Based on the index, the refinement vector is extracted from the modified set of refinement vectors.[6] This Summary is provided to introduce a selection of concepts in a simplified form that is further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure.BRIEF DESCRIPTION OF THE DRAWINGS[7] FIG. 1 is a block diagram of an example system, according to which aspects of the present embodiments can be implemented.[8] FIG. 2 is a functional block diagram of an example video encoder, according to which aspects of the present embodiments can be implemented.[9] FIG. 3 is a functional block diagram of an example video decoder, according to which aspects of the present embodiments can be implemented.
[0010] FIG. 4 is a diagram illustrating an example of translational modeling of motion associated with a video block, according to which aspects of the present embodiments can be implemented.
[0011] FIG. 5 is a diagram illustrating an example refinement set used in an MMVD mode, according to which aspects of the present embodiments can be implemented.
[0012] FIG. 6 is a flowchart illustrating an example method for modifying motion refinement data, according to which aspects of the present embodiments can be implemented.
[0013] FIG. 7 is a diagram illustrating the operation of a motion refinement data modifier, according to which aspects of the present embodiments can be implemented.
[0014] FIG. 8 is a flowchart of an example method for encoding video data, according to which aspects of the present embodiments can be implemented.
[0015] FIG. 9 is a flowchart of an example method for decoding video data, according to which aspects of the present embodiments can be implemented.DETAILED DESCRIPTION
[0016] Apparatuses and methods are presented herein for encoding and decoding of video data. According to aspects, the coding of motion associated with a video block is improved by adaptively modifying the manner in which motion refinement is encoded in an MMVD mode. To that end, a set of refinement vectors can be adaptively modified based on related probabilities that are dynamically generated. Likewise, a merge candidate list (from which a motion vector is to be refined) can be adaptively modified based on related probabilities that are dynamically generated. Traditional systems and methods for predictive video coding are described next in reference to FIGS. 1-3, followed by a description of aspects of the present disclosure, described in reference to FIGS. 4-9.
[0017] FIG. 1 illustrates a block diagram of an example system 100. System 100 can be embodied as a device and can be configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, can be embodied in an integrated circuit, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, the processing 110 and encoder / decoder 130 elements of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports.
[0018] The system 100 includes at least one processor 110 that can be configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 can include embedded memory, input and output interfaces, and various other circuitries as known in the art. The system 100 includes at least one memory 120, such as a volatile memory device and / or a non-volatile memory device. System 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 can be an internal storage device, an attached storage device, and / or a network accessible storage device, for example.
[0019] System 100 includes an encoder / decoder module 130 configured to process data to provide encoded video data or decoded video data. The encoder / decoder module 130 caninclude its own processor and memory. The encoder / decoder module 130 can be implemented as a separate element of system 100 or can be incorporated within processor 110 as a combination of hardware and / or software as known to those skilled in the art. Additionally, the encoder / decoder module 130 represents module(s) that can be implemented in a separate device to perform encoding and / or decoding functions.
[0020] Program code that is to be loaded into processor 110 or into encoder / decoder 130 to perform the various aspects described in this application can be stored in a storage device 140 and subsequently loaded into memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 can store one or more of various items during the performance of the processes described in this application. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, operational logic, and intermediate or final results from the processing of equations, formulas, operations.
[0021] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing functions that are needed during encoding or decoding. In other embodiments, however, memory external to the processing device (where, for example, the processing device can be either the processor 110 or the encoder / decoder module 130) can be used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140 that may comprise, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations.
[0022] The input to the elements of system 100 can be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal (COMP), (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0023] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion can be associatedwith elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select, for example, a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs some of these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to a baseband. In one set-top box embodiment, the RF portion and its associated input processing element receive an RF signal transmitted over a wired (for example, cable) medium, and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Added elements can include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0024] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, can be implemented, for example, within a separate input processing integrated circuit or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface integrated circuits or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.
[0025] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0026] The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 can include, but is not limited to, a modem or network card. The communication channel 190 can be implemented, for example, within a wired and / or a wireless medium.
[0027] Data can be streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signal of these embodiments is received over the communication channel 190 and the communication interface 150 which can be adapted for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. In other embodiments, data can be streamed to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105 or data can be streamed to the system 100 using the RF connection of the input block 105.
[0028] The system 100 can provide an output signal to various output devices, including a display device 165, an audio device (e.g., speaker(s)) 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display device 165, the audio device 175, or the other peripheral devices 185 using signaling such as AV.link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 using the communication channel 190 via the communication interface 150. The display device 165 and the audio device 175 can be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0029] Alternatively, the display device 165 and the audio device 175 can be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display device 165 and the audiodevice 175 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0030] FIG. 2 illustrates a functional block diagram of an example video encoder 200. The video encoder 200 can be employed by the system 100 described in reference to FIG. 1. For example, the video encoder 200 can be an encoder that operates according to coding standards such as Advanced Video Coding (AVC, H.264 / MPEG-4 | ISO / IEC 14496-10), High Efficiency Video Coding (HEVC, ITU-T H.265 | ISO / IEC 23008-2), or VVC (Standard ITU- T H.266, ISO / IEC 23090-3, 2020).
[0031] Prior to undergoing encoding, the video data can be pre-processed by a precoding processor (not shown). Such pre-processing can include applying a color model transform to the color components of the input video frames (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0) or mapping the color components of the input video frames to obtain a signal distribution that is more resilient to compression (for instance, applying a histogram equalizer and / or a denoising filter to one or more of the video frames’ color components). The preprocessing can also include associating metadata with the video data that can be attached to the coded video bitstream.
[0032] In the encoder 200, a video frame is encoded by the encoder elements as generally described below. A picture (frame) of the original video to be encoded is partitioned into coding units (namely, original blocks) by an image partitioner 202. Typically, a coding unit (CU) contains a luminance block and respective chroma blocks, and so, generally, operations described herein as applied to a CU are applied to the luminance block and to the respective chroma blocks. Following partition 202, each CU can be encoded using an intra-prediction mode or an inter-prediction mode. In an intra-prediction mode, a prediction of the CU is performed by an intra-predictor 260. In the intra-prediction mode, the content of a CU in a frame is predicted based on content from one or more other CUs of the same frame, using the other CUs’ reconstructed version (available from the adder 255 output). In an inter-prediction mode, motion estimation and motion compensation are performed by a motion estimator 275 and a motion compensator 270, respectively. In the inter-prediction mode, the content of a CU in a frame is predicted based on content from one or more other CUs of neighboring frames, using the other CUs’ reconstructed versions (available from the reference picture buffer 280). The encoder decides 205 which prediction result (one obtained through operations in the intraprediction mode 260 or one obtained through operations in the inter-prediction mode 270, 275) to use for encoding a CU, and indicates the selected prediction mode by a prediction modeflag, for example. The selected prediction result may then be enhanced (e.g., filtered) by a prediction enhancer 285, outputting a respective prediction block. Once a prediction block is generated for each CU, a respective residual block is calculated, for example, by subtracting 210 the predicted CU (i.e., prediction block) from the CU (i.e., original block).
[0033] A CU’s respective residual block or a partition thereof (i.e., a transform block) is then transformed into a coefficient block by a transformer 220 - that is, residual samples of the transform block are transformed into transform coefficients of the coefficient block. The resulting coefficient block is quantized by a quantizer 230. An entropy encoder 245 is next employed to entropy-encode the quantized coefficient block and respective coding parameters (e.g., syntax elements including motion vectors and other control data). Hence, the entropy- encoded quantized coefficient blocks and respective encoding parameters associated with each video frame of the original video are packed into the bitstream of the coded video data.
[0034] Along with the coding of original blocks (CUs), as described above, the encoder 200 reconstructs the coded original blocks to provide references for future predictions. Accordingly, quantized coefficient blocks (provided by the quantizer 230) are de-quantized, by an inverse quantizer 240, and then inverse transformed, by an inverse transformer 250, to reconstruct (decode) the residual blocks of respective original blocks. Adding 255 the reconstructed residual blocks to respective prediction blocks results in respective reconstructed original blocks. In-loop filters 265 can then be applied to the reconstructed picture (formed by the reconstructed original blocks), performing, for example, deblocking filtering and / or sample adaptive offset (SAO) filtering to reduce encoding artifacts. The filtered reconstructed picture can then be stored in the reference picture buffer 280, available for future predictions in an inter-prediction mode. Thus, the encoder 200 also performs decoding operations 240, 250 through which the encoded pictures (frames) are reconstructed. The reconstructed pictures can then be stored in the reference picture buffer 280 and be used to facilitate motion estimation 275 and compensation 270, as explained above.
[0035] FIG. 3 illustrates a functional block diagram of an example video decoder 300. The video decoder 300 can be employed by the system 100 described in reference to FIG. 1. Generally, operational aspects of the video decoder 300 are reciprocal to operational aspects of the video encoder 200. In the decoder 300, the bitstream of coded video data, generated by the video encoder 200, is first entropy-decoded by an entropy decoder 330, decoding from the bitstream the quantized coefficient blocks and various coding parameters. The quantized coefficient blocks are de-quantized, by an inverse quantizer 340, and then are inversetransformed, by an inverse transformer 350, to decode (reconstruct) respective residual blocks. Adding 355 the reconstructed residual blocks to respective prediction blocks results in respective reconstructed original blocks. Depending on the selected prediction mode, a predicted original block can be obtained 370 from an intra-predictor 360 or from a motion compensator 375 and may then be enhanced (e.g., filtered) by a prediction enhancer 390, generating a prediction block. In-loop filters 365 can be applied to the reconstructed picture (formed by the reconstructed original blocks), outputting a reconstructed (decoded) video frame. The filtered reconstructed picture is also stored in a reference picture buffer 380 to facilitate motion compensation 375.
[0036] A post-decoding processor (not shown) can further process the reconstructed video. For example, post-decoding processing can include an inverse color model transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse mapping to reverse the mapping process performed by the pre-encoding processor. The post-decoding processor can use metadata that were derived by the pre-encoding processor and / or were signaled in the video bitstream.
[0037] Aspects disclosed herein are described in reference to a video block for which interprediction or intra-prediction can be applied. However, the described aspects are similarly applicable to any region of the video frame to which coding tools may be applied to by an encoder 200 or by a decoder 300. Generally, aspects described herein may be applied to a video data region, formed by a video partition, of any shape or size. The video region may be a CU or a partition thereof, including a luma component, F, and chroma components, Cr and Cb.
[0038] Aspects described herein improve the video coding efficiency of video compression systems (e.g., the systems described with respect to FIGS. 1-3). According to aspects, techniques are provided herein to code motion in an MMVD mode utilizing probabilities that are dynamically generated, for example, by context-based adaptive binary arithmetic coding (CAB AC). The MMVD mode, adopted by the VVC standard, offers a refinement of the motion vector provided when operating in a merge mode (or a merge and skip mode) used in the HEVC standard, as further described below.
[0039] As defined by the HEVC standard, motion data coding (in an inter-prediction mode) can be performed in two modes of operation: an advanced motion vector prediction (AMVP) mode and a merge mode (or a merge and skip mode). Both operational modes are designed to exploit spatial and temporal correlations among motion vectors of adjacent video blocks. InHEVC, as in prior standards, a motion vector associated with a video block to be encoded (referred to herein also as the current video block) points to a matching (reference) video block from a reference picture. That reference video block is used to predict the video block (serving as its respective prediction block), as explained above with respect to the encoder 200 of FIG. 2 and the decoder 300 of FIG. 3. In an AMVP mode, a motion vector of the current video block is coded relative to a motion vector of another video block in the spatiotemporal neighborhood of the current video block, as further explained with respect to FIG. 4.
[0040] FIG. 4 is a diagram illustrating an example of translational modeling of motion associated with a video block 400. As illustrated, the motion of content within a video block 455 (from a current picture 450) can be represented by a motion vector 440 that points to a reference block 435 (from a reconstructed reference picture 430). The pointed-to reference block 435 has content that matches the content of the video block 455. In predictive coding, such a reference block 435 can be used to predict the video block 455 and constitutes a respective motion compensated prediction block. Generally, more than one reference block can be used for prediction of a video block. For example, a second motion vector 420 that points to a second reference block 415 (from a second reference picture 410) can be determined and a combination of the two reference blocks 415, 435 can be used to predict the video block 455. Thus, to encode the video block 455 in an inter-prediction mode of operation, the motion vector(s) 420, 440 and the respective reference picture(s) 410, 430 (i.e., the motion information) as well as the residual (i.e., the prediction error provided by the difference between the original video block and its prediction) are encoded into the bitstream.
[0041] Since motion vectors of adjacent video blocks are correlated (especially when the motion that these motion vectors represent is a result of a rigid-object’s motion or the camera’s motion), one motion vector can efficiently predict another. Thus, a video block’s motion vector can be represented relative to a neighboring block’s motion vector that serves as its predictor, namely, a motion vector predictor (MVP). When coding a motion vector of a video block, what is signaled in the bitstream is a respective motion vector difference (MVD):MVDx = dx - MVPx (1)MVDy = dy - MVPy, (2) where (dx, dy) are the motion vector components (e.g., as illustrated in FIG. 4), MVPx, MVPy) are the motion vector predictor components, and MVDx, MVDy) are the motion vector difference components. To decode a motion vector (i.e., (dx, dy)), the decodercan extract from the bitstream the MVD (i.e., (MVDx, MV Dy ) and infer (i.e., independently derive) the corresponding MVP (i.e., (MVPx, MVPy)). In the AVC standard, the MVP is derived from the median of three motion vectors of predetermined blocks that are spatially adjacent to the video block (thus, no signaling of the MVP is needed). However, in the HEVC standard, the MVP is explicitly signaled in the bitstream as being one of a list of candidate MVPs. To that end, the HEVC standard defines a process in which a list of candidate MVPs is constructed (that for simplicity includes two MVP candidates of two blocks situated in the spatiotemporal neighborhood of the video block). This MVP candidate list is constructed independently by the encoder and the decoder. Thus, in HEVC, to code the motion of a video block in an AMVP mode, the encoder signals the motion information that in this case include: an index to the reference picture, denoted refldx, an index to the MVP candidate list, denoted mvpldx, and a motion vector difference MVD. In turn, the decoder: 1) extracts the MVP from the MVP candidate list using the signaled mvpldx, 2) computes the motion vector (dx, dy) based on the signaled MVD and the extracted MVP (according to equations 1-2); and 3) accesses the reference block pointed-to by the computed motion vector in the reference picture indicated by the signaled refldx. The accessed reference block is then used by the decoder to perform motion compensated prediction of the corresponding video block.
[0042] To reduce the bitrate associated with signaling motion information of motion that is coded in an AMVP mode, the motion can instead be coded in a merge mode. In a merge mode, the decoder can perform motion compensated prediction of a video block based on the decoded motion information of a neighboring block (namely, a merge video block). In this case, the only information that needs to be signaled by the encoder is a pointer (an index) to the merge video block whose motion information should be used. In the HEVC standard, the encoder constructs a list of five merge candidates and signals an index, mergeldx, to this list to indicate the merge candidate (i.e., the merge video block) whose motion information should be used for the current video block’s motion compensated prediction. Note that in an AMVP mode only a motion vector associated with a candidate of the MVP candidate list is used, while in a merge mode all motion information (needed to perform motion compensated prediction) associated with a candidate in the merge candidate list is used.
[0043] Hence, when a merge mode is used to code the motion of a video block, the encoder needs to merely signal an index, denoted mergeldx, to the merge candidate list. The decoder constructs the merge candidate list in the same manner as is constructed by the encoder. Based on a decoded index mergeldx) that is associated with a current video block, the decoder usesthe already decoded motion information of a video block that is pointed to by that index in the merge candidate list.
[0044] In the HEVC standard, a merge mode can be combined with a skip mode (i.e., a merge and skip mode). In a merge and skip mode, the prediction error (i.e., the residual) is not coded into the bitstream (e.g., by the transformer 220, quantizer 230, and entropy encoder 245 of FIG. 2), and so reconstruction of a video block is performed only based on its motion compensated prediction. Generally, a merge and skip mode is most suitable to code motion of content within video blocks coming from static regions of the pictures. Note that aspects described herein with respect to a merge mode are similarly applicable to a merge and skip mode.
[0045] In the recent VVC standard and ECM, several enhancements have been made to the coding of motion in an inter-prediction mode relative to the HEVC standard. The introduced enhancements include extensions to the AMVP and the merge modes (see, W. J. Chien et al., "Motion Vector Coding and Block Merging in the Versatile Video Coding Standard," in IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 10, pp. 3848-3861, Oct. 2021, doi: 10.1109 / TCSVT.2021.3101212). Relevant herein, is the introduction of the MMVD mode that allows for refinement of motion vectors derived in a merge mode (see, S. Jeong et al., “Merge mode with motion vector difference,” in Proc. IEEE Int. Conf. Image Process. (ICIP), Abu Dhabi, United Arab Emirates, Oct. 2020, pp. 1157-1160, and S. Jeong et al., "CE4 Ultimate motion vector expression (Test 4.5.4)," JVET-L0054, 12th JVET meeting: Macao, CN, 3-12 Oct. 2018). The MMVD mode (applied in combination with the merge mode or with the merge and skip mode) is a tradeoff between motion coding accuracy and low bitrate - it can be viewed as an intermediate approach to motion coding situated between the merge mode and the AMVP mode. In an MMVD mode, one of the first two candidates in the merge candidate list is selected to be used as a base motion vector. That base motion vector is adjusted by a refinement vector, as illustrated in FIG. 5.
[0046] FIG. 5 is a diagram illustrating an example refinement set used in an MMVD mode 500. The refinement set 500 includes refinement vectors that define respective refinement positions (denoted by white circles) relative to a position 510 (denoted by a black circle) of a base motion vector. The base motion vector is a decoded motion vector of a merge video block candidate selected from the merge candidate list for the current video block. Each refinement vector (or refinement step) is represented by its magnitude value (i.e., a distance between the respective refinement position and position 510 pointed to by the base motion vector) and adirection value. For example, four refinement positions 520 are shown to be positioned in increasing distances in a direction of 0 = 0 degrees. Likewise, four refinement positions 530 are shown to be positioned in increasing distances in a direction of 0 = - degrees. Generally, a predetermined number of refinement positions (defined by combinations of distances and directions) are tested for each base motion vector (of a merge video block candidate) by the encoder, and the refinement position that leads to the best refined base motion vector is selected to perform the motion compensated prediction for the current video block.
[0047] Hence, when the merge and the MMVD modes are used to code the motion of a video block, the base motion vector (the motion vector of a merge video block) is adjusted by a refinement vector. The refinement vector is determined based on a direction value and a distance value indexed, respectively, in a direction table and in a distance table. For example, in the VVC standard, four directions (including directions in 0 = n ■ - degrees, for n = 0, ... ,3) and eight distances provided by two distance tables (including distances { 1 / 4, 1 / 2, 1, 2, 4, 8, 16, 32} and { 1, 2, 4, 8, 16, 32, 64, 128} in pixel units) are defined. The encoder can select what distance table is used at a picture level. The merge and the MMVD modes are also offered in the ECM. Therein, up to 16 directions (0 = n ■ - for n = 0, ... ,15) are defined. A similar87T concept is used in ECM when using an Affine MMVD, where four directions (0 = n ■ - for n = 0, ... ,3) and 5 distances depending on the frame size ({ 1 / 16, 1 / 8, 14, 14,1} for frame size < 720p, { 1 / 4, 14, 1, 2, 4} for frame size between 720p to 1600p, and { 1 / 2, 1, 2, 4 8} for frame size >= 1600p) are used. Thus, to code the motion of a video block in the merge and MMVD modes, in addition to the merge candidate index (mergeldx) that indicates the base motion vector, a distance index, denoted distldx, and a direction index, denoted dirldx, are signaled to indicate a corresponding refinement vector.
[0048] Signaling motion information in an AMVP mode requires more bits compared to signaling motion in a merge mode. As discussed above, in an AMVP mode, to code the motion of a video block, the reference picture index (refldx), the MVP candidate list index (mvpldx), and the motion video difference (MVDx, MV Dy) should be signaled for each reference picture used. Where to code the motion of the video block in a merge mode only the merge candidate list index (mergeldx) should be signaled. However, when an MMVD mode is applied to code the motion of the video block in a merge mode, the distance index (distldx) and the direction index dirldx have to be signaled in addition to the merge candidate list index (mergeldx). Thus, an MMVD mode can be considered as a tradeoff between motion information accuracy(associated with an AMVP mode) and low bitrate representation (associated with a merge mode).
[0049] Hence, when an MMVD mode is used, the base motion vector is obtained from a merge candidate selected out of the merge candidate list (i.e., selected out of the first of two candidates in VVC and selected out of the first of up to six candidates in ECM). This base motion vector is refined by a selected refinement position (e.g., one of the refinement positions illustrated in the example of FIG. 5). The selected refinement position is indicated by a distance index and by a direction index, respectively, indexing a distance table and a direction table. Table 1 and Table 2 show, respectively, the distance table and the direction table that are defined in VVC. The distance is coded using truncated unary coding to allow efficient coding of distances that are most frequently used; and the direction is coded using fixed length coding. In addition to signaling the distance index (distldx) and the direction index (dirldx), a merge candidate index (mergeldx) is signaled to indicate one of two merge candidates. The motion vector of the indicated merge candidate is used as the base motion vector that is then adjusted by the refinement vector (defined by distldx and dirldx).
[0050] Table 1 : MMVD Distance Table
[0051] Table 2: MMVD Direction Table
[0052] Different distance tables that are applicable for different types of content can be used (see, H. Liu, et al., "AHG11 : MMVD without Fractional Distances for SCC," JVET-M0255, 13th JVET meeting: Marrakech, Morocco, 9-18 Jan. 2019, hereinafter “Liu”). In Liu, integer distances are proposed to be used for screen content coding (SCC) because, as shown therein,using fractional distances in the case of SCC may be ineffective. Table 3 shows the distance tables that can be used in VVC for different scene categories.
[0053] Table 3 : MMVD Distance Tables for Different Scene Categories
[0054] To perform motion compensated prediction in an MMVD mode, the encoder searches for the refinement vector that refines the base motion vector into a vector that points to a reference block (in the reference picture) that best represents the currently encoded video block. Thus, the encoder needs to search for that best refinement vector in a predetermined space of refinement positions for each merge video block candidate. The space of refinement positions is of a size equal to the number of distances times the number of directions. For example, in VVC, the searching space can reach a size of 32 (as shown in Tables 1 and 2, that is, 8 times 4 refinement positions). Likewise, in ECM, the searching space can reach a size of 96 refinement positions.
[0055] Finding the best (optimal) refinement position (for each merge video block candidate) requires an exhaustive search that is time consuming, especially when a large number of refinement positions is used. However, it is possible to simplify the search process by restricting the searching space and thereby reducing complexity in the encoder. Furthermore, the distance table, the direction table, and / or the list of merge candidates can be dynamically modified, for example, by reducing, re-ordering, and / or shifting the elements in the original table or list. As a result, the indexed elements (e.g., a subset of elements at the beginning of the table or list) can be reduced and / or can be changed at the block level or the slice / picture level. Consequently, a reduced bitrate can be maintained while effectively allowing the indexing of more elements overall.
[0056] Several variations to the MMVD mode have been proposed, including the following:• extending the direction table to include diagonal directions (see, T. Hashimoto, et al., "CE4-3.4: Diagonal direction candidate for MMVD," JVET-N0142, 14th JVET meeting: Geneva, Switzerland, 19-27 Mar. 2019; and Y. Kidani, et al., "AHG12: Diagonal MMVD with ARMC," JVET-W0112, 23rd JVET meeting: Online, 07-16 Jul 2021);• reordering of the merge candidate list or reordering of the refinement set based on template matching (see, Y. Kidani, et al. " AHG12: Diagonal MMVD with ARMC,"JVET-W0112, 23rd JVET meeting: Online, 07-16 Jul 2021; and F. Galpin, et al. "non-CE4 flexible MMVD candidates," JVET-P0384, 16th JVET meeting: Geneva, 01-11 Oct 2019);• signaling and / or reordering of the distance table at slice or picture level (see, J. Li, et al., "CE4-related: MMVD improving with signaling distance table," JVET-M0314, 13th JVET meeting: Marrakech, Morocco, 9-18 Jan. 2019; and S. Jeong, et al., "CE4- 3.1 : MMVD binarization," JVET-N0125, 14th JVET meeting: Geneva, Switzerland, 19-27 Mar. 2019);• using all the merge candidates instead of only the first two (see, F. Galpin, et al., "Non- CE4 flexible MMVD candidates," JVET-P0384, 16th JVET meeting: Geneva, 01-11 Oct 2019); and• reordering of refinement positions based on template costs (see, M. Salehifar, et al., "Non-EE2: Template Matching-based Reordering for Extended MMVD Design," JVET-X0085, 24th JVET meeting: Online, 06-15 Oct 2021).These variations aim to either reduce the cost of signaling indices associated with an MMVD mode or to enhance the accuracy of coding the motion in the MMVD mode by increasing the searched refinement space. However, often, these variations add more computational complexity on both the encoder and the decoder ends.
[0057] As described in reference to FIG. 2 and FIG. 3, in addition to prediction of video blocks (either in inter-prediction mode or in intra-prediction mode) predictive video coding generally includes generating respective residuals (i.e., prediction errors). Transform-based coding is used to code these residuals 220, 230, resulting in quantized transform coefficients. The quantized transform coefficients and coding parameters (e.g., motion information) - typically represented by non-binary symbols - are then entropy-based coded 245 into a bitstream. A CABAC is an entropy-based coding technique used in the VVC standard (and in the AVC and the HEVC standards) to code non-binary symbols (see, Marpe, D., et al., “Context-Based Adaptive Binary Arithmetic Coding in the H.264 / AVC Video Compression Standard,” IEEE Trans. Circuits and Systems for Video Technology, Vol. 13, No. 7, pp. 620-636, July, 2003). To code a symbol, a CABAC engine combines binary arithmetic coding with context modeling. The former uses a probability model (the underlying probability distribution) of a symbol to efficiently code the symbol and the latter updates that probability model according to changesin the statistics of the symbol that may occurred during the coding of the video. Thus, probabilities - namely, mState values - of respective symbols (generated by the encoder to represent a video stream being encoded) are computed and adapted overtime by the CAB AC engine during the coding of the video stream. In order to ensure the correct decoding of the symbols, the respective mState values are fully synchronized and available at the encoder side and at the decoder side. For example, when coding motion in an MMVD mode, the CAB AC engine codes (at the encoder end) and decodes (at the decoder end) indices such as distldx, dirldx, and mergeldx based on their respective mState values that are being dynamically generated and updated in the same manner at both ends.
[0058] Aspects described herein utilize the mState values (or other probability metrics) associated with symbols generated by motion coding - such as indices generated by motion coding in an MMVD mode - to improve motion coding efficiency. According to aspects, the refinement of a base motion vector and the representation of related symbols can be performed based on internal CABAC state of related probability models. Thus, aspects disclosed herein can be used to accelerate the encoding time (e.g., to reduce the time spent on searching for a refinement vector) and / or to reduce the signaling overhead (e.g., to reduce the bitrate used to signal symbols associated with coding of motion refinement).
[0059] As mentioned above, operating in an MMVD mode requires searching for the best refinement position out of P • A possible steps with respect to a given base motion vector, where P denotes the number of distances and A denotes the number of directions (angles). According to aspects, a modified distance table can be determined. The modified distance table can be: 1) a reduced distance table with k candidates where k < P 2) a reordered distance table; and / or 3) a shifted distance table where distance values are changed by shifting the distance values by a shift factor sdstand / or adding an offset. The modified distance table can also be a completely new distance table. When the probability, as captured for example by a respective mState value, of a small distance value is high, it reflects that the content of the corresponding video block changes slowly (low motion region) and hence using smaller distance values (e.g., fractional refinement positions) should be favored. On the other hand, if this probability is low, then using larger distance values (e.g., integer refinement positions) should be favored. Alternatively, or in addition, just as with the distance table, a modified direction table can be determined. The modified direction table can be 1) a reduced direction table with n candidates where n < A; 2) a reordered direction table; and / or 3) a shifted direction table where direction values are changed by shifting the direction values by a shiftfactor sdirand / or adding an offset. The modified direction table can also be a completely new direction table. When the probability, as captured for example by a respective mState value, of a certain direction value is high, it reflects that content of the corresponding video block changes in that direction and hence nearby direction values should be favored. The above approach is further explained with respect to FIG. 6.
[0060] FIG. 6 is a flowchart illustrating an example method for modifying motion refinement data 600. According to aspects, the motion refinement data can include refinement vectors and merge video block candidates. The method 600 begins, in step 610, by deriving a motion activity indicator (MAI). Generally, the MAI can be designed to characterize the motion of the content in the spatiotemporal neighborhood of the current video block. The MAI can be predicted by the encoder based on dynamic coding data generated and adapted during the encoding of a video block and / or its spatiotemporal neighborhood. Similarly, the MAI can be predicted by the decoder based on the same dynamic coding data generated and adapted during the decoding of the video block and / or its spatiotemporal neighborhood. In an aspect, the dynamic coding data can include coding parameters such as quantization parameters (QP) and / or spatial dimensions associated with the video block and its neighboring blocks. In a further aspect, the dynamic coding data can include probability data. The probability data can be related to symbols generated in conjunction with motion coding. In this aspect, mState values dynamically generated by the CAB AC engine (in both encoder and decoder ends) can be used to derive the MAI. The derived MAI is then compared to a threshold T in step 620. If the MAI is above the threshold T, in step 630, motion refinement data are modified as further explained below.
[0061] The threshold T can represent a set of thresholds that may be compared 620 to respective MAI values (derived in step 610). The threshold(s) can be fixed threshold(s), for example, empirically determined based on a training dataset. The threshold(s) can also be explicitly signaled in a picture header or in a slice header when method 600 for modifying the motion refinement data is applied. For example, a flag can be signaled at a high-level to indicate whether the threshold(s) are explicitly signaled or not. In a case where this flag is set to 0, default threshold value(s) (e.g., determined empirically) can be used.
[0062] The MAI can be derived using an MAI predictor that is modeled based on dynamic coding data. In an aspect, the MAI predictor is a linear combination of features expressed as follows:MAI = a ■ Ps + (3 ■ QP + y ■ size + 6, (3) where Ps is a probability feature, QP is a quantization parameter feature, and size is a feature associated with a spatial dimension. The MAI predictor is determined by estimation parameters (weights): a, fl, y, and 8 . The Ps feature may represent one or more mState values of respective symbols (e.g., distldx, dirldx, and / or mergeldx). The mState values are probabilities that can be dynamically generated by the CABAC engine. The QP feature can represent the quantization parameter(s) signaled with respect to the video block and / or its neighboring blocks. The size feature can represent a dimension (e.g., an area) of the video block and / or dimensions of its neighboring blocks. In an aspect, the features used by the MAI predictor can be empirically obtained based on a training dataset. During the encoding process, if a feature is not available (or otherwise is not used), the corresponding weight can be set to 0.
[0063] For example, when the Ps feature represents the mState value of distldx, method 600 can be applied to determine a possible modification of the distance table. Alternatively, or in addition, when the Ps feature represents the mState value of dirldx, method 600 can be applied to determine a possible modification of the direction table. Likewise, when the Ps feature represents the mState value of mergeldx, method 600 can be applied to determine a possible modification of the merge candidate list. Such modifications 630 may include reducing, reordering, and / or shifting a respective table / list.
[0064] In an aspect, modifying the motion refinement data 630 can include reducing the number of distances, directions, merge candidates, or a combination thereof to be tested by the encoder in an MMVD mode, leading to a reduced searching space that, in turn, reduces related computational complexity and thus the overall encoding time.
[0065] FIG. 7 is a diagram illustrating the operation of a motion refinement data modifier 700. As shown in the example of FIG. 7, an encoder 710 and a decoder 730, communicatively linked 720, are each configured to employ a motion refinement data modifier 715, 735. Both modifiers 715, 735 can apply aspects of method 600 for modifying the motion refinement data. Since features needed to compute the MAI (e.g., see equation (3)) are available at both the encoder 710 and the decoder 730 ends, the modifier 735 at the decoder can determine whether and how the motion refinement data has been modified by the modifier 715 at the encoder. That is, the decoder 730, after constructing the distance table, the direction table, and the merge candidate list, can apply method 600 to modify these distance table, direction table, and mergecandidate list in the same manner as done by the encoder when applying method 600.
[0066] Hence, a default distance table, direction table and merge candidate list can be initially constructed by the encoder and by the decoder in the same manner and independently. These initial tables and list can be modified 630 based on dynamic coding data, as explained above. Indices to the modified distance table, direction table, and / or merge candidate list can be explicitly signaled in the bitstream. When the modification of the motion refinement data includes reducing the number of the indexed elements in the distance / direction table or merge candidate list, such a modification leads to an encoder speed up (due to a reduced searching space) and also to a reduced signaling bitrate (as fewer bits are required to signal the reduced number of elements being indexed). For example, if by default only the first four elements of the distance table are indexed and the modification includes reducing the indexed number of elements to two, then less bits will be required to signal the symbol distldx (see Table 1). When the modification of the motion refinement data includes reordering the elements of the distance table, direction table, and / or merge candidate list, the most probable elements can be positioned first. For example, if by default only the first four distance values in the distance table are indexed out of a maximum of eight distance values (as in Table 1), the most probable distance values can be dynamically placed at the beginning of the distance table.
[0067] In another aspect, a high-level syntax element (e.g., an SPS flag) can be signaled by the encoder to indicate that a table or a list has been modified (e.g., by the motion refinement data modifier 715 according to method 600). Based on such a signal, the same modification will be generated by the decoder (e.g., by the motion refinement data modifier 735 according to method 600). According to another aspect, the application of method 600 to modify the motion refinement coding can be activated / deactivated at the picture level, through a flag coded in a picture header (or alternatively in a slice header). In a variant, a flag can be coded (e.g., at a picture level, a slice level, or an SPS / PPS level) to indicate whether any of the distance table, direction table, and merge candidate list is modified by method 600 or whether the default respective table / list should be used. In another variant, the usage of method 600 is only signaled when the usage of the MMVD tool is enabled in the high-level syntax.
[0068] FIG. 8 is a flowchart of an example method for encoding video data 800. The method 800 is applied to code a video block into a bitstream. Generally, the encoded video block is a partition of a currently encoded video frame (e.g., by encoder 200) of video data that can be obtained via streaming or can be retrieved from storage. The coding of the video block into the bitstream can include operations carried out by steps 810-830, as follows. In step 810, aset of refinement vectors is adaptively modified based on respective probability values. In step 820, a refinement vector from the modified set of refinement vectors is selected. The selected refinement vector is applied to refine a base motion vector (e.g., a motion vector of a merge video block). The refined base motion vector can then be used for motion compensated prediction of the video block, as generally done by predictive coding. In step 830, an index is coded into the bitstream to indicate the selected refinement vector (e.g., the index may be representing dstldx and dirldx). A flag can also be coded into the bitstream to indicate that the set of refinement vectors is modified.
[0069] To adaptively modify the set of refinement vectors, the method 800 can be further applied to derive a motion activity indicator. Generally, the motion activity indicator characterizes motion of content in a neighborhood of the video block. Based on the motion activity indicator, the set of refinement vectors can be modified, if, for example, the motion activity indicator is above a predetermined threshold (e.g., see FIG. 6). According to aspects, the motion activity indicator can be determined based on dynamic coding data. The dynamic coding data can include probability values associated with the set of refinement vectors, a quantization parameter associated with the video block, a size associated with the video block, or a combination thereof (e.g., see equation (3)).
[0070] In an aspect, a refinement vector in the set of refinement vectors is defined by a distance value from a distance table (e.g., see Table 1) and by a direction value from a direction table (e.g., see Table 2). In this aspect, adaptively modifying the set of refinement vectors can include modifying the distance table. For example, modifying the distance table may be by reducing the number of elements of the distance table, by reordering the elements of the distance table, by changing one or more values of the elements of the distance table, or by a combination thereof. Likewise, adaptively modifying the set of refinement vectors can include modifying the direction table. For example, modifying the direction table may be by reducing the number of elements of the direction table, by reordering the elements of the direction table, by changing one or more values of the elements of the direction table, or by a combination thereof.
[0071] In an aspect, the method 800 may be further applied to adaptively modify a list of merge candidates based on one or more of respective probability values, where the merge video block is selected from the modified list of merge candidates. And an index indicating the selected merge video block is coded into the bitstream. In this aspect, a flag can be coded into the bitstream to indicate that the list of merge candidates is modified.
[0072] FIG. 9 is a flowchart of an example method for decoding video data 900. The method 900 is applied to decode a video block from a bitstream. Generally, the decoded video block is a partition of a currently decoded video frame (e.g., by decoder 300) of video data coded into the bitstream (e.g., by encoder 200) and obtained via streaming or retrieved from storage. The decoding of the video block from the bitstream can include operations carried out by steps 910-930, as follows. In step 910, a set of refinement vectors is adaptively modified based on respective probability values, in the same manner performed by the encoder in step 810. In step 920, an index, indicating a refinement vector from the modified set of refinement vectors, is decoded from the bitstream. Based on the index, in step 930, the refinement vector from the modified set of refinement vectors is extracted. The extracted refinement vector is applied to refine a base motion vector (e.g., a motion vector of a merge video block). The refined base motion vector can then be used for motion compensated prediction of the video block, as generally done by predictive coding. In an aspect, a flag indicating whether the set of refinement vectors should be modified is decoded from the bitstream and if that flag indicates that the set of refinement vectors should be modified, then the set of refinement vectors is adaptively modified, generally, in the same manner it is performed at the encoder, as described above.
[0073] We have described several aspects and embodiments in the present disclosure. These aspects and embodiments provide at least the following outputs and results, including all combinations, across different claim categories and types:• Encoding, into coded video data, syntax elements that can enable the decoder to decode the coded video data, according to any of the aspects described herein.• A bitstream that includes one or more of the described syntax elements, or variations thereof. A bitstream can be any set of data whether transmitted, stored, or otherwise made available.• Creating, transmitting, receiving, and / or decoding of the bitstream.• An electronic device (e.g., a TV, a set-top box, a cell phone, or a tablet) that tunes (e.g., using a tuner) a channel to receive the bitstream or that receives (e.g., using an antenna) the bitstream over the air. The electronic device decodes the syntax elements from the bitstream, and, optionally, displays (e.g., using a monitor, screen, or any other type of display) a resulting image.Various other generalized, as well as particularized, outputs, results, implementations, and claims are also supported and contemplated throughout this disclosure.
[0074] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as “first”, “second”, etc. can be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and can occur, for example, before, during, or in an overlapping time period with the second decoding.
[0075] Various methods and other aspects described in this application can be used to modify modules, for example, the modules of the video encoder 200 and the video decoder 300 as shown in FIG. 2 and FIG. 3. Moreover, the present aspects are not limited to a specific standard (such as VVC or HEVC) and can be applied, for example, to other standards and recommendations, as well as extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0076] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0077] Various implementations involve decoding. “Decoding,” as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0078] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video data in order to produce an encoded bitstream. Additionally, the terms “reconstructed” and “decoded” can be used interchangeably,the terms “encoded” or “coded” can be used interchangeably, and the terms “image,” “picture,” and “frame” can be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used on the encoder side while the term “decoded” is used on the decoder side.
[0079] Note that the syntax elements as used herein are descriptive terms. As such, they do not preclude the use of other syntax element names.
[0080] This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example. This information can be packaged or arranged in a variety of manners, including, for example, manners that are common in video standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, or a slice header), or an SEI message. Other manners are also available, including, for example, manners common for system level or application level standards such as signaling the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example, as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. DASH MPD (Media Presentation Description) Descriptors, for example, as used in DASH and transmitted over HTTP. A descriptor is associated with a Representation or collection of Representations to provide additional characteristics to the content Representation. c. RTP header extensions, for example, as used during RTP streaming. d. ISO Base Media File Format, for example, as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length (also known as 'atoms' in some specifications). e. HLS (HTTP live Streaming) manifest transmitted over HTTP. A manifest can be associated, for example, with a version or collection of versions of content to provide the characteristics of the version or collection of versions.
[0081] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, forexample, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (PDAs), and other devices that facilitate communication of information between end-users.
[0082] Reference to “one / an aspect” or “one / an embodiment” or “one / an implementation,” as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the aspect / embodiment / implementation is included in at least one embodiment. Thus, the appearances of the phrase “in one / an aspect” or “in one / an embodiment” or “in one / an implementation,” as well any other variations, appearing in various places throughout this application, are not necessarily all referring to the same embodiment.
[0083] Additionally, this application can refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0084] Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0085] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing,” intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0086] It is to be appreciated that the use of any of the following “ / ”, “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B,” is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example,in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This can be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0087] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a quantization parameter for de-quantization. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual data, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0088] As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
CLAIMS1. A method for encoding video data, comprising: coding, into a bitstream, a video block of the video data, the coding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, selecting a refinement vector from the modified set of refinement vectors, the selected refinement vector is applied to refine a base motion vector, and coding, into the bitstream, an index indicating the selected refinement vector.
2. The method according to claim 1, further comprising: coding, into the bitstream, a flag indicating that the set of refinement vectors is modified.
3. The method according to claim 1 or 2, wherein the adaptively modifying the set of refinement vectors further comprises: deriving a motion activity indicator, characterizing content motion in a neighborhood of the video block; and modifying the set of refinement vectors based on the motion activity indicator.
4. The method according to claim 3, wherein the set of refinement vectors is modified if the motion activity indicator is above a predetermined threshold.
5. The method according to any one of claims 3 to 4, wherein the deriving of the motion activity indicator comprises: determining the motion activity indicator based on dynamic coding data, including one or more of the respective probability values, a quantization parameter associated with the video block, a size associated with the video block, or a combination thereof.
6. The method according to any one of claims 1 to 5, wherein a refinement vector in the set of refinement vectors is defined by a distance value from a distance table and by a direction value from a direction table.
7. The method according to claim 6, wherein the adaptively modifying the set of refinement vectors comprises: modifying the distance table, the modifying includes reducing a number of elements of the distance table, reordering the elements of the distance table, changing one or more values of the elements of the distance table, or a combination thereof.
8. The method according to claim 6, wherein the adaptively modifying the set of refinement vectors comprises: modifying the direction table, the modifying includes reducing a number of elements of the direction table, reordering the elements of the direction table, changing one or more values of the elements of the direction table, or a combination thereof.
9. The method according to any one of claims 1 to 8, further comprising: adaptively modifying a list of merge candidates based on one or more of respective probability values, wherein a merge video block is selected from the modified list of merge candidates, and wherein the base motion vector is a motion vector of the merge video block; and coding, into the bitstream, an index indicating the merge video block.
10. The method according to claim 9, further comprising: coding, into the bitstream, a flag indicating that the list of merge candidates is modified.
11. A method for decoding video data, comprising: decoding, from a bitstream, a video block of the video data, the decoding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, decoding, from the bitstream, an index indicating a refinement vector from the modified set of refinement vectors, and extracting, based on the index, the refinement vector from the modified set of refinement vectors, the extracted refinement vector is applied to refine a base motion vector.
12. The method according to claim 11, further comprising: decoding, from the bitstream, a flag indicating whether the set of refinement vectors should be modified; and performing the adaptively modifying of the set of refinement vectors if the flag indicates that the set of refinement vectors should be modified.
13. The method according to claim 11 or 12, wherein the adaptively modifying the set of refinement vectors further comprises: deriving a motion activity indicator, characterizing content motion in a neighborhood of the video block; and modifying the set of refinement vectors based on the motion activity indicator, wherein the set of refinement vectors is modified if the motion activity indicator is above a predetermined threshold.
14. The method according to claim 13, wherein the deriving of the motion activity indicator comprises: determining the motion activity indicator based on dynamic coding data, including one or more of the respective probability values, a quantization parameter associated with the video block, a size associated with the video block, or a combination thereof.
15. The method according to any one of claims 11 to 14, wherein a refinement vector in the set of refinement vectors is defined by a distance value from a distance table and by a direction value from a direction table.
16. The method according to claim 15, wherein the adaptively modifying the set of refinement vectors comprises: modifying the distance table, the modifying includes reducing a number of elements of the distance table, reordering the elements of the distance table, changing one or more values of the elements of the distance table, or a combination thereof.
17. The method according to claim 15, wherein the adaptively modifying the set of refinement vectors comprises:modifying the direction table, the modifying includes reducing a number of elements of the direction table, reordering the elements of the direction table, changing one or more values of the elements of the direction table, or a combination thereof.
18. The method according to any one of claims 11 to 17, further comprising: adaptively modifying a list of merge candidates based on one or more of respective probability values; decoding, from the bitstream, an index indicating the merge video block; and extracting, based on the index, a merge video block from the modified list of merge candidate, wherein the base motion vector is a motion vector of the merge video block.
19. The method according to claim 18, further comprising: decoding, from the bitstream, a flag indicating whether the list of merge candidates should be modified; and performing the adaptively modifying of the list of merge candidates if the flag indicates that the list of merge candidates should be modified.
20. An apparatus for encoding video data, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to: code, into a bitstream, a video block of the video data, the coding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, selecting a refinement vector from the modified set of refinement vectors, the selected refinement vector is applied to refine a motion vector of a merge video block, and coding, into the bitstream, an index indicating the selected refinement vector.
21. The apparatus according to claim 20, wherein the instructions further cause the apparatus to: code, into the bitstream, a flag indicating that the set of refinement vectors is modified.
22. The apparatus according to claim 20 or 21, wherein the adaptively modifying the set of refinement vectors further comprises: deriving a motion activity indicator, characterizing content motion in a neighborhood of the video block; and modifying the set of refinement vectors based on the motion activity indicator.
23. The apparatus according to claim 22, wherein the set of refinement vectors is modified if the motion activity indicator is above a predetermined threshold.
24. The apparatus according to any one of claims 20 to 23, wherein the deriving of the motion activity indicator comprises: determining the motion activity indicator based on dynamic coding data, including one or more of the respective probability values, a quantization parameter associated with the video block, a size associated with the video block, or a combination thereof.
25. An apparatus for decoding video data, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to: decode, from a bitstream, a video block of the video data, the decoding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, decoding, from the bitstream, an index indicating a refinement vector from the modified set of refinement vectors, and extracting, based on the index, the refinement vector from the modified set of refinement vectors, the extracted refinement vector is applied to refine a base motion vector.
26. The apparatus according to claim 25, wherein a refinement vector in the set of refinement vectors is defined by a distance value from a distance table and by a direction value from a direction table.
27. The apparatus according to claim 26, wherein the adaptively modifying the set of refinement vectors comprises: modifying the distance table, the modifying includes reducing a number of elements of the distance table, reordering the elements of the distance table, changing one or more values of the elements of the distance table, or a combination thereof.
28. The apparatus according to claim 26, wherein the adaptively modifying the set of refinement vectors comprises: modifying the direction table, the modifying includes reducing a number of elements of the direction table, reordering the elements of the direction table, changing one or more values of the elements of the direction table, or a combination thereof.
29. The apparatus according to any one of claims 25 to 28, wherein the instructions further cause the apparatus to: adaptively modify a list of merge candidates based on one or more of respective probability values; decode, from the bitstream, an index indicating a merge video block; and extract, based on the index, the merge video block from the modified list of merge candidate, wherein the base motion vector is a motion vector of the merge video block.
30. The apparatus according to claim 29, wherein the instructions further cause the apparatus to: decode, from the bitstream, a flag indicating whether the list of merge candidates should be modified; and perform the adaptively modifying of the list of merge candidates if the flag indicates that the list of merge candidates should be modified.
31. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for encoding video data, the method comprising: coding, into a bitstream, a video block of the video data, the coding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, selecting a refinement vector from the modified set of refinement vectors, the selected refinement vector is applied to refine a base motion vector, and coding, into the bitstream, an index indicating the selected refinement vector.
32. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for decoding video data, the method comprising: decoding, from a bitstream, a video block of the video data, the decoding of the video block comprises: adaptively modifying a set of refinement vectors based on respective probability values, decoding, from the bitstream, an index indicating a refinement vector from the modified set of refinement vectors, and extracting, based on the index, the refinement vector from the modified set of refinement vectors, the extracted refinement vector is applied to refine a base motion vector.