Method, apparatus and computer program for coding using motion vectors

By calculating and decoding motion vector differentials using context models, the method enhances video coding efficiency and accuracy, addressing inefficiencies in existing technologies.

JP2025534737APending Publication Date: 2025-10-17TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025521445
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-18
Filing Date
2023-10-19
Publication Date
2025-10-17

Smart Images

  • Figure 2025534737000001_ABST
    Figure 2025534737000001_ABST
Patent Text Reader

Abstract

The processing circuit receives coding information of a motion vector differential (MVD). The processing circuit calculates cost values ​​associated with value combinations for a plurality of bits in the coding bits of the MVD, at least one of the plurality of bits being a bit in a codeword for indicating the magnitude of the MVD. The processing circuit determines a predicted value combination for the plurality of bits from the value combinations, the predicted value combination being associated with the lowest cost value among the cost values. The processing circuit decodes the coding information of the MVD to obtain one or more indicators for the predicted value combination, the one or more indicators indicating whether the plurality of bits are correctly predicted by the predicted value combination. The processing circuit determines the MVD based on the predicted value combination and the one or more indicators.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Incorporated by reference] This application claims the benefit of priority to U.S. Patent Application No. 18 / 381,496, entitled "METHOD AND APPARATUS FOR MOTION VECTOR CODING," filed October 18, 2023, which claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 417,651, entitled "Method and Apparatus for Motion Vector Coding," filed October 19, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.

[0002] [Technical field] This disclosure describes embodiments generally related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. The work of the presently named inventors, to the extent that that work is described in this background discussion, along with aspects of the description that would not normally be considered prior art at the time of filing, is not admitted explicitly or implicitly as prior art to the present disclosure.

[0004] Image / video compression helps transmit image / video data across different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit receives coding information for a current block in a current picture, where the coding information includes coding information for a motion vector differential (MVD). The processing circuit calculates cost values ​​respectively associated with multiple value combinations for multiple bits in coding bits of the MVD, where the multiple bits include a partial codeword for the MVD, and at least one of the multiple bits is a bit in the codeword for indicating a magnitude for the MVD. The processing circuit determines a predicted value combination for the multiple bits from the multiple value combinations, where the predicted value combination is associated with the lowest cost value among the cost values. The processing circuit decodes the MVD coding information to obtain one or more indicators for the predicted value combination, where the one or more indicators indicate whether the multiple bits are correctly predicted by the predicted value combination. The processing circuit determines an MVD based on the combination of the predictor value and one or more indicators. Further, the processing circuit determines a motion vector for the current block based on the motion vector predictor (MVP) and the MVD, and reconstructs the current block based on a reference block in the reference picture, the reference block being indicated by the motion vector.

[0006] In some examples, the plurality of bits comprises the first N bins of a suffix of a codeword to indicate the magnitude of at least one of a horizontal component and / or a vertical component of the MVD, where N is a positive integer.

[0007] In some examples, the one or more indicators include binary values ​​associated with the first N bins of the suffix, respectively, and the binary values ​​associated with a bin of the suffix indicate whether the bin of the suffix is ​​correctly predicted in the combination of predictors.

[0008] In some examples, the one or more indicators are context coded within the MVD coding information, and the processing circuitry decodes the one or more indicators from the MVD coding information according to one or more context models.

[0009] In some examples, the one or more indicators are context coded within the MVD coding information based on the respective context models, and the processing circuitry decodes the one or more indicators from the MVD coding information in accordance with the respective context models.

[0010] In some examples, the remaining bins of the codeword are decoded from the MVD coding information according to equal probability bins.

[0011] In some examples, the plurality of bits includes at least the MVD sign and the first N bins of a suffix of the codeword to indicate the magnitude of at least one of the horizontal and / or vertical components of the MVD.

[0012] In some examples, the processing circuitry calculates template matching cost values ​​associated with each of the plurality of value combinations.

[0013] In some examples, the processing circuitry calculates a smoothness cost value associated with each of the plurality of value combinations.

[0014] In some examples, the processing circuit determines N according to syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header.

[0015] In some examples, the processing circuit decodes a prefix value of the codeword and determines that the decoded prefix value belongs to a subset of prefix values. In response to the decoded prefix value belonging to the subset of prefix values, the processing circuit calculates cost values ​​respectively associated with a plurality of value combinations for a plurality of bits comprising the first N bins of the suffix of the codeword. In one example, the subset of prefix values ​​has more than M bins for each prefix value, where M is a positive integer. In one example, the processing circuit determines M according to syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header.

[0016] In some examples, the processing circuit decodes the MVD coding information to obtain indicator bits associated with the combination of predictors, the indicator bits indicating whether the multiple bits are correctly predicted by the combination of predictors.

[0017] In some examples, the processing circuit decodes a flag bin from the coding information of the current block, the flag bin indicating whether a predictor is applied to at least one of the horizontal and / or vertical components of the MVD.

[0018] In one example, the processing circuit determines a context model for the flag bin and decodes the flag bin according to the context model.

[0019] In some examples, the processing circuit sorts the value combinations according to a cost value respectively associated with the value combinations, decodes an index from the coding information of the current block, and selects a combination from the sorted value combinations according to the index.

[0020] In some examples, the processing circuitry determines at least a context model for coding one or more bins in a prefix of the codeword, and decodes the one or more bins in the prefix based on at least the context model.

[0021] In some examples, the processing circuitry determines a context model for each of the first K bins in the prefix of the codeword and decodes each of the first K bins according to the context model.

[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding / encoding. [Brief explanation of the drawings]

[0023] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.

[0024] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a video processing system (100).

[0025] [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder.

[0026] [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder.

[0027] [Figure 4] FIG. 10 illustrates pseudocode for motion vector differential (MVD) signaling in some examples.

[0028] [Figure 5] FIG. 10 shows a table illustrating remainder values ​​and corresponding codewords in some examples.

[0029] [Figure 6] FIG. 10 is a diagram illustrating an example of template matching.

[0030] [Figure 7] FIG. 10 is a diagram for template matching based search of MVD code and magnitude combinations.

[0031] [Figure 8] FIG. 10 shows a flowchart outlining another process according to an embodiment of the present disclosure.

[0032] [Figure 9] FIG. 10 shows a flowchart outlining another process according to an embodiment of the present disclosure.

[0033] [Figure 10] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0034] 1 illustrates a block diagram of a video processing system 100 in accordance with some examples. The video processing system 100 is an example of an application of the disclosed subject matter, a video encoder and decoder in a streaming environment. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, etc.

[0035] The video processing system (100) includes a capture subsystem (113), which may include a video source (101), such as a digital camera, that creates a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102) is depicted as a thick line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream) and may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream), depicted as a thin line to emphasize its low data volume compared to the stream of video pictures (102), may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of Figure 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include a video decoder (110), for example, within an electronic device (130). The video decoder (110) decodes an input copy (107) of the encoded video data and creates an output stream (111) of video pictures that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0036] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and similarly, electronic device 130 may also include a video encoder (not shown).

[0037] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0038] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210), e.g., contained in a bitstream. In one embodiment, the coded video sequences are received one at a time, where the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (231) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be external to the video decoder 210 (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder 210, for example, to combat network jitter, and there may be another buffer memory 215 internal to the video decoder 210, for example, to handle playback timing. When the receiver 231 is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory 215 may not be needed or may be small.For use in best-effort packet networks such as the Internet, a buffer memory (215) may be required, the size of which may be relatively large, advantageously adaptively sized, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (210).

[0039] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information to control a rendering device, such as a render device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) as shown in FIG. 2 but may be coupled to the electronic device (230). The rendering device control information may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) may also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0040] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0041] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter- and intra-picture, inter- and intra-block) and other factors. Which units are involved and how can be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the units below is not shown for clarity.

[0042] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into functional units is appropriate below.

[0043] The first unit may be a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220), as well as control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) may output blocks containing sample values ​​that can be input to an aggregator (255).

[0044] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258) may be, for example, a partially reconstructed and / or fully reconstructed current picture. The aggregator (255) optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0045] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to a block that is inter-coded and potentially motion-compensated. In such cases, the motion-compensated prediction unit (253) can access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples can be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches the prediction samples can be controlled by a motion vector, which is available to the motion-compensated prediction unit (253) in the form of a symbol (221) that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0046] The output samples of the aggregator (255) may be subjected to various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.

[0047] The output of the loop filter unit (256) can be a sample stream that can be output to a render device such as a display (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0048] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and the fresh current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.

[0049] The video decoder (210) may perform decoding operations according to a given video compression technology or standard, such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and a profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from all tools available in the video compression technology or standard as the only tools available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level may constrain the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference pixel size, etc. The limits set by the level may be further constrained, in some cases, through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled with the coded video sequence.

[0050] In one embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0052] The video encoder (303) may receive video samples from a video source (301) (which, in the example of FIG. 3, is not part of the electronic device (320)) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0053] The video source (301) may provide a source video sequence to be coded by the encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may comprise one or more samples depending on the sampling structure, color space, etc., in use.

[0054] According to one embodiment, the video encoder (303) can code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Enforcing the appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) can also control and be functionally coupled to other functional units, as described below. This coupling is not shown for clarity. Parameters set by the controller (350) can include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured with other appropriate functions that may be associated with an optimized video encoder (303) for a particular system design.

[0055] In some embodiments, the video encoder is configured to operate within a coding loop. As an oversimplified explanation, in one example, the coding loop can include a source coder (330) (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture, for example) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that which a (remote) decoder would also create. The reconstructed sample stream (sample data) can be input to a reference picture memory (334). Because decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents in the reference picture memory (334) are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained due to, for example, channel errors) is used in several related fields as well.

[0056] The operation of the "local" decoder (333) may be the same as a "remote" decoder, such as the video decoder (210), already described above in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0057] In one embodiment, decoder techniques other than analysis / entropy decoding present in a decoder are present in the same or substantially the same functional form in the corresponding encoder. Therefore, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, as opposed to the decoder techniques, which are exhaustively described. Only in certain areas will more detailed descriptions be provided below.

[0058] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with respect to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0059] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures with common content as reconstructed reference pictures to be obtained by the far-end video decoder (without transmission errors).

[0060] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., which may serve as appropriate prediction references for the new picture. The predictor (335) may operate on a sample block-by-pixel block basis to find appropriate prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0061] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0062] The outputs of all of the aforementioned functional units may be subject to entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0063] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0064] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign each coded picture a particular coded picture type, which may affect the coding that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:

[0065] An intra picture (I picture) can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0066] Predictive pictures (P pictures) can be coded and decoded using intra- or inter-prediction, using motion vectors and reference indices to predict the sample values ​​of each block.

[0067] Bidirectionally predictive pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, using two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0068] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded relative to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded relative to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction relative to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction relative to one or two previously coded reference pictures.

[0069] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.

[0070] In one embodiment, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include other types of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0071] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded is called the current picture and is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0072] In some embodiments, bi-prediction techniques can be used for inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used that are located before the current picture in decoding order (but past and future, respectively, in display order) with respect to the current picture in the video. A block in the current picture can be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.

[0073] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0074] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0075] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors executing software instructions.

[0076] Aspects of the present disclosure provide techniques for coding motion vector differentials (MVDs), which are used in various inter-prediction tools such as merged motion vector differentials (MMVDs), affine MVDs, and the like.

[0077] For example, in VVC, various inter-prediction modes can be used. For inter-predicted CUs, motion parameters can include MVs, one or more reference picture indices, reference picture list usage indices, and additional information for specific coding features used for inter-prediction sample generation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU and cannot have significant residual coefficients, coded motion vector deltas or MV differences (e.g., MVDs), or reference picture indices. Merge mode can be specified when motion parameters for the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally additional information such as those introduced in VVC. Merge mode can be applied not only to skip mode but also to inter-predicted CUs. In one example, an alternative to merge mode is explicit transmission of motion parameters, in which the MVs, corresponding reference picture indices and reference picture list usage flags for each reference picture list, and other information are explicitly signaled for each CU.

[0078] In an embodiment such as VVC, the VVC Test Model (VTM) reference software includes one or more inter-prediction coding tools, including: enhanced merge prediction, merge motion vector differential (MMVD) mode, adaptive motion vector prediction with symmetric MVD signaling (AMVP) mode, affine motion compensation prediction, sub-block-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU level weights (BCW), bidirectional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Some inter-prediction and related methods are described in detail below.

[0079] In the following description, the MMVD mode is used to illustrate the use of MVD coding. In some examples, MMVD reuses merge candidates. Candidates can be selected from among the merge candidates and further extended with a motion vector representation. In some examples, the motion vector representation includes a starting point, a motion magnitude, and a motion direction.

[0080] In some instances (e.g., VVC), the MMVD technique may use a merge candidate list to select candidate starting points, however, in one instance, only candidates that are of the default merge type (MRG_TYPE_DEFAULT_N) are considered for MMVD expansion.

[0081] In some examples, a base candidate index is used to define a starting point. The base candidate index indicates the best candidate among the candidates in the list shown in Table 1. For example, the list is a merge candidate list with a motion vector predictor (MVP). The base candidate index can indicate the best candidate in the merge candidate list. [Table 1]

[0082] Note that in one example, when the number of base candidates is equal to 1, the base candidate IDX is not signaled.

[0083] In MMVD mode, after a merge candidate (also called an MV basis or MV starting point) is selected, the merge candidate can be refined by additional information such as signaled MVD information. The additional information can indicate the MVD (or motion offset) relative to the MV basis. For example, the motion magnitude indicates the magnitude of the MVD, and the motion direction indicates the direction of the MVD.

[0084] In some examples (e.g., in VVC), for MVD coding, the MVD magnitude value (absolute value of MVD) is signaled first, followed by the code value. For example, for each horizontal and vertical component of MVD, a first flag (e.g., indicated by abs_mvd_greater0_flag) is coded first to indicate whether the absolute value of the component is greater than 0. When the absolute value of the component is greater than 0, a second flag (e.g., indicated by abs_mvd_greater1_flag) is coded to indicate whether the absolute value of the component is greater than 1. When the absolute value of the component is greater than 1, a remainder value (absolute value -2) is coded, as indicated by abs_mvd_minus2.

[0085] In some examples, the flags abs_mvd_greater0_flag and abs_mvd_greater1_flag are both context coded, the remainder value abs_mvd_minus2 is binarized using a constrained first-order Exp-Golomb binarization process, and each resulting bin is coded using bypass coding.

[0086] Figure 4 shows pseudocode (400) for MVD signaling in some examples. In the example of Figure 4, for the horizontal and vertical components of MVD, a first flag (abs_mvd_greater0_flag[0] for the horizontal component and abs_mvd_greater0_flag[1] for the vertical component) is coded to indicate whether the absolute value of the horizontal and vertical components is greater than 0. For example, this is shown by (410) in Figure 4.

[0087] For the horizontal component, when the absolute value of the horizontal component is greater than 0, a second flag (abs_mvd_greater1_flag[0]) is coded to indicate whether the absolute value of the horizontal component is greater than 1, as shown by (421) in Figure 4. When the absolute value of the horizontal component is greater than 1, a remainder value (absolute value - 2 indicated by abs_mvd_minus2[0]) is coded, as shown by (431) in Figure 4.

[0088] For the vertical component, when the absolute value of the vertical component is greater than 0, a second flag (abs_mvd_greater1_flag[1]) is coded to indicate whether the absolute value of the vertical component is greater than 1, as shown by (422) in Figure 4. When the absolute value of the vertical component is greater than 1, a remainder value (the absolute value indicated by abs_mvd_minus2[1] - 2) is coded, as shown by (432) in Figure 4.

[0089] Note that each remainder value is coded by a codeword of the first-order Exp-Golomb binarization process, including a prefix and a suffix.

[0090] 5 shows a table 500 illustrating some example remainder values ​​and corresponding codewords. The remainder values ​​are coded by codewords of a first-order Exp-Golomb binarization process, where each codeword includes a prefix and a suffix.

[0091] In one example, the remainder value is converted to a codeword. The remainder is first converted to binary. Then, the number of binary bits is determined. The number of binary bits minus one is the number of leading zeros in the prefix of the codeword.

[0092] In some examples (such as VVC), sign bits such as the horizontal sign bit indicated by mvd_sign_flag[0] and the vertical sign bit indicated by mvd_sign_flag[1] in Figure 4 are bypass coded with bins having a probability of 0.5.

[0093] Entropy coding techniques, such as those for context-based adaptive binary arithmetic coding (CABAC), can also reduce signaling costs. For example, a context model in CABAC for coding a current block can be determined based on information from the current block's temporally neighboring blocks (also called temporally co-located blocks) and / or the current block's spatially neighboring blocks. CABAC can be used to code various syntax elements, such as an inter-affine flag (e.g., inter_affine_flag), a sub-block merging flag (e.g., merge_subblock_flag), a local illumination compensation (LIC) flag (e.g., lic_flag), etc. The inter-affine flag associated with a coding block is used to indicate whether the coding block is coded using affine motion compensation prediction. The sub-block merging flag associated with a coding block is used to indicate whether the coding block is coded using a sub-block motion compensation mode. The LIC flag associated with a coding block is used to indicate whether the coding block is coded using local illumination compensation.

[0094] CABAC is a coding technique used for entropy coding. Generally, the encoding process of CABAC includes a binarization step, a context modeling step, and an arithmetic coding step.

[0095] In the binarization step for a CABAC-based encoding process, non-binary valued syntax elements can be mapped to bin binary sequences, also called bin strings. When syntax elements are provided with values ​​in binary form (e.g., binary sequences), the binarization step can be bypassed.

[0096] In the context modeling step for a CABAC-based encoding process, a probability model is determined depending on previously encoded syntax elements. In some examples, the probability model (also referred to as a context or context model in CABAC) can be represented by a probability state (also referred to as a context state in some examples) and a Most Probable Symbol (MPS) value. The probability state is associated with a probability value and can implicitly represent that the probability that a particular symbol (e.g., a bin) is the Least Probable Symbol (LPS) is equal to the probability value. A symbol can be an LPS or an MPS. For binary symbols, the MPS and LPS can be 0 or 1. For example, if the LPS is 1, the MPS is 0, and if the LPS is 0, the MPS is 1. The probability value can be estimated for the corresponding context and used to entropy code the symbol using an arithmetic coder.

[0097] The arithmetic coding step of the CABAC-based encoding process is based on the principle of recursive interval subdivision according to a probability model. In some examples, the arithmetic coding step is processed by a state machine having a range parameter and a low parameter. The state machine can change the values ​​of the range parameter and the low parameter based on the context (probability model) and the sequence of bins to be coded. The value of the range parameter indicates the size of the current range into which the coded value (of the bin) falls, and the value of the low parameter indicates the lower limit of the current range. In one example, according to the probability state (e.g., related to the probability value), the current range (CurrRange) is divided into a first subrange (MpsRange) (also referred to as the MPS range of the current state) and a second subrange (LpsRange) (also referred to as the LPS range of the current state). In one example, the second subrange can be calculated by multiplication using Equation (1): LpsRange = CurrRange × ρ Equation (1) where ρ is the probability value that the current bin is an LPS. The probability that the current bin is an MPS can be calculated by (1-ρ). The first subrange can be calculated by formula (2): MpsRange=CurrRange-LpsRange Formula (2)

[0098] In one example, when the current bin is an MPS, the value of the low parameter is maintained and the value of the range parameter is updated to MpsRange, and when the current bin is an LPS, the value of the low parameter is updated to (low+MpsRange) and the value of the range parameter is updated to LpsRange. The CABAC encoding process can then continue to the next bin in the bin sequence.

[0099] In some examples (e.g., HEVC), the value of the range parameter is represented by 9 bits, and the value of the low parameter is represented by 10 bits. Furthermore, to maintain sufficient precision for the range and low values, a renormalization process can be performed. For example, whenever the value of the range parameter is less than 256, renormalization can be performed. Thus, the range parameter becomes equal to or greater than 256 after renormalization.

[0100] In some examples (e.g., HEVC), 64 possible probability values ​​for the LPS may be used, and each MPS may be 0 or 1. In one example, the probability model may be stored as 7-bit entries corresponding to 64 probability values ​​(64 probability states) and two possible values ​​for the MPS (0 or 1). In each of the 7-bit entries, 6 bits may be allocated to represent the probability states, and 1 bit may be allocated to the MPS.

[0101] According to one aspect of the present disclosure, a technique called template matching (TM) can be used in video / image coding to improve coding efficiency. For example, to further improve compression efficiency in the VVC standard, TM can be used to refine the MV. In one example, TM is used on the decoder side. In TM mode, the MV can be refined by constructing a template (e.g., a current template) for a block (e.g., a current block) in a current picture and determining the closest match between the template for the block in the current picture and multiple possible templates (e.g., multiple possible reference templates) in a reference picture. In one embodiment, the template for a block in the current picture can include reconstructed samples adjacent to the left of the block and reconstructed samples adjacent to the top of the block.

[0102] 6 shows an example of template matching (600). Using a TM, motion information for a current CU (e.g., a current block) (601) can be derived (e.g., deriving final motion information from initial motion information such as initial MV 602) by determining the closest match between a template (e.g., current template) (621) of a current CU (601) in a current picture (610) and a template (e.g., reference template) of multiple possible templates (e.g., one of the multiple possible templates is template (625)) in a reference picture (611). The template (621) of the current CU (601) can have any suitable shape and any suitable size.

[0103] In one embodiment, the template 621 of the current CU 601 includes a top template 622 and a left template 623. Each of the top template 622 and the left template 623 can have any suitable shape and any suitable size.

[0104] The top template (622) can include samples in one or more upper neighboring blocks of the current CU (601). In one example, the top template (622) includes four rows of samples in one or more upper neighboring blocks of the current CU (601). The left template (623) can include samples in one or more left neighboring blocks of the current CU (601). In one example, the left template (623) includes four columns of samples in one or more left neighboring blocks of the current CU (601).

[0105] Each one of the multiple possible templates in the reference picture (611) (e.g., template (625)) corresponds to the template (621) in the current picture (610). In one embodiment, the initial MV (602) points from the current CU (601) to the reference block (603) in the reference picture (611). Each one of the multiple possible templates in the reference picture (611) (e.g., template (625)) and the template (621) in the current picture (610) may have the same shape and size. For example, the template (625) for the reference block (603) includes a top template (626) in the reference picture (611) and a left template (627) in the reference picture (611). The top template (626) may include samples in one or more upper neighboring blocks of the reference block (603). The left template (627) may include samples in one or more left neighboring blocks of the reference block (603).

[0106] A TM cost can be determined based on a pair of templates, such as a template (e.g., a current template) (621) and a template (e.g., a reference template) (625). The TM cost can indicate a match between the template (621) and the template (625). An optimized MV (or final MV) can be determined based on a search around the initial MV (602) of the current CU (601) within a search range (615). The search range (615) can have any suitable shape and any suitable number of reference samples. In one example, the search range (615) in the reference picture (611) includes a [-L,L]-pel range, where L is a positive integer such as 8 (e.g., 8 samples). For example, a difference (e.g., [0,1]) is determined based on the search range (615), and an intermediate MV is determined by adding the initial MV (602) and the difference (e.g., [0,1]). Based on the intermediate MV, an intermediate reference block and a corresponding template in the reference picture (611) can be determined. A TM cost can be determined based on the template (621) and an intermediate template in the reference picture (611). The TM cost can correspond to a difference (e.g., [0,0] corresponding to the initial MV (602) [0,1]) determined based on the search range (615). In one example, the difference corresponding to the smallest TM cost is selected, and the optimized MV is the sum of the difference corresponding to the smallest TM cost and the initial MV (602). As described above, the TM can derive final motion information (e.g., the optimized MV) from initial motion information (e.g., the initial MV 602).

[0107] In the example of FIG. 6, a better MV can be searched for around the initial motion vector of the current CU within a search range such as [-8pel, +8pel].

[0108] In some cases, the MVD code can be predicted. For example, TM techniques can be used to explore possible MVD code combinations. The possible MVD code combinations can be sorted according to template matching cost, and an index corresponding to the true MVD code is derived and used for context coding.

[0109] In some examples, at the decoder side, the MVD code is suitably derived, for example, in the following five steps: In the first step, syntax elements for the magnitude of the MVD components are parsed from the coded video bitstream; In the second step, context-coded MVD code prediction indices are parsed from the coded video bitstream; In the third step, MV candidates are constructed, for example, by creating combinations between possible MVD codes and absolute MVD values ​​and adding the combinations to an MV predictor list; In the fourth step, an MVD code prediction cost for each combination in the MV predictor list is derived based on the template matching cost, and the combinations in the MV predictor list are sorted according to the MVD code prediction costs associated with these combinations; In the fifth step, the MVD code prediction index is used to select a combination in the MV prediction list, which includes the true MVD code to be combined with the absolute MVD value.

[0110] The MVD code prediction technique can be applied to inter-AMVP, affine AMVP, MMVD and affine MMVD modes.

[0111] It should be noted that template matching-based MVD code prediction can reduce the signaling cost of MVD coding. Some aspects of the present disclosure provide additional prediction techniques applied to MVD magnitude coding to further reduce the signaling cost. In some examples, the encoder / decoder can calculate cost values ​​associated with each possible combination of values ​​of multiple bits in the MVD coding bits. At least one of the multiple bits is a bit in the codeword for indicating the MVD magnitude. The encoder / decoder can determine a combination of predicted values ​​for the multiple bits from the possible combinations, and the predicted value combination is associated with the lowest cost value among the cost values. The encoder can encode one or more indicators (also referred to as predicted values ​​in some examples) for the predicted value combination in the MVD coding information. The decoder can decode the one or more indicators from the MVD coding information. The one or more indicators indicate whether the multiple bits are correctly predicted by the predicted value combination.

[0112] Some aspects of the present disclosure provide techniques for predicting the values ​​of the first N bins of a suffix of a codeword in motion vector differential (MVD) coding. In some examples, first-order Exp-Golomb binarization is used to generate codewords for residual values ​​in MVD coding of horizontal and vertical component magnitudes. Each codeword includes a prefix and a suffix. The present disclosure provides techniques for predicting the values ​​of the first N bins of the suffix within the codeword.

[0113] In some embodiments, to signal the first N bins of the suffix of the MVD, instead of signaling actual values, a predicted value is coded into the coded video bitstream. For each bin of the first N bins, the predicted value is a binary value that indicates whether the actual value of the bin is correctly predicted or not. In some examples, the predicted value is context coded. In some examples, different contexts are used for entropy coding for different bins of the first N bins of the suffix.

[0114] In some embodiments, the remaining MVD bits, including the remaining bins in the codeword prefix and suffix after the first N bins, are signaled in the coded video bitstream. In some examples, the remaining MVD bits are treated as a regular MVD magnitude and signaled using the same method as signaling the MVD magnitude bits. In some examples, the remaining MVD bits are signaled using equal probability (EP) bins, where each bin has an equal probability of being "1" or "0." Note that bins coded using EP are also referred to as bypass coded.

[0115] In some embodiments, the predicted values ​​of the first N bins of the suffix (along with the predicted values ​​of the MVD code) are derived by a template matching process. For example, different combinations of the predicted values ​​and the MVD code are used to form possible MVDs. The possible MVDs can point to possible reference blocks. Then, the template matching cost between each possible reference block and the current block is calculated, and the one with the smallest template matching cost is selected as the predicted value of the first N bins of the suffix.

[0116] FIG. 7 shows a diagram (700) for a template matching-based search of an MVD code and magnitude combination. In the example of FIG. 7, a current block (701) is in the current picture. Neighboring samples of the current block (701) can form a template associated with the current block (701). For example, the templates associated with the current block (701) include a top template (702) and a left template (703). Further, in FIG. 7, a motion vector predictor (MVP) can indicate a base motion vector.

[0117] In one example, different values ​​of the first bin of the suffix for the horizontal component codeword, different values ​​of the first bin of the suffix for the vertical component codeword, and different values ​​of the MVD code can form 16 combinations of MVD candidates, each combination corresponding to a motion vector differential (MVD). Using MVP, the 16 combinations of MVD candidates can indicate 16 possible reference blocks in the reference picture. Neighboring samples of each possible reference block can form possible reference templates associated with the possible reference block. Furthermore, for each of the 16 possible reference templates, a template matching cost relative to the template of the current block is calculated. Then, in one example, the combination with the lowest template matching cost can be selected to generate a predicted value of the first bin of the suffix for the horizontal component codeword and a predicted value of the first bin of the suffix for the vertical component codeword.

[0118] As an example, eight of the sixteen combinations are shown in Figure 7, with the other eight omitted for ease and clarity. For example, the first combination corresponds to a first MVD (711), where the first MVD (711) using an MVP can indicate a first possible reference block in a reference picture, and a first possible reference template associated with the first possible reference block includes a top template (712) and a left template (713). The second combination corresponds to a second MVD (721), where the second MVD (721) using an MVP can indicate a second possible reference block in a reference picture, and a second possible reference template associated with the second possible reference block includes a top template (722) and a left template (723). The third combination corresponds to a third MVD (731), where the third MVD (731) using the MVP can indicate a third possible reference block in the reference picture, and the third possible reference template associated with the third possible reference block includes a top template (732) and a left template (733). The fourth combination corresponds to a fourth MVD (741), where the fourth MVD (741) using the MVP can indicate a fourth possible reference block in the reference picture, and the fourth possible reference template associated with the fourth possible reference block includes a top template (742) and a left template (743). The fifth combination corresponds to a fifth MVD (751), where the fifth MVD (751) using the MVP can indicate a fifth possible reference block in the reference picture, and the fifth possible reference template associated with the fifth possible reference block includes a top template (752) and a left template (753). The sixth combination corresponds to a sixth MVD (761), and the sixth MVD (761) using the MVP can indicate a sixth possible reference block in the reference picture, and the sixth possible reference templates associated with the sixth possible reference block include a top template (762) and a left template (763).The seventh combination corresponds to a seventh MVD (771), where the seventh MVD (771) using the MVP can indicate a seventh possible reference block in the reference picture, and the seventh possible reference templates associated with the seventh possible reference block include the top template (772) and the left template (773). The eighth combination corresponds to an eighth MVD (781), where the eighth MVD (781) using the MVP can indicate an eighth possible reference block in the reference picture, and the eighth possible reference templates associated with the eighth possible reference block include the top template (782) and the left template (783).

[0119] In some embodiments, the predicted values ​​of the first N bins of the suffix (along with the predicted value of the MVD code) are derived by a smoothness cost value that measures the smoothness of the block boundary for the candidate MVD value. Note that the smoothness cost value can be measured using any suitable smoothness cost metric. In one example, a smoothness cost metric in the spatial domain is used. For example, the smoothness cost metric (also called a discontinuity cost metric in one example) can be calculated based on the sum of absolute sample differences along the block boundaries of the reconstructed block, for example, along the top and left block boundaries.

[0120] In another example, a smoothness cost metric in the transform domain is used. The smoothness cost metric in the transform domain can be estimated as the energy of discontinuities on block boundaries in the transform domain. In one example, the smoothness cost metric used to predict the sign values ​​of the transform coefficients can be used to evaluate smoothness.

[0121] In some embodiments, the value of N is signaled in a high-level syntax (HLS), such as a video parameter set (VPS), a sequence parameter set (SPS), a picture header, a slice header, or the like.

[0122] Note that in some examples, prediction of the first N bits of the codeword suffix is ​​used only for a selected subset of prefix values. In some examples, prediction of the first N bits of the codeword suffix is ​​applied only to codewords associated with prefixes that have more than M bins. Example values ​​of M include, but are not limited to, 1, 2, 3, 4, ..., 16, etc.

[0123] In some examples, prediction of the first N bins of the codeword's suffix is ​​applied only for codewords associated with prefixes with more than M bins, and the value of M is signaled in an HLS such as a VPS, SPS, picture header, slice header, etc.

[0124] Some aspects of the present disclosure provide techniques for predicting a combination of code and MVD value, and whether the prediction is correct or not is signaled by one bit.

[0125] In some embodiments, the code and MVD combination with the best template matching cost based on the remaining signaled MVD bits is used as the predictor.

[0126] In one embodiment, one flag bin in the bitstream is signaled to indicate whether the predictor is used for the MVD horizontal or vertical component. If the flag is not used, normal MVD coding is applied to MVD. In one example, this flag bin is coded using the CABAC context model. In another example, this flag bin is signaled as an EP bin.

[0127] In one embodiment, one flag bin in the bitstream is signaled. When the flag bin is true, it indicates that both horizontal and vertical component signs and significant bits predictors are used. In one example, when the flag bin is false, regular MVD coding is applied to both horizontal and vertical MVD components. In another example, when the flag bin is false, each horizontal or vertical component may have a flag bit signaled to indicate whether a predictor is used for this component.

[0128] In some embodiments, the code and significant MVD combinations are listed in ascending order of template matching cost, and an index is signaled in the bitstream to indicate which prediction is used by the decoder's reconstruction process.

[0129] According to some aspects of the present disclosure, instead of bypass coding, context coding can be used for the first K bins of the MVD codeword prefix. In one embodiment, different contexts are applied to different bins of the first K bins of the MVD codeword prefix.

[0130] 8 is a flowchart outlining a process (800) according to an embodiment of the present disclosure. The process (800) may be used in a video decoder. In various embodiments, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some embodiments, the process (800) is implemented with software instructions, and thus, the processing circuit performs the process (800) when it executes the software instructions. Processing begins at (S801) and proceeds to (S810).

[0131] At (S810), coding information for a current block in a current picture is received, the coding information including coding information for a motion vector difference (MVD) between a motion vector prediction (MVP) and a (true) motion vector used to reconstruct the current block.

[0132] At step S820, cost values ​​associated with respective combinations of values ​​of bits in the coding bits of the MVD are calculated, the bits comprising a partial codeword for the MVD, and at least one of the bits is a bit in the codeword for indicating the magnitude of the MVD.

[0133] In (S830), a combination of predicted values ​​of a plurality of bits is determined from the plurality of value combinations, and the combination of predicted values ​​is associated with the lowest cost value among the cost values.

[0134] At (S840), the MVD coding information is decoded to obtain one or more indicators for the predictor combination, where the one or more indicators indicate whether a plurality of bits are correctly predicted by the predictor combination.

[0135] At (S850), the (true) MVD is determined based on a combination of the predicted value and one or more indicators.

[0136] At (S860), a motion vector for the current block is determined based on the motion vector predictor (MVP) and the MVD.

[0137] At (S870), the current block is reconstructed based on a reference block in a reference picture, where the reference block is indicated by a motion vector.

[0138] In some examples, the plurality of bits includes the first N bins of a codeword suffix to indicate a magnitude, such as the magnitude of the horizontal component and / or the magnitude of the vertical component of the MVD, where N is a positive integer.

[0139] In some examples, the one or more indicators include binary values ​​associated with the first N bins of the suffix, each of which indicates whether a prediction in the combination of predictions for that bin is correct.

[0140] In some examples, the one or more indicators are context coded within the MVD coding information, and the one or more indicators can be decoded from the MVD coding information according to one or more context models.

[0141] In some examples, the one or more indicators are context coded within the MVD coding information based on the respective context models, and the one or more indicators can be decoded from the MVD coding information according to the respective context models.

[0142] In some examples, the remaining bins of the codeword are decoded from the MVD coding information according to equal probability (EP) bins.

[0143] In some examples, the plurality of bits includes the first N bins of a codeword suffix and at least an MVD code, where the codeword indicates the magnitude of a horizontal component of the MVD and / or the magnitude of a vertical component of the MVD.

[0144] In some examples, a template matching cost value associated with each of the plurality of value combinations is calculated as the cost value.

[0145] In some examples, a smoothness cost value associated with each of the plurality of value combinations is calculated as the cost value.

[0146] In some examples, N is determined according to values ​​of syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header.

[0147] In some examples, a prefix value of the codeword is decoded from the coding information of the MVD. Then, it is determined whether the decoded prefix value belongs to a subset of prefix values. In response to the decoded prefix value belonging to the subset of prefix values, cost values ​​respectively associated with a plurality of value combinations for a plurality of bits are calculated, and a plurality of bits including the first N bins of the suffix of the codeword are predicted according to the cost values. In one example, the subset of prefix values ​​has more than M bins for each prefix value, where M is a positive integer. In one example, M is determined according to values ​​of syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header.

[0148] In some examples, the MVD coding information is decoded to obtain indicator bits associated with the predictor combination, which indicate whether multiple bits are correctly predicted by the predictor combination.

[0149] In some examples, flag bins are decoded from the coding information of the current block, and the flag bins indicate whether a predictor (e.g., a combination of predictors) is used for at least one of the horizontal and / or vertical components of the MVD. In one example, the flag bins are context coded. Then, a context model is determined for the flag bins, and the flag bins are decoded according to the context model.

[0150] In some examples, the multiple value combinations are sorted according to cost values ​​associated with each of the multiple value combinations. An index is decoded from the coding information of the current block, and a combination is selected from the sorted value combinations according to the index. The selected value combination is used as a combination of predicted values ​​for the multiple bits.

[0151] In some examples, at least a context model for coding one or more bins in a prefix of the codeword is determined. The one or more bins in the prefix are decoded based on at least the context model.

[0152] In some examples, a context model is determined for each of the first K bins in the prefix of the codeword, and each of the first K bins is decoded according to the context model.

[0153] The process then proceeds to (S899) and ends.

[0154] The process 800 may be adapted as appropriate. Steps of the process 800 may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0155] 9 is a flowchart outlining a process (900) according to an embodiment of the present disclosure. The process (900) may be used in a video encoder. In various embodiments, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), etc. In some embodiments, the process (900) is implemented with software instructions, and thus, the processing circuit performs the process (900) when it executes the software instructions. Processing begins at (S901) and proceeds to (S910).

[0156] At (S910), it is decided to code a motion vector differential (MVD), which is between a motion vector predictor and a motion vector used to reconstruct a current block in a current picture.

[0157] At (S920), cost values ​​associated with each of a plurality of value combinations for a plurality of bits in the coding bits of the MVD are calculated, the plurality of bits comprising a partial codeword for the MVD, and at least one of the plurality of bits is a bit in the codeword for indicating the magnitude of the MVD.

[0158] In (S930), a combination of predicted values ​​for a plurality of bits is determined from the plurality of value combinations, and the predicted value combination is associated with the lowest cost value among the cost values.

[0159] At (S940), one or more indicators for the combination of predictors are determined relative to the true value of the MVD, the one or more indicators indicating whether a plurality of bits are correctly predicted by the combination of predictors.

[0160] At (S950), one or more indicators are encoded into the coding information for the current block.

[0161] In some examples, the plurality of bits includes the first N bins of a codeword suffix to indicate a magnitude, such as the magnitude of the horizontal component and / or the magnitude of the vertical component of the MVD, where N is a positive integer.

[0162] In some examples, the one or more indicators include binary values ​​associated with the first N bins of the suffix, respectively, and the binary value associated with a bin of the suffix indicates whether a prediction in the set of predictions for that bin is correct.

[0163] In some examples, one or more indicators are context coded within the coding information of the MVD. The one or more indicators may be coded according to one or more context models.

[0164] In some examples, the one or more indicators are context coded within the coding information of the MVD based on the respective context models. The one or more indicators can be encoded into the coding information of the MVD according to the respective context models.

[0165] In some examples, the remaining bins of the codeword are encoded into the MVD coding information according to equal probability (EP) bins.

[0166] In some examples, the plurality of bits includes the first N bins of a codeword suffix and at least an MVD code, where the codeword indicates the magnitude of a horizontal component of the MVD and / or the magnitude of a vertical component of the MVD.

[0167] In some examples, a template matching cost value associated with each of the plurality of value combinations is calculated as the cost value.

[0168] In some examples, a smoothness cost value associated with each of the plurality of value combinations is calculated as the cost value.

[0169] In some examples, the value of a syntax element in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header may indicate N.

[0170] In some examples, when a prefix value belongs to a subset of prefix values, coding of the suffix into the first N bits can be predicted in accordance with this disclosure. In one example, the subset of prefix values ​​has more than M bins for each prefix value, where M is a positive integer. In one example, a value of a syntax element in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header indicates M.

[0171] In some examples, indicator bits associated with the predictor combination are coded in the MVD coding information, and the indicator bits indicate whether multiple bits are correctly predicted by the predictor combination.

[0172] In some examples, the flag bin is coded in the coding information of the current block, and the flag bin indicates whether a predictor (e.g., a combination of predictors) is used for at least one of the horizontal and / or vertical components of the MVD. In one example, the flag bin is context coded. For example, a context model is determined for the flag bin, and the flag bin is coded according to the context model.

[0173] In some examples, the multiple value combinations are sorted according to cost values ​​respectively associated with the multiple value combinations, and an index is encoded in the coding information of the current block, where the index indicates that a combination from the possible combinations sorted according to the index corresponds to multiple bits in the true MVD.

[0174] In some examples, at least a context model for coding one or more bins in a prefix of the codeword is determined, and the one or more bins in the prefix are encoded based on at least the context model.

[0175] In some examples, a context model is determined for each of the first K bins in the prefix of the codeword, and each of the first K bins is encoded according to the context model.

[0176] The process then proceeds to (S999) and ends.

[0177] The process 900 may be adapted as appropriate. Steps of the process 900 may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0178] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 10 illustrates a computer system (1000) suitable for implementing embodiments of the disclosed subject matter.

[0179] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through interpretation, microcode execution, etc.

[0180] The instructions may be executed in various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0181] 10 for computer system 1000 are exemplary in nature and are not intended to suggest any limitation regarding the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 1000.

[0182] The computer system 1000 may include certain human interface input devices that may respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface input devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic vision).

[0183] The human interface input devices may include one or more of a keyboard (1001), a mouse (1002), a trackpad (1003), a touch screen (1010), a data glove (not shown), a joystick (1005), a microphone (1006), a scanner (1007), and a camera (1008) (only one of each is shown).

[0184] The computer system (1000) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1010), data gloves (not shown), or joystick (1005), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (1010), headphones (not shown), etc.), visual output devices (e.g., screens (1010), including CRT, LCD, plasma, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may provide two-dimensional visual output or greater than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0185] The computer system (1000) may also include human-accessible storage devices and their associated media, such as optical media or similar media (1021), including CD / DVD ROM / RW (1020) with CDs / DVDs, thumb drives (1022), removable hard drives or solid-state drives (1023), legacy magnetic media such as tape and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0186] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0187] The computer system 1000 may also include an interface 1054 to one or more communication networks 1055. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial TV; and vehicular and industrial networks including CANBus, etc. Certain networks generally require an external network interface adapter (e.g., a USB port on the computer system 1000) attached to a particular general-purpose data port or peripheral bus 1049; others are generally integrated into the core of the computer system 1000 by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system (1000) can communicate with other entities. Such communications can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., from a particular CANbus to a particular CANbus device), or bidirectional to other computer systems, using, for example, local or wide-area digital networks. As noted above, specific protocols and protocol stacks can be used in each of these networks and network interfaces.

[0188] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1040) of the computer system (1000).

[0189] The core (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1043), task-specific hardware accelerators (1044), graphics adapters (1050), etc. These devices may be connected through a system bus (1048), along with read-only memory (ROM) (1045), random access memory (RAM) (1046), and internal mass storage (1047), such as an internal non-user-accessible hard drive or SSD. In some computer systems, the system bus (1048) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1048) or via a peripheral bus (1049). In one example, a screen (1010) may be connected to the graphics adapter (1050). Peripheral bus architectures include PCI, USB, and the like.

[0190] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) can execute specific instructions, which, in combination, can constitute the aforementioned computer code. The computer code can be stored in ROM (1045) or RAM (1046). Temporary data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1041), GPU (1042), mass storage (1047), ROM (1045), RAM (1046), etc.

[0191] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0192] By way of example and not limitation, a computer system (1000) having an architecture, and in particular a core (1040), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as introduced above, as well as media associated with specific storage of the core (1040) that is non-transitory in nature, such as the core's internal mass storage (1047) or ROM (1045). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1040). The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core (1040) and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0193] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, when applicable, such as when the elements are not mutually exclusive.

[0194] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. 1. A processor-implemented method of video decoding, comprising: receiving coding information of a motion vector differential (MVD) associated with a current block in a current picture; calculating cost values ​​respectively associated with a plurality of value combinations of a plurality of bits from the coding bits of the MVD, the plurality of bits comprising a partial codeword for the MVD, and at least one of the plurality of bits being a bit in a codeword for indicating a magnitude of the MVD; determining a predicted value combination for the plurality of bits from the plurality of value combinations, the predicted value combination being associated with a lowest one of the cost values; decoding the coding information of the MVD to obtain one or more indicators for the combination of predictors, the one or more indicators indicating whether the plurality of bits are correctly predicted by the combination of predictors; determining the MVD based on the combination of the predictors and the one or more indicators; determining a motion vector for the current block based on a motion vector predictor (MVP) and the MVD; reconstructing the current block based on a reference block in a reference picture, the reference block being pointed to by the motion vector; A method comprising:

2. the plurality of bits comprises the first N bins of a codeword suffix for indicating the magnitude of at least one of a horizontal component and / or a vertical component of the MVD, where N is a positive integer; The method of claim 1.

3. the one or more indicators include binary values ​​associated with the first N bins of the suffix, respectively, wherein the binary value associated with a bin of the suffix indicates whether the bin of the suffix is ​​correctly predicted in a combination of predictors. The method of claim 2.

4. The one or more indicators are context coded within the coding information of the MVD, and decoding the coding information of the MVD comprises: and decoding the one or more indicators from the coding information of the MVD according to one or more context models. The method of claim 2.

5. The one or more indicators are context coded in the coding information of the MVD based on respective context models, and decoding the coding information of the MVD comprises: and decoding the one or more indicators from the coding information of the MVD according to a respective context model. The method of claim 2.

6. Determining the MVD based on the combination of the predictor values ​​and the one or more indicators includes: and decoding the remaining bins of the codeword from the coding information of the MVD according to equal probability bins. The method of claim 2.

7. the plurality of bits includes at least an MVD code and the first N bins of a suffix of the codeword to indicate the magnitude of at least one of a horizontal component and / or a vertical component of the MVD; The method of claim 1.

8. Calculating the cost values ​​associated with each of the plurality of value combinations includes: calculating a template matching cost value associated with each of the plurality of value combinations. The method of claim 1.

9. Calculating the cost values ​​associated with each of the plurality of value combinations includes: calculating a smoothness cost value associated with each of the plurality of value combinations; The method of claim 1.

10. determining N according to syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header; The method of claim 2 further comprising:

11. decoding a prefix value of the codeword; determining that the decoded prefix value belongs to a subset of prefix values; responsive to the decoded prefix value belonging to the subset of prefix values, calculating the cost values ​​respectively associated with the plurality of value combinations for the plurality of bits comprising the first N bins of the suffix of the codeword; The method of claim 2 further comprising:

12. the subset of prefix values ​​has more than M bins for each prefix value, where M is a positive integer; The method of claim 11.

13. determining M according to syntax elements in at least one of a video parameter set, a sequence parameter set, a picture header, and a slice header; The method of claim 12.

14. Decoding the coding information of the MVD to obtain one or more indicators for the combination of predictors includes: and decoding the coding information of the MVD to obtain indicator bits associated with the combination of predictors, the indicator bits indicating whether the plurality of bits are correctly predicted by the combination of predictors. The method of claim 1.

15. and decoding a flag bin from the coding information of the current block, the flag bin indicating whether a predictor is applied to at least one of a horizontal component and / or a vertical component of the MVD.

15. The method of claim 14.

16. The step of decoding the flag bin comprises: determining a context model for the flag bin; decoding the flag bins according to the context model; 16. The method of claim 15, further comprising:

17. sorting the plurality of value combinations according to the cost value associated with each of the plurality of value combinations; decoding an index from the coding information of the current block; selecting a combination from the sorted value combinations according to the index; The method of claim 1 further comprising:

18. determining at least a context model for coding one or more bins within a prefix of said codeword; decoding one or more bins in the prefix based at least on the context model; The method of claim 2 further comprising:

19. determining a context model for each of the first K bins in the prefix of the codeword; decoding each of the first K bins according to the context model; The method of claim 2 further comprising:

20. 1. An apparatus for video decoding, comprising: receiving coding information of a motion vector differential (MVD) associated with a current block in a current picture; calculating cost values ​​respectively associated with a plurality of value combinations of a plurality of bits from the coding bits of the MVD, the plurality of bits comprising a partial codeword for the MVD, and at least one of the plurality of bits being a bit in a codeword for indicating a magnitude of the MVD; determining a combination of predicted values ​​for the plurality of bits from the plurality of value combinations, the combination of predicted values ​​being associated with a lowest one of the cost values; decoding the coding information of the MVD to obtain one or more indicators for the combination of predictors, the one or more indicators indicating whether the plurality of bits are correctly predicted by the combination of predictors; determining the MVD based on the combination of the predictors and the one or more indicators; determining a motion vector for the current block based on a motion vector predictor (MVP) and the MVD; reconstructing the current block based on a reference block in a reference picture pointed to by the motion vector; The apparatus comprises a processing circuit configured to:

21. A computer program which, when executed by a processor, causes the processor to carry out the method of any one of claims 1 to 19.