METHOD AND APPARATUS FOR ENCODING / DECODING VIDEO - Patent application

JP2025510835A5Pending Publication Date: 2026-04-03INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in achieving high compression efficiency and accurate motion compensation, particularly when dealing with complex motion patterns and block discontinuities.

Method used

The use of an adaptive interpolation motion compensation filter, which determines the optimal filter length and type based on block discontinuities, location within the coding unit, and motion differences between blocks, to improve prediction accuracy and compression efficiency.

Benefits of technology

This approach enhances interprediction accuracy and compression efficiency by adaptively selecting the most suitable motion compensation interpolation filter for each block, effectively addressing the limitations of fixed filter lengths and motion models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding or decoding video is provided. At least one motion compensated interpolation filter for a block of the video is determined. In an embodiment, the length of the motion compensated interpolation filter is adapted based on a block condition. A predictive block is determined based on the at least one motion compensated interpolation filter and a reference block determined for the block, and the block of the video is encoded or decoded based on at least the predictive block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THE DISCLOSURE The present embodiments generally relate to methods and apparatus for video encoding or decoding. Some embodiments relate to methods and apparatus for video encoding or decoding using adaptive interpolation motion compensation filters. [Background technology]

[0002] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transformation to exploit spatial and temporal redundancy in video content. In general, intra- or inter-prediction is used to exploit intra- or inter-picture correlation, and then the difference between the original block and the predicted block, often called the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to the entropy encoding, quantization, transformation, and prediction. Summary of the Invention

[0003] According to a first aspect, a method for encoding video is provided, comprising determining at least one motion compensated interpolation filter for a block of the video, and determining a predictive block based on the at least one motion compensated interpolation filter and a reference block, wherein the block is encoded based on at least the predictive block.

[0004] An apparatus for encoding video is provided, the apparatus comprising one or more processors configured to encode a block of the video, whereby the one or more processors are configured to determine at least one motion compensated interpolation filter for the block, determine a predictive block based on the at least one motion compensated interpolation filter and a reference block, and encode the block based on at least the predictive block.

[0005] According to another aspect, a method for decoding video is provided, wherein at least one motion compensated interpolation filter for a block of the video is determined, and a predictive block is determined based on the at least one motion compensated interpolation filter and a reference block, and the block is decoded based on at least the predictive block.

[0006] An apparatus for decoding video is provided, the apparatus comprising one or more processors configured to decode a block of the video, whereby the one or more processors are configured to determine at least one motion compensated interpolation filter for the block, determine a predictive block based on the at least one motion compensated interpolation filter and a reference block, and decode the block based on at least the predictive block.

[0007] One or more embodiments also provide a computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to the methods described above. One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described above. One or more embodiments also provide methods and apparatus for transmitting or receiving a bitstream generated according to the methods described above. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 illustrates a block diagram of a system in which aspects of the present embodiments may be implemented. [Diagram 2] FIG. 2 shows a block diagram of an embodiment of a video encoder. [Diagram 3] FIG. 3 shows a block diagram of an embodiment of a video decoder. [Figure 4]FIG. 4 shows the control-point based affine motion models, shown for a four-parameter affine model on the left and a six-parameter affine model on the right. [Diagram 5] FIG. 5 shows an affine motion vector field for each sub-block. [Figure 6] FIG. 6 shows spatially neighboring blocks used in the sub-block temporal motion vector prediction mode in VVC (SbTMVP). [Figure 7] FIG. 7 shows an example of deriving a sub-CU (sub-coding unit) motion field by applying motion shift from spatial neighbors and scaling motion information from corresponding co-located sub-CUs in SbTMVP mode. [Figure 8] FIG. 8 shows an example of the frequency response of a 12-tap interpolation filter (proposed) and a VVC interpolation filter at half-pixel phase. [Figure 9] FIG. 9 illustrates an example of a method for encoding a block in a video according to an embodiment. [Figure 10] FIG. 10 illustrates an example of a method for decoding a block in a video according to an embodiment. [Figure 11] FIG. 11 shows an example of sub-block positions in a coding unit (CU). [Figure 12] FIG. 12 illustrates an example of a motion difference based MCIF selection method according to one embodiment. [Figure 13] FIG. 13 illustrates an example of a method for selecting an MCIF based on the coding modes of neighboring blocks, according to an embodiment. [Figure 14] FIG. 14 illustrates an example of an MCIF selection method based on a four-parameter affine model condition, according to an embodiment. [Figure 15] FIG. 15 illustrates an example of creating an asymmetric filter where the left portion of the filter is shorter than the right portion of the filter, according to one embodiment. [Figure 16] FIG. 16 illustrates an example of creating an asymmetric filter where the right portion of the filter is shorter than the left portion of the filter, according to one embodiment. [Figure 17] FIG. 17 illustrates an example of an asymmetric filter selection method, according to one embodiment. [Figure 18] FIG. 18 shows two remote devices communicating over a communications network, in accordance with an example of the present principles. [Figure 19] FIG. 19 shows the syntax of a signal, in accordance with an example of the present principles. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] In this application, various aspects are described, including tools, features, embodiments, models, methods, and the like. Many of these aspects are described with specificity and in what may often sound definitive terms to at least indicate their individual characteristics. However, this is for purposes of clarity of description and does not limit the applicability or scope of the aspects. In fact, all of the different aspects can be combined and substituted to provide further aspects. Moreover, these aspects can also be combined and substituted with aspects described in previous applications as well.

[0010] The aspects described and contemplated in this application can be implemented in many different forms. Figures 1, 2, and 3 below provide some embodiments, but other embodiments are contemplated, and discussion of Figures 1, 2, and 3 is not intended to limit the scope of implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.

[0011] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably.

[0012] Various methods are described herein, each of which includes one or more steps or acts for achieving the described method. Unless a specific order of steps or acts is required for proper operation of the method, the order and / or use of specific steps and / or acts may be modified or combined. It is noted that terms such as "first", "second", etc. may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decode" and "second decode". The use of such terms does not imply any ordering to the modified operations, unless specifically required. Thus, in this embodiment, the first decode need not be performed before the second decode, but may occur, for example, before, during, or during a period of overlap with the second decode.

[0013] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include various electronic devices, such as, but not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected appliances, and servers. Elements of system 100, alone or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, through a communication bus or dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described herein.

[0014] The system 100 includes at least one processor 110 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., volatile and / or non-volatile memory devices). The system 100 includes a storage device 140, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0015] The system 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded or decoded video, which may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. It is noted that the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.

[0016] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described herein may be stored in the storage device 140 and then loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, expressions, operations, and operational logic.

[0017] In some embodiments, memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, 13818-1 also known as H.222, and 13818-2 also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, H.265 and MPEG-H Part 2 are also known), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Team (JVET)).

[0018] Inputs to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices may include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcast station, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video, not shown in FIG. 1.

[0019] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band that may be referred to (for example) as a channel in a particular embodiment, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs these various functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a nearby baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and refiltering RF signals transmitted over a wired (e.g., cable) medium to a desired frequency band. Various embodiments rearrange the order of the above-mentioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0020] It should be noted that the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or in the processor 110, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, for example, in a separate interface IC or in the processor 110, as desired. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 operating in combination with memory and storage elements that process the data streams as desired for presentation on an output device.

[0021] The various elements of system 100 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between each other using suitable connection arrangements 115, e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards.

[0022] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.

[0023] Data is streamed to the system 100 using a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), in various embodiments. The Wi-Fi signal in these embodiments is received by a communication channel 190 and communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the streamed data is provided to the system 100 using a set-top box that delivers data through an HDMI connection of the input block 105. In still other embodiments, the streamed data is provided to the system 100 using an RF connection of the input block 105. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0024] The system 100 may provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 in various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 may be for a television, a tablet, a laptop, a mobile phone, or other device. The display 165 may also be integrated with other components (e.g., as in a smart phone) or may be separate (e.g., an external monitor for a laptop). The other peripheral devices 185, in various examples of embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (DVR in both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functionality based on the output of the system 100. For example, a disc player performs the functionality of playing the output of the system 100.

[0025] In various embodiments, control signals are communicated between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that allow inter-device control with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 may be integrated into a single unit with other components of system 100, for example, in an electronic device such as a television. In various embodiments, display interface 160 includes a display driver, for example, a Timing Controller (T Con) chip.

[0026] Display 165 and speakers 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where display 165 and speakers 175 are external components, the output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0027] The embodiments may be performed by computer software implemented by the processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 120 may be of any type suitable for the technology environment, and may be implemented using any suitable data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories, and removable memory devices. The processor 110 may be of any type suitable for the technology environment, and may include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0028] 2 shows a block diagram of an embodiment of a video encoder 200. Variations of this encoder 200 are contemplated, but for clarity, the encoder 200 is described below without describing all possible variations.

[0029] Before being encoded, a video sequence may undergo encoding pre-processing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components to obtain a signal distribution that is more resistant to compression (e.g., using histogram equalization of color components) or resizing (ex: downscaling) the picture. Metadata may be associated with the pre-processing and attached to the bitstream.

[0030] In the encoder 200, a picture is coded by the encoder elements as described below. The picture to be coded is divided (202) into units, e.g., CUs, and processed. Each unit is coded, e.g., using either intra mode or inter mode. When a unit is coded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides (205) whether to use intra mode or inter mode to code the unit, and indicates the intra / inter decision, e.g., by a prediction mode flag. The encoder may also mix (263) intra and inter prediction results, or mix results from different intra / inter prediction methods. A prediction residual is calculated (210), e.g., by subtracting the predicted block from the original image block.

[0031] The motion refinement module (272) uses already available reference pictures to refine the motion field of a block without referring to the original block. The motion field for a region can be considered as a collection of motion vectors for all pixels that comprise the region. If the motion vectors are subblock-based, the motion field can also be represented as a collection of all subblock motion vectors in the region (all pixels in a subblock have the same motion vector, and the motion vectors can be different for each subblock). If a single motion vector is used for a region, the motion field for the region can also be represented by a single motion vector (the same motion vector for all pixels in the region).

[0032] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.

[0033] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / Sample Adaptive Offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0034] Figure 3 shows a block diagram of a video decoder 300. In the decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs a decoding path that is the inverse of the encoding path described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.

[0035] In particular, the decoder's input includes a video bitstream, which may be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. The decoder may then partition the picture according to the decoded picture partition information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (355) to reconstruct an image block.

[0036] A prediction block can be obtained (370) from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375). The decoder may mix (373) intra and inter prediction results, or mix results from multiple intra / inter prediction methods. Before motion compensation, the motion field may be improved (372) by using already available reference pictures. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0037] The decoded picture may further undergo post-decoding processing (385), such as an inverse color conversion (e.g., from YCbCr 4:2:0 to RGB 4:4:4), or an inverse remapping or resizing (ex: upscaling) of the reconstructed picture that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0038] In some embodiments, Figures 2 and 3 also show encoders / decoders that improve upon the HEVC or VVC standard, or that employ techniques similar to VVC, such as encoders under development by the Joint Video Exploration Team (JVET). Various methods and other aspects described herein can be used to modify modules, such as the motion compensation module (270) of the video encoder 200 shown in Figure 2 and the motion compensation module (375) of the video decoder 300 shown in Figure 3. Furthermore, aspects of the present disclosure are not limited to VVC or HEVC, but can be applied to other standards and recommendations, such as existing or future developments, and extensions of any such standards and recommendations, including VVC and HEVC. Unless otherwise specified or technically precluded, aspects described herein can be used individually or in combination.

[0039] Inter prediction is a coding tool in video compression that uses a reference picture of a video to predict a current block of a current picture to be encoded / decoded. The encoder selects the best block in a reference frame after applying a motion model (e.g., translation or sub-block-based motion warping). The best block can be understood, for example, in the sense of rate / distortion for predicting the current block. Hereinafter, this selected block is called the reference block.

[0040] In current video codecs such as VVC, two types of motion compensation (MC) filters are available: an 8-tap filter based on DCT-IF and a smoother filter for 1 / 2 MC. In new video codec software, larger MC filters are used (e.g., 12-tap filters).

[0041] Some video codecs also propose several filters (eg, a normal filter, a smooth filter and a sharp filter), possibly combined for the horizontal and vertical directions.

[0042] In modern codecs, MC filters are becoming longer, improving the accuracy of the reconstruction. However, in some cases, for example when the filter is taking samples from different moving regions, a longer filter is not desirable.

[0043] High-precision (1 / 16 pixel) motion compensation and motion vector storage VVC increases the MV precision to 1 / 16 luma samples to improve prediction efficiency for slow-motion videos. This higher motion precision is especially useful for video content with locally varying non-translational motion, such as in the case of affine mode (warping). For higher MV precision fractional position sample generation, the 8-tap luma and 4-tap chroma interpolation filters of HEVC are extended to 16 phases for luma and 32 phases for chroma. This extended filter set is applied to MC processing of inter-coded CUs (coding units) other than affine mode CUs. For affine mode, a set of 6-tap luma interpolation filters with 16 phases is used for lower computational complexity as well as memory bandwidth savings.

[0044] In VVC, the highest precision of explicitly signaled motion vectors for non-affine CUs is a quarter luma sample. In some inter prediction modes, such as affine modes, motion vectors may be signaled at 1 / 16 luma sample precision. For all inter coded CUs with implicitly estimated MVs, the MVs are derived at 1 / 16 luma sample precision and motion compensated prediction is performed at 1 / 16 sample precision. For intra motion field storage, all motion vectors are stored at 1 / 16 luma sample precision.

[0045] For the temporal motion field storage used by TMVP (Temporal Motion Vector Prediction) and SbTMVP (Sub-block Temporal Motion Vector Prediction), motion field compression is performed at 8x8 size granularity, as opposed to 16x16 size granularity in HEVC.

[0046] Affine motion compensation prediction In HEVC, only translation motion model is applied to motion compensated prediction (MCP). In the real world, there are many kinds of motion, such as zoom in / out, rotation, perspective motion, and other irregular motion. In VVC, block-based affine transformation motion compensated prediction allows warping to be performed using motion compensation. As shown in Figure 4, the affine motion field of a block is described by the motion information of two control points (four-parameter affine model shown on the left) or three control point motion vectors (six-parameter affine model shown on the right).

[0047] For the four-parameter affine motion model, a motion vector (v,v) at a sample position (x,y) in a block (Cur) is derived as follows:

[0048]

number

[0049] For a six-parameter affine motion model, the motion vector at sample position (x,y) within a block (Cur) is derived as follows:

[0050]

number

[0051] To simplify the motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector of each 4x4 luma sub-block, the motion vector of the center sample of each sub-block is calculated according to the above formula and rounded to 1 / 16 fractional precision, as shown in Figure 5. Then, a motion compensation interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chroma components is also set to 4x4. The MV of a 4x4 chroma sub-block is calculated as the average of the MVs of the top-left and bottom-right luma sub-blocks in the collocated 8x8 luma region.

[0052] As is done for translational motion inter prediction, there are also two affine motion inter prediction modes: affine merge mode and affine AMVP (Advanced Motion Vector Prediction) mode.

[0053] Sub-block-based Temporal Motion Vector Prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses motion fields in a co-located picture to improve motion vector prediction and merge modes for CUs in a current picture. The same co-located picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following ways: TMVP predicts motion at the CU level, while SbTMVP predicts motion at the sub-CU level. TMVP fetches temporal motion vectors from a co-located block in a co-located picture (the co-located block is the bottom-right or center block with respect to the current CU), while SbTMVP applies a motion shift before fetching temporal motion information from the co-located picture, where the motion shift is obtained from a motion vector from one of the spatial neighboring blocks of the current CU.

[0054] The SbTMVP process is shown in Figure 6 and Figure 7. SbTMVP predicts the motion vectors of sub-CUs in the current CU in two steps. In the first step, the spatial neighborhood A1 in Figure 6 is examined. If A1 has a motion vector that uses a co-located picture as its reference picture, this motion vector is selected to be the motion shift to be applied. If no such motion is identified, the motion shift is set to (0, 0).

[0055] In the second step, the motion shift identified in the first step is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the co-located picture as shown in FIG. 7. The example of FIG. 7 assumes that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-located picture is used to derive the motion information for the sub-CU. After the motion information of the co-located sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector to the reference picture of the current CU.

[0056] In VVC, a combined sub-block-based merge list containing both SbTMVP candidates and affine merge candidates is used to signal the sub-block-based merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block-based merge candidates, followed by the affine merge candidates. The size of the sub-block-based merge list is signaled in the SPS, and the maximum allowed size of the sub-block-based merge list is 5 in VVC.

[0057] The sub-CU size used in SbTMVP is fixed to be 8x8, and as is done for the affine merge mode, the SbTMVP mode is only applicable to CUs with width and height greater than or equal to 8.

[0058] The encoding logic of the additional SbTMVP merge candidate is the same as that of the other merge candidates, i.e., for each CU in a P slice or B slice, an additional RD check is performed to determine whether the SbTMVP candidate should be used.

[0059] Interpolation Filters in VVC Table 1 below shows the 8-tap luma interpolation filter coefficients f for each 1 / 16 fractional sample position p. L [p] specifications are shown below.

[0060] [Table 1]

[0061] Interpolation Filters in ECM Software In the ECM software, the 8-tap interpolation filter used in VVC is replaced by a 12-tap filter. The interpolation filter is derived from a sinc function and its frequency response is cut off at the Nyquist frequency and truncated by a cosine window function. Table 2 below gives the filter coefficients for all 16 phases for the 12-tap interpolation filter. Figure 8 compares the frequency response of the interpolation filters, all at half-pixel phase, with the VVC interpolation filter.

[0062] [Table 2]

[0063] According to an aspect of the present principles, there is provided a method for encoding or decoding video, in which an adapted motion compensation filter is used for inter prediction.

[0064] 9 illustrates an example of a method 900 for encoding a block in a video, according to an embodiment. The block is encoded using inter prediction, for example, using the encoder described in relation to FIG. 2. It is assumed here that a reference block has already been identified for the block to be encoded. The reference block is a block area in a reference picture used by a current block for encoding, and the block area in the reference picture is determined by a motion vector associated with the current block.

[0065] At 910, a motion compensation interpolation filter (MCIF) is determined for the current block. In some variations, when the MCIF is separable, the horizontal MCIF and the vertical MCIF are determined separately. Some embodiments for determining the MCIF for the current block are described below. In some embodiments, the MICF is adapted according to a current block discontinuity condition. At 920, a prediction block is determined based on the determined MCIF and a reference block from the reference picture. The reference block is interpolated using the determined MCIF to provide the prediction block. At 930, the current block is encoded based on the prediction block.

[0066] Figure 10 shows an example of a method 1000 for decoding a block in a video according to an embodiment. The block is encoded using inter prediction and decoded, for example, using the decoder described with reference to Figure 3. As in the encoder, at 1010, an MCIF is determined for the block to be decoded. At 1020, a predictive block is determined based on the determined MCIF and a reference block, and at 1030, the current block is decoded / reconstructed using the predictive block.

[0067] According to the present principles, to improve inter prediction, a motion compensated interpolation filter (MCIF) is adapted based on a discontinuity condition as described below. According to a variant, the length of the MCIF is adapted at 910 in Fig. 9 or 1010 in Fig. 10. This can be done by selecting the MCIF from a set of filters with different lengths, for example an 8-tap filter for VVC and a 12-tap filter for ECM. However, other variants and filters are possible.

[0068] In another variation, the filter response is adapted using an alternative filter, typically smoother, for sub-blocks at discontinuous positions.

[0069] Location-based selection According to an embodiment, selecting the MCIF is based on the position of the block within an area larger than the block. When the block is a sub-block of a larger block or coding unit (CU) having a plurality of sub-blocks, the selection of the MCIF is based on the position of the sub-block within the coding unit.

[0070] FIG. 11 shows a 32×32 CU composed of 4×4 sub-blocks. It is assumed that the sub-block is a unit entity for motion compensation. In this embodiment, the filter used for motion compensation is adapted depending on the sub-block position. For the central sub-block (the white square in FIG. 11), a default filter, for example, an N-tap filter (N is an integer), is used. For sub-block "a", vertical filtering takes samples outside the CU, which is likely to be in an area that is not the same "object", so vertical filtering uses a shorter filter, for example, an M-tap filter (M is an integer where M < N). The horizontal filter is the default filter (N-tap filter). For the sub-block within "c", it is the reverse, the horizontal filter is shorter (M-tap), and the vertical filter is the default filter (N-tap). For the sub-block within "b", both the horizontal filter and the vertical filter are shorter than the default filter.

[0071] As an example, the default filter is the 12-tap filter described above, and the shorter filter is the 8-tap filter. Other filters can also be used.

[0072] In this embodiment, when the block is located at the inner boundary of the coding unit, at least one of the horizontal MCIF or the vertical MCIF has a length shorter than the MCIF used for the block when the block is inside the coding unit and not at the inner boundary of the coding unit (i.e., when the block is the white sub-block in FIG. 11).

[0073] In Figure 11, the inner boundary is one subblock wide. In other variations, the inner boundary may be thicker and have a width of more than one subblock. In this variation, the band with the shorter MCIF is larger than one subblock.

[0074] In another variation, several short filters are used, the filters becoming shorter when they are close to the block boundaries in one or both directions, in which the length of at least one of the horizontal or vertical MCIFs varies with the distance of the block to the center of the coding unit.

[0075] Motion-Based Filter Selection In another embodiment, the MCIF is selected based on the difference in motion between the motion of the block and the motion of at least one of the block's horizontal and vertical neighboring blocks.

[0076] The concept in this embodiment is generalized to any 4x4 sub-block based on the motion difference with that sub-block's neighbors, as shown in FIG.

[0077] Depending on the motion difference with its horizontal (and vertical) neighboring pixels, the length of the filter is adapted. When the motion is not available, the motion difference is assumed to be null. Method 1200 is an example of a method for MCIF selection based on the motion difference between a block's motion and its neighboring blocks. At 1210, the motion vector (ux, uy) of the current block is obtained. At 1220, the motion vector (lx, ly) of the block to the left of the block is obtained, and the motion vector (rx, ry) of the block to the right of the block is obtained if available. At 1230, the difference between the motion vector of the current block and the motion vector of its neighboring block is compared with a threshold th, for example as follows: if |ux-lx| is above a threshold, if |ux-rx| is above a threshold, if |uy-ly| is above a threshold, or if |uy-ry| is above a threshold, then at 1240, a short MCIF is selected for the horizontal direction. Otherwise, at 1250, a default filter is selected for the horizontal direction. The same logic can be applied to the vertical direction. At 1220, the motion vector (tx,ty) of the block above the block and the motion vector (bx,by) of the block below the block are obtained if available. At 1270, the difference between the motion vectors is compared to a threshold th. When the difference is above the threshold for at least one of the x or y coordinates, at 1280, a short MCIF is selected for the vertical direction. Otherwise, at 1290, a default filter is selected for the vertical direction.

[0078] When no motion is available, e.g., for the right and bottom blocks, the difference is considered null. In some cases, the motion of the right and bottom sub-blocks is available for sub-blocks within a coding unit for coding units with sub-block motion such as affine or SbTMVP. For sub-blocks at the right or bottom boundary of a coding unit, the motion of the right and bottom neighboring blocks is not available.

[0079] In a variant, for such sub-blocks (sub-blocks at the right or bottom boundary of a coding unit) a short filter is always selected.

[0080] In one variant, the filter selection is performed using neighboring reconstructed pixels of the current block to be coded or decoded. In this variant, a determination is made to determine whether an edge exists in the current block, for example by applying edge detection to neighboring reconstructed pixels of the current block. If an edge exists, a filter shorter than the default filter is selected to filter the reference block.

[0081] In another variant, all filters are applied to the neighboring templates of the current block, and the filter that allows the best reconstruction of the neighboring templates is selected for the current block.

[0082] The filter is considered as the best filter, for example, in terms of the quality reconstruction provided by the filter considered. The selection of the filter in this variant is therefore performed by selecting the filter that provides the smallest distortion between the pixels of the neighboring template of the current block and the pixels of the neighboring template of the reference block filtered by the filter considered. The neighboring template is, for example, a template from a template matching method, which includes neighboring reconstructed pixels in the band to the left of the current block and in the band above the current block.

[0083] In a variant, only the sub-blocks at the boundary of a coding unit use the above embodiment. The boundary may have a width of one or more sub-blocks.

[0084] Mode-Based Selection In another embodiment, the MCIF is selected based on the coding mode of at least one of the block's neighboring blocks. In this variant, the filter selection is based on the mode of the neighboring (sub)blocks, for example as shown in FIG. 13. The process starts at 1310. At 1320, the coding mode C0 of the block's left neighbor and the coding mode C1 of the block's right neighbor are obtained. At 1330, it is determined whether at least one of C0 or C1 is intra-coded. When at least one of C0 and C1 is intra-coded, at 1340, a short MCIF is selected for the horizontal direction. Otherwise, at 1350, a default filter is selected for the horizontal direction. The same logic can be applied to the vertical direction. At 1360, the coding mode C0 of the block's top neighboring block and the coding mode C1 of the block's bottom neighboring block are obtained. At 1370, it is determined whether at least one of C0 or C1 is intra-coded. When at least one of C0 and C1 is intra-coded, a short MCIF is selected for the vertical direction, at 1380. Otherwise, a default filter is selected for the vertical direction, at 1390.

[0085] In this embodiment, the length of the filter is adapted depending on the mode of the neighborhood: when a mode is not available, the mode is assumed to be non-intra.

[0086] In a variant, other modes, for example the presence of residuals (non-zero transform coefficients) in the neighboring sub-blocks are compared. If a residual exists for the left or right neighboring sub-block, a filter shorter than the default filter is selected for the horizontal direction. The same logic is applied vertically with the neighboring blocks above and below.

[0087] In a variant, only sub-blocks at block boundaries of a coding unit use the above embodiment.

[0088] In a variation, a combination of at least two of the above conditions (position, motion, mode) is used to determine the MCIF, in which a short filter is selected when at least one of the above conditions is true.

[0089] Deblocking filter condition selection

[0090] [Table 3]

[0091] Table 3 above shows the deblocking filter strength derivation for luma samples.

[0092] In this embodiment, the same logic is reused to select the MCIF of a block. The condition for selecting a filter shorter than the default filter can be inferred from at least one of the following: whether the neighboring block is intra-coded or not, whether the neighboring block has residuals (i.e., non-zero transform coefficients), whether the motion vectors of the current block and its neighboring block are different, and whether the reference pictures of the current block and its neighboring block are different.

[0093] The short filter may be selected when at least one of the conditions is met.

[0094] Selection based on affine models According to another embodiment, the MCIF is selected based on the motion difference between the control motion vectors when the block is coded in an affine model based coding mode.

[0095] In the case of affine inter prediction, the filter used for each subblock depends on the amplitude of the non-translated part of the motion. If the amplitude of the non-translated part is large, the motion difference between each subblock is large. Instead of testing the motion difference between each subblock of an affine block, a test is performed for the whole block based on the motion model coded with CPMV(v0,v1) for the affine 4-parameter model, or based on the motion model coded with (v0,v1,v2) for the affine 6-parameter model, as shown in FIG. 14. FIG. 14 shows an example of an MCIF selection method 1400 based on a 4-parameter affine model condition according to an embodiment. In 1410 and 1420, the control motion vectors (v0x,v0y) and (v1x,v1y) of the top-left and top-right of the block are obtained. In 1430, the amplitude of the control point motion vector is checked against a threshold th. If the amplitude is above the threshold for at least one of the coordinates, in 1440, a short MCIF is selected, otherwise in 1450, a default MCIF is used.

[0096] The same logic can also be applied to the control point motion vector cpmv v2 if a six-parameter affine motion model is available.

[0097] In a variation, the non-default MCIF selected in 1440 is a smoother filter rather than a shorter filter.In a variation, only sub-blocks on coding unit boundaries use the above embodiment.

[0098] Asymmetric MCIF In another embodiment, the MCIF is adapted more specifically to the discontinuity location by taking into account not only the direction (horizontal or vertical) but also the left / right (or top / bottom) discontinuities.

[0099] In this embodiment, an MCIF is determined for a block by creating a new MCIF from two filters having different lengths, at 910 in FIG. 9 or 1010 in FIG.

[0100] Phase 1 / 16 Filter Example Below, examples of asymmetric filter configurations using 8-tap and 12-tap filters are described. This example is intended to illustrate the principles. Other filters of the same or other lengths can also be used.

[0101] [Table 4]

[0102] Table 4 above gives examples of 8 tap (top row) and 12 tap (bottom row) filtering for phase 1 / 16 MCIF using the same scaling.

[0103] To create a filter with a shorter left part, two filters are concatenated using a short filter for the left part and a long filter for the right part (see Figure 15). The central coefficient is adjusted to obtain a unit filter (the sum of all coefficients should be 1, here it is 256 after quantization).

[0104] The same logic is applied to create a filter with a shorter right portion, as shown in FIG.

[0105] Using Asymmetric Filters In FIG. 17, an example of a method 1700 for horizontal filter selection is provided.

[0106] At 1710, a motion vector (ux, uy) of a current block is obtained. At 1720, a motion vector (lx, ly) of a block to the left of the current block and a motion vector (rx, ry) of a block to the right of the current block are obtained. At 1730, an asymmetric filter is selected based on the conditions provided in Table 5 below.

[0107] [Table 5]

[0108] The condition to be checked (|ux-lx|>th or |uy-ly|>th) is shown in the first row of Table 5. As shown in Table 5, if both conditions are false, then a long filter or a default filter is selected for the block. When both conditions are true, then a short filter is selected. Otherwise, an asymmetric filter is selected, where the left or right part is shorter than the other part, depending on the condition.

[0109] The same logic as shown in FIG. 17 and Table 5 is applied to vertical filter selection by checking sub-blocks above and below instead of left and right.

[0110] The same logic extends to other criteria for filter selection.

[0111] syntax A flag sps_adapted_mcif can be coded to signal whether an adapted filter is used at the CU / region / slice / picture / sequence level, possibly also inherited at the CU level. The following table shows an example of a syntax for signaling whether any one of the embodiments allowing to use an adapted MICF is enabled.

[0112] Sequence Parameter Set RBSP Syntax

[0113] [Table 6]

[0114] Slice Header Syntax

[0115] [Table 7]

[0116] In the slice header of a non-intra slice, a flag sh_adapted_mcif is signaled that controls the usage of the tool.

[0117] In one embodiment shown in FIG. 18, in a transmission context between two remote devices A and B over a communications network NET, device A comprises a processor associated with memory RAM and ROM configured to implement a method for encoding video as described in FIG. 1 to FIG. 17, and device B comprises a processor associated with memory RAM and ROM configured to implement a method for decoding video as described in relation to FIG. 1 to FIG. 17.

[0118] According to one embodiment, the network is a broadcast network adapted to broadcast / transmit encoded data representing video from device A to decoding devices, including device B.

[0119] The signal intended to be transmitted by device A carries at least one bitstream including coded data representing video. The bitstream may be generated from any embodiment of the present principles.

[0120] 19 shows an example of the syntax of such a signal transmitted over a packet-based transmission protocol. Each transmission packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include coded video data encoded according to any one of the embodiments described above.

[0121] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence to produce a final output suitable for, e.g., a display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders in various implementations described herein, such as, e.g., decoding resampling filter coefficients, resampling a decoded picture, etc.

[0122] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstruction picture process including entropy decoding. Whether the phrase "decoding process" is intended to refer to a task subset specifically or to the broader decoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.

[0123] Various implementations involve encoding. Similar to the above discussion regarding "decoding," "encoding" as used herein can encompass all or part of the processes performed on an input video sequence to produce, for example, an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by encoders in various implementations, such as those described herein, such as, for example, determining resampling filter coefficients, resampling decoded pictures, etc.

[0124] As a further example, in one embodiment, "encoding" refers to entropy encoding only, in another embodiment, "encoding" refers to differential encoding only, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.

[0125] Please note that syntax elements as used herein are descriptive terms, and therefore do not exclude the use of other syntax element names.

[0126] This disclosure has described various information, such as syntax, that may be transmitted or stored. This information may be packaged or arranged in various manners, including, for example, manners common in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or slice header), or SEI message. Other manners are also available, including, for example, manners common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for purposes of session announcement and session invitation, e.g., as described in the RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport. b. DASH MPD (Media Presentation Description) Descriptors, e.g., as used in DASH and transmitted over HTTP, Descriptors are associated with a representation or a collection of representations to provide additional characteristics to the content representation. c. RTP header extensions, for example as used during RTP streaming. d. The ISO Base Media File Format, such as that used in OMAF and in some specifications, which uses boxes, which are object-oriented building blocks defined by a unique type identifier and a length, also known as "atoms". e. HTTP Live Streaming (HLS) manifests transmitted over HTTP. A manifest can be associated with a version or collection of versions of content, for example to provide characteristics of the version or collection of versions.

[0127] If a diagram is presented as a flow diagram, it should be understood that the diagram also provides a block diagram of the corresponding device. Similarly, if a diagram is presented as a block diagram, it should be understood that the diagram also provides a flow diagram of the corresponding method / process. Some embodiments refer to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is usually considered, often due to computational complexity constraints. Rate-distortion optimization is usually formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these techniques can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, but with a full evaluation of their encoding costs and associated distortion of the reconstructed signal after encoding and decoding. To keep the encoding complexity down, more rapid techniques can also be used, especially with the calculation of approximate distortion based on the prediction or prediction residual signal rather than the reconstructed signal. A mixture of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete assessment of both the coding cost and the associated distortion.

[0128] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed in the context of only a single implementation form (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). For example, an apparatus may be implemented in appropriate hardware, software, and firmware. The method may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0129] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation," as well as other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0130] It should be noted that the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0131] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0132] It should be noted that the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some manner an operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0133] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of only the third enumerated alternative (C), or the selection of only the first and second enumerated alternatives (A and B), or the selection of only the first and third enumerated alternatives (A and C), or the selection of only the second and third enumerated alternatives (B and C), or the selection of all three alternatives (A and B and C). This can be expanded as many times as the items listed, as would be apparent to one of ordinary skill in this and related arts.

[0134] Also, as used herein, the term "signaling" means to indicate something to a corresponding decoder, among other things. For example, in a particular embodiment, an encoder signals a particular one of a number of resampling filter coefficients. Thus, in an embodiment, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder can transmit a particular parameter to a decoder (explicit signaling) so that the decoder can use the same particular parameter. In contrast, if the decoder already has the particular parameter as well as other parameters, a non-transmitting signaling (implicit signaling) can be used to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual function, bit savings are realized in various embodiments. It will be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to a corresponding decoder. The above relates to the verb form of the word "signal", which may also be used as a noun in this specification.

[0135] As will be apparent to one skilled in the art, implementations can result in a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bit stream of the described embodiments. For example, such a signal can be formatted as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information that the signal carries can be, for example, analog information or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored in a processor-readable medium.

[0136] Many embodiments are described herein. The features of these embodiments may be provided alone or in any combination across various claim categories and types. Furthermore, the embodiments may include one or more of the following features, devices, or aspects across various claim categories and types, alone or in any combination: A video encoding / decoding method comprising adapting a motion compensated interpolation filter based on a block discontinuity condition. Encoding / decoding a block of video comprising adapting a motion compensated interpolation filter based on a position of the block in a coding unit. · Encoding / decoding a block of video comprising adapting a motion compensated interpolation filter based on the motion of the block and the difference in motion between its neighboring blocks. · Encoding / decoding a block of a video comprising adapting a motion compensated interpolation filter based on a coding mode of a neighboring block. · Encoding / decoding a block of video comprising adapting a motion compensated interpolation filter based on neighboring reconstructed pixels of the block. · Encoding / decoding a block of video comprising adapting a motion compensated interpolation filter based on a motion difference between control motion vectors when the block is coded in an affine model based coding mode. · Encoding / decoding a block of video comprising creating an asymmetric motion compensated interpolation filter based on a discontinuity condition of the block. · Encoding / decoding a block of video comprising creating an asymmetric motion compensated interpolation filter from two filters having different lengths. In response to determining that a block is at a coding unit boundary, adapting a motion compensated interpolation filter based on a discontinuity condition of the block. In response to determining that the block is at a coding unit boundary, generating an asymmetric motion compensated interpolation filter based on a discontinuity condition of the block. · Encoding / decoding video comprising signaling information enabling / disabling use of an adapted motion compensation interpolation filter. A bitstream or signal comprising one or more of the described syntax elements or variations thereof. A bitstream or signal including syntax carrying information generated according to any of the described embodiments. Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal comprising one or more of the described syntax elements or variations thereof. Creating and / or transmitting and / or receiving and / or decoding according to any of the described embodiments. A method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device performing video decoding according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that performs decoding of video according to any of the described embodiments and displays the resulting images (e.g., using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal containing encoded images and performs video decoding according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that receives an over-the-air signal (e.g., using an antenna) containing encoded images and performs video decoding according to any of the described embodiments.

Claims

1. A method comprising decoding a block of a video coding unit, wherein the coding unit has a plurality of blocks, and the method Determining at least one motion compensation interpolation filter for the block from a set of motion compensation interpolation filters having different lengths, wherein the length of the motion compensation interpolation filter depends on the position of the block in the coding unit, Determining a prediction block based on the aforementioned at least one motion compensation interpolation filter and reference block, Decoding the block based on the prediction block, A method that includes this.

2. A device comprising one or more processors, wherein the one or more processors are configured to decode blocks of a video coding unit, and the coding unit has a plurality of blocks, and the decoding of the blocks is Determining at least one motion compensation interpolation filter for the block from a set of motion compensation interpolation filters having different lengths, wherein the length of the motion compensation interpolation filter depends on the position of the block in the coding unit, Determining a prediction block based on the aforementioned at least one motion compensation interpolation filter and reference block, Decoding the block based on the prediction block, A device that includes this.

3. A method comprising encoding a block of a video coding unit, wherein the coding unit has a plurality of blocks, and the method Determining at least one motion compensation interpolation filter for the block from a set of motion compensation interpolation filters having different lengths, wherein the length of the motion compensation interpolation filter depends on the position of the block in the coding unit, Determining a prediction block based on the aforementioned at least one motion compensation interpolation filter and reference block, Encoding the block based on the prediction block, A method that includes this.

4. A device comprising one or more processors, wherein the one or more processors are configured to encode blocks of a video coding unit, and the coding unit has a plurality of blocks, and encoding the blocks is: Determining at least one motion compensation interpolation filter for the block from a set of motion compensation interpolation filters having different lengths, wherein the length of the motion compensation interpolation filter depends on the position of the block in the coding unit, Determining a prediction block based on the aforementioned at least one motion compensation interpolation filter and reference block, Encoding the block based on the prediction block, A device that includes this.

5. The method according to claim 1 or 3, wherein determining the at least one motion compensation interpolation filter includes selecting the at least one motion compensation interpolation filter from the set of motion compensation interpolation filters.

6. Selecting at least one motion compensation interpolation filter means Select a horizontal motion compensation interpolation filter to filter the aforementioned reference block horizontally, The method according to claim 5, comprising selecting a vertical motion compensation interpolation filter for vertically filtering the reference block.

7. The method according to claim 6, wherein, when the block is located at the internal boundary of the coding unit, at least one of the horizontal motion compensation interpolation filter or the vertical motion compensation interpolation filter has a shorter length than the motion compensation interpolation filter selected for the block when the block is inside the coding unit and the block is not at the internal boundary of the coding unit.

8. The method according to claim 7, wherein, if the block is not located on the internal boundary of the coding unit, the motion compensation interpolation filter selected for the block is a 12-tap filter.

9. The method according to claim 6, wherein, when the block is located at the internal boundary of the coding unit, at least one of the horizontal motion compensation interpolation filters or vertical motion compensation interpolation filters selected for the block is an 8-tap filter.

10. The method according to claim 7, wherein the internal boundary of the coding unit is a bandwidth having the width of at least one subblock.

11. The method according to claim 6, wherein the length of at least one of the horizontal motion compensation interpolation filter or the vertical motion compensation interpolation filter varies depending on the distance of the block to the center of the coding unit.

12. The method according to claim 11, wherein the length of the motion compensation interpolation filter is shorter for blocks closer to the center of the coding unit.

13. The method according to claim 1 or 3, further comprising decoding or encoding information to enable or disable the use of the motion compensation interpolation filter.

14. The apparatus according to claim 2 or 4, wherein determining the at least one motion compensation interpolation filter comprises selecting the at least one motion compensation interpolation filter from the set of motion compensation interpolation filters.

15. Selecting at least one motion compensation interpolation filter means that Select a horizontal motion compensation interpolation filter to filter the aforementioned reference block horizontally, The apparatus according to claim 14, further comprising selecting a vertical motion compensation interpolation filter for filtering the reference block vertically.

16. The apparatus according to claim 15, wherein when the block is located at the internal boundary of the coding unit, at least one of the horizontal motion compensation interpolation filter or the vertical motion compensation interpolation filter has a shorter length than the motion compensation interpolation filter selected for the block when the block is inside the coding unit and the block is not at the internal boundary of the coding unit.

17. The apparatus according to claim 16, wherein, if the block is not located on the internal boundary of the coding unit, the motion compensation interpolation filter selected for the block is a 12-tap filter.

18. The apparatus according to claim 15, wherein, when the block is located at the internal boundary of the coding unit, at least one of the horizontal motion compensation interpolation filters or vertical motion compensation interpolation filters selected for the block is an 8-tap filter.

19. The apparatus according to claim 16, wherein the internal boundary of the coding unit is a bandwidth having the width of at least one subblock.

20. The apparatus according to claim 16, wherein the length of at least one of the horizontal motion compensation interpolation filter or the vertical motion compensation interpolation filter varies depending on the distance of the block to the center of the coding unit.

21. The apparatus according to claim 20, wherein the length of the motion compensation interpolation filter is shorter for blocks closer to the center of the coding unit.

22. The apparatus according to claim 2 or 4, further comprising decoding or encoding information to enable or disable the use of the motion compensation interpolation filter.

23. A computer-readable storage medium storing instructions for causing one or more processors to perform the method described in claim 1 or 3.

24. It is a device, The apparatus according to claim 2, A device comprising: (i) an antenna configured to receive a signal, wherein the signal includes data representing video; (ii) a band limiter configured to restrict the signal to a frequency band including the data representing video; or (iii) a display configured to display the video.