Encoders, decoders, methods, and video bitstreams for hybrid video coding, as well as computer programs
The encoder and decoder system adapts interpolation filters based on video sequence characteristics to optimize encoding efficiency and quality, addressing the trade-off challenges in video coding standards by reducing bitstream size and computational effort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
- Filing Date
- 2024-08-28
- Publication Date
- 2026-04-27
Smart Images

Figure 0007851998000004 
Figure 0007851998000005 
Figure 0007851998000006
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to encoders and decoders for hybrid video coding. Further embodiments of the present invention relate to a method for hybrid video coding. Further embodiments of the present invention relate to video bitstreams for hybrid video coding. [Background technology]
[0002] All modern video coding standards use motion-compensated prediction (MCP) with fractional sample precision for motion vectors (MV). Interpolation filters are used to determine the sample value of the reference picture at fractional sample positions.
[0003] The High Efficiency Video Coding (HEVC) standard, which uses 1 / 4 lumar sample motion vector accuracy, has 15 fractional lumar sample positions that are computed using various one-dimensional finite impulse response (FIR) filters. However, for each fractional sample position, the interpolation filter is fixed.
[0004] Alternatively, instead of having a fixed interpolation filter for each fractional sample position, it is also possible to select an interpolation filter for each fractional sample position from a set of interpolation filters. Techniques that allow switching between different interpolation filters at the slice level include Non-separable Adaptive Interpolation Filter (AIF)[1], Separable Adaptive Interpolation Filter (SAIF)[2], Switched Interpolation Filtering with Offset (SIFO)[3], and Enhanced Adaptive Interpolation Filter (EAIF)[4]. The interpolation filter coefficients are explicitly signaled for each slice (e.g., AIF, SAIF, EAIF), or (as in the case of SIFO) it is indicated which interpolation filter from a given set of interpolation filters will be used.
[0005] The current draft of the latest Versatile Video Coding (VVC) standard[5] supports so-called Adaptive Motion Vector Resolution (AMVR). Using AMVR, the accuracy of both motion vector predictor (MVP) and motion vector difference (MVD) can be selected together from the following possibilities: -Quarter sample (QPEL): Motion vector accuracy of one-quarter lumens. - Full sample level (FPEL): Motion vector accuracy of lumens sample, -4 samples (4PEL): Motion vector accuracy of 4 lumens samples.
[0006] Additional syntax elements are sent in the bitstream to select the desired precision. In the case of biprediction, the indicated precision applies to both list0 and list1. AMVR is not available in skip mode and merge mode. Furthermore, AMVR is not available if all MVDs are equal to 0. If a precision other than QPEL is used, the MVP is rounded to the initially specified precision before adding the MVDs to obtain the MV.
[0007] Some embodiments of this disclosure may implement or be used in the context of one or more of the concepts described above. For example, embodiments may implement the HEVC or VVC standard and may use AMVR and the MV accuracy described above.
[0008] A concept for encoding, decoding, and transmitting video data is desired that offers an improved trade-off between bitrate, complexity, and the achievable quality of hybrid video coding. [Overview of the project]
[0009] Aspects of this disclosure relate to encoders for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the encoder provides an encoded representation of a video sequence based on input video content, the encoder determines one or more syntax elements (e.g., motion vector precision syntax elements) associated with a portion of the video sequence (e.g., a block), and, based on the characteristics described by the one or more syntax elements (e.g., decoder settings signaled by the one or more syntax elements, or characteristics of the portion of the video sequence signaled by the one or more syntax elements), a processing scheme (e.g., interpolation filter) applied to the portion of the video sequence (e.g., by the decoder). The selected processing method is for obtaining samples (e.g., samples, e.g., fractional samples) for motion compensation prediction at integer and / or fractional positions within a portion of a video sequence, and encodes an index indicating the selected processing method (depending on one or more syntax elements) such that a given encoded index value represents a different processing method (e.g., related) depending on the characteristics described by one or more syntax elements, and is configured to provide a bitstream containing one or more syntax elements and the encoded index as an encoded representation of the video sequence (e.g., send, upload, save as a file).
[0010] By selecting a processing method to apply to a part based on its characteristics, it becomes possible to adapt the applied processing method to the part according to its characteristics. For example, different processing methods can be selected for different parts of a video sequence. For example, a particular processing method may be particularly beneficial for encoding a part having given characteristics, because it can be adapted to facilitate particularly efficient encoding and / or to facilitate particularly high quality of the encoded representation of a part having given characteristics. Therefore, by selecting a processing method according to the characteristics of a part (for example, by adapting a set of possible processing methods to the characteristics of the part), the overall efficiency and quality of encoding a video sequence can be improved simultaneously. Encoding efficiency may refer to low computational effort in encoding, or it may refer to a high compression ratio of the encoded information.
[0011] Proper decoding of a portion can be ensured by providing an encoded index indicating the selected processing method. Since a given value of the encoded index can represent different processing methods depending on the characteristics of the portion, the encoder can adapt the encoding of the index according to the characteristics. For example, the number of possible values of the encoded index may be limited to the number of processing methods applicable to a portion having that characteristic. Therefore, even if the total number of selectable processing methods (e.g., related to different characteristics) is large, the encoded index may require fewer bits (this can be adapted, for example, to the number of selectable processing methods in terms of the actual characteristics of the portion currently under consideration). Thus, the encoder can leverage its knowledge of the portion's characteristics by determining one or more syntax elements to efficiently encode the index. Therefore, the combination of selecting a processing method according to a characteristic (a characteristic can reduce the number of possible processing methods for the portion currently under consideration compared to the total number of different processing methods available) and encoding the index according to the characteristic can improve the overall efficiency of encoding, while the selected processing method may be indicated within the bitstream.
[0012] According to one embodiment, the encoder is configured to adapt mapping rules that map index values to encoded index values according to a characteristic described by one or more syntax elements. By adapting the mapping rules according to the characteristic, it can be ensured that an index can be encoded such that a given encoded index value represents a different processing method depending on the characteristic.
[0013] According to one embodiment, the encoder is configured to adapt the number of bins used to provide an encoded index indicating a selected processing method, depending on a characteristic described by one or more syntax elements. Thus, the encoder can adapt the number of bins to the number of processing methods acceptable for the characteristic, and as a result, the encoder can use fewer bins when a small number of processing methods are acceptable for the characteristic, thus reducing the size of the bitstream.
[0014] According to one embodiment, the encoder is configured to determine a set of acceptable processing methods in terms of characteristics described by one or more syntax elements, and to select a processing method from the determined set. By determining a set of acceptable processing methods in terms of characteristics, the encoder can infer information about selectable processing methods from one or more syntax elements. Thus, the encoder uses information already available from determining one or more syntax elements to select a processing method, and therefore increases the efficiency of encoding. Furthermore, by relying on a set of processing methods to select a processing method, the number of acceptable processing methods can be limited to the set of processing methods, thereby enabling the encoding of the index to a smaller size of the encoded index.
[0015] According to one embodiment, the encoder is configured to determine whether a single processing method is acceptable for a portion of a video sequence, and to selectively omit the inclusion of encoded indices if it is found that only a single processing method (e.g., a spatial filter or an interpolation filter) is acceptable for the portion of the video sequence. Thus, the encoder can, for example, infer from the characteristics of the portion that a single processing method is acceptable for the portion, and select a single processing method as the processing method to be applied to the portion. Omitting the inclusion of encoded indices in the bitstream can reduce the size of the bitstream.
[0016] According to one embodiment, the encoder is configured to select a mapping rule for mapping index values to encoded index values such that the number of bins representing encoded index values matches the number of acceptable processing methods for a portion of the video sequence. Thus, the encoder can adjust the number of bins to match the number of acceptable processing methods for a portion, and as a result, the encoded index may have fewer bins if fewer processing methods are acceptable for a portion, thus reducing the size of the bitstream.
[0017] According to one embodiment, in an encoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0018] According to one embodiment, in the encoder, the filters in the set of filters are separably applicable to parts of the video sequence (e.g., one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or y-direction). Thus, filters can be selected individually, increasing the flexibility for filter selection and, as a result, the set of filters can be very precisely fitted to the characteristics of the part.
[0019] According to one embodiment, the encoder is configured to select different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or picture). Thus, the selection of processing methods can take into account the individual characteristics of different blocks, and therefore facilitate improvements in the encoding of blocks, or coding efficiency, or reduction in the size of the bitstream.
[0020] According to one embodiment, the encoder is configured to encode an index indicating a selected processing method such that a given encoded index value represents (e.g., is associated with) different processing methods according to motion vector precision (which can be signaled using, for example, syntax elements within a bitstream and can take on, for example, one of the following values: QPEL, HPEL, FPEL, 4PEL). Thus, the encoding of the index can depend on the MV precision, e.g., the number of acceptable processing methods for the MV precision. For some settings of the MV precision, the number of acceptable processing methods may be less than others, which is an efficient way to reduce the size of the bitstream.
[0021] According to one embodiment, in the encoder, one or more syntax elements include at least one of motion vector precision (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translational inter, affine inter, translational merge, affine merge, combined inter / intra), availability of the coded residual signal (e.g., coded block flag), spectral characteristics of the coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edge from co-located prediction unit PUs within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, determination of deblocking filter and boundary strength), motion vector length (e.g., enabling only long motion vectors or additional smoothing filters for a particular direction), or adaptive motion vector resolution mode.
[0022] According to one embodiment, the encoder is configured to encode an index indicating a selected processing method (e.g., according to one or more syntax elements) such that a given encoded index value represents (e.g., is associated with) different processing methods according to the fractional part (or parts) of the motion vector. (For example, different encoding methods are used depending on whether all motion vectors in the part considered point to integer sample positions, e.g., luma sample positions.) Thus, the encoder can adapt the encoding of the index to the number of acceptable processing methods for the fractional part of the motion vector. Since the fractional part of the motion vector can be determined from the part, this is an efficient way to reduce the size of the bitstream.
[0023] According to one embodiment, the encoder is configured to selectively determine a set of processing methods (e.g., "encoding of multiple possibilities") for a motion vector accuracy (e.g., sub-sample motion vector accuracy, e.g., HPEL) between a maximum motion vector accuracy (e.g., a finer sub-sample motion vector resolution than the aforementioned motion vector accuracy, e.g., QPEL) and a minimum motion vector accuracy (e.g., 4PEL), or for a motion vector accuracy (e.g., HPEL) between a maximum motion vector accuracy (e.g., QPEL) and a full-sample motion vector accuracy (e.g., FPEL), and to select a processing method from the determined set. For example, at maximum MV accuracy, there may be one particularly efficient processing method, so it may be beneficial not to determine a set of processing methods, but at maximum or full-sample MV accuracy, none of the selectable processing methods may be applicable, so the selection of a processing method may not be necessary. In contrast, in the case of MV accuracy between maximum MV accuracy and full-sample or minimum MV accuracy, it may be particularly beneficial to adapt the processing method to the characteristics of the part.
[0024] According to one embodiment, the encoder is configured to encode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering), a second FIR filtering, and a third FIR filtering.
[0025] According to one embodiment, the encoder is configured to encode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering) and a second FIR filtering.
[0026] According to one embodiment, the encoder is configured to select (e.g., switch) between processing methods (e.g., interpolation filters) having different characteristics (e.g., stronger low-pass characteristics versus weaker low-pass characteristics). For example, a processing method or filter having strong low-pass characteristics can attenuate high-frequency noise components, thereby improving the quality of the encoded representation of the video sequence.
[0027] According to one embodiment, the encoder is configured to select a mapping rule depending on the motion vector precision and the fractional sample position or fractional part (or multiple fractional parts) of the motion vector. Thus, the encoder can adapt the encoding of the index to the number of acceptable processing methods for the MV precision and fractional part of the motion vector. Since the MV precision and fractional part of the motion vector can be determined from one or more syntax elements, it can be very efficient to use these characteristics to determine the number of acceptable processing methods and select a mapping rule accordingly. According to one embodiment, the encoder is configured to determine a set of available processing methods (e.g., interpolation filters) depending on the motion vector accuracy.
[0028] According to one embodiment, the encoder is configured to select between a quarter-sample motion vector resolution, a half-sample motion vector resolution, a full-sample motion vector resolution, and a four-sample motion vector resolution (for example, an index element or interpolation filter describing the processing method is selectively included when the half-sample motion vector resolution is selected, and the index element describing the processing method is omitted in other cases, for example).
[0029] According to one embodiment, the encoder comprises one or more FIR filters. The FIR filters are particularly stable filters for processing video sequences.
[0030] Another aspect of the present disclosure relates to an encoder for hybrid video coding (e.g., video coding with prediction and transformation coding of prediction error), wherein the encoder provides an encoded representation of a video sequence based on input video content, the encoder determines one or more syntax elements (e.g., motion vector precision syntax elements) associated with a portion of the video sequence (e.g., a block), and selects a processing method (e.g., interpolation filter) to be applied to the portion of the video sequence (e.g., by the decoder) based on the characteristics described by the one or more syntax elements (e.g., decoder settings signaled by the one or more syntax elements, or characteristics of the portion of the video sequence signaled by the one or more syntax elements), the processing method being applied to the portion of the video sequence at integer and / or fractional positions within the portion of the video sequence. The system is configured to obtain samples for compensation prediction (e.g., samples, fractional samples, etc., related to a video sequence), and to select an entropy coding scheme used to provide an encoded index indicating a selected processing scheme depending on a characteristic described by one or more syntax elements (e.g., depending on one or more syntax elements) (e.g., the number of bins described by the entropy coding scheme may be equal to zero if, in terms of one or more syntax elements considered, only one processing scheme is acceptable), and to provide a bitstream (e.g., send, upload, save as file) as an encoded representation of the video sequence, which includes one or more syntax elements and an encoded index (e.g., if the number of selected bins is greater than zero).
[0031] Selecting a processing method to apply to a portion based on its characteristics provides the same functionality and advantages as described in previous embodiments. Furthermore, by selecting an entropy coding method according to its characteristics, the provision of the encoded index can be adapted to those characteristics. For example, the number of possible values for the encoded index may be limited to the number of processing methods applicable to the portion having that characteristic. Thus, the encoded index may require a small number of bits in the bitstream, or even no bits at all, even if there are many processing methods available for different portions. Therefore, selecting an entropy coding method based on knowledge of the portion's characteristics available to the encoder by determining one or more syntax elements can provide an efficient provision of the encoded index. Thus, selecting both a process method and an entropy coding method according to its characteristics may improve the overall efficiency of encoding, even if the selected processing method is shown in the bitstream.
[0032] According to one embodiment, in an encoder, the entropy coding scheme includes a binarization scheme for providing an encoded index. Since the entropy coding scheme is selected according to its characteristics, the binarization scheme can be adapted according to its characteristics, and as a result, the binarization scheme can provide a particularly short representation of the encoded index.
[0033] According to one embodiment, in an encoder, the binarization scheme includes the number of bins used to provide an encoded index. Therefore, the number of bins used to provide an encoded index can be adapted according to characteristics, and as a result, the encoder may select a small number of bins or be unable to select any bins to provide an encoded index in order to provide efficient encoding.
[0034] According to one embodiment, the encoder is configured to adapt mapping rules that map index values to encoded index values, depending on a characteristic described by one or more syntax elements.
[0035] According to one embodiment, the encoder is configured to adapt the number of bins used to provide an encoded index indicating a selected processing method, depending on a characteristic described by one or more syntax elements.
[0036] According to one embodiment, the encoder is configured to determine a set of acceptable processing methods in terms of characteristics described by one or more syntax elements, and to select a processing method from the determined set.
[0037] According to one embodiment, the encoder is configured to determine whether a single processing method is acceptable for a portion of a video sequence, and to selectively omit the inclusion of an encoded index if it is found that only a single processing method (e.g., a spatial filter or an interpolation filter) is acceptable for the portion of the video sequence.
[0038] According to one embodiment, the encoder is configured to select a mapping rule that maps index values to encoded index values such that the number of bins representing encoded index values is suitable for a number of acceptable processing methods for a portion of the video sequence.
[0039] According to one embodiment, in an encoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0040] According to one embodiment, in the encoder, the filters in the set of filters are separably applicable to parts of the video sequence (for example, one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or the y-direction).
[0041] According to one embodiment, the encoder is configured to select different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or "picture").
[0042] According to one embodiment, the encoder is configured to encode an index indicating a selected processing method (e.g., depending on one or more syntax elements) such that a given encoded index value represents (e.g., associated with) a different processing method depending on motion vector precision (which can be signaled, for example, using syntax elements in a bitstream, and can take, for example, one of the following values: QPEL, HPEL, FPEL, 4PEL).
[0043] According to one embodiment, the encoder includes at least one of the following syntax elements: motion vector accuracy (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translation inter, affine inter, translation merge, affine merge, combined inter / intra), availability of coded residual signal (e.g., coded block flag), spectral characteristics of coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edges from co-located prediction unit PU within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, deblocking filter determination and boundary intensity), motion vector length (e.g., enable only long motion vectors or additional smoothing filters for specific directions), or adaptive motion vector resolution mode.
[0044] According to one embodiment, the encoder is configured to encode an index indicating a selected processing method (for example, depending on one or more syntax elements) such that a given encoded index value represents (for example, associated with) a different processing method depending on the fractional part (or multiple fractional parts) of the motion vector (for example, the different encoding methods used depend on whether all motion vectors of the part under consideration point to integer sample positions, e.g., lumen sample positions).
[0045] According to one embodiment, the encoder is configured to selectively determine a set of processing methods (e.g., "encoding of multiple possibilities") for motion vector precision between maximum motion vector precision (e.g., subsample motion vector resolution finer than the aforementioned motion vector precision, e.g., QPEL) and minimum motion vector precision (e.g., 4PEL) (e.g., subsample motion vector precision, e.g., HPEL), or for motion vector precision between maximum motion vector precision (e.g., QPEL) and full sample motion vector precision (e.g., FPEL) (e.g., HPEL), and to select a processing method from the determined set.
[0046] According to one embodiment, the encoder is configured to encode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering), a second FIR filtering, and a third FIR filtering.
[0047] According to one embodiment, the encoder is configured to encode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering) and a second FIR filtering.
[0048] According to one embodiment, the encoder is configured to select (e.g., switch) between processing methods (e.g., interpolation filters) having different characteristics (e.g., stronger low-pass characteristics versus weaker low-pass characteristics).
[0049] According to one embodiment, the encoder is configured to select a mapping rule depending on the motion vector accuracy and the fractional sample position or fractional part (or multiple fractional parts) of the motion vector. According to one embodiment, the encoder is configured to determine a set of available processing methods (e.g., interpolation filters) depending on the motion vector accuracy.
[0050] According to one embodiment, the encoder is configured to select between a quarter-sample motion vector resolution, a half-sample motion vector resolution, a full-sample motion vector resolution, and a four-sample motion vector resolution (for example, an index element or interpolation filter describing the processing method is selectively included when the half-sample motion vector resolution is selected, and the index element describing the processing method is omitted in other cases, for example). According to one embodiment, the encoder includes one or more FIR filters in its processing method.
[0051] Another aspect of the present disclosure relates to a decoder for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the decoder provides output video content based on an encoded representation of a video sequence, the decoder takes a bitstream as an encoded representation of the video sequence (e.g., receive, download, open from file), identifies one or more syntax elements (e.g., motion vector precision syntax elements) from the bitstream relating to a portion of the video sequence (e.g., a block) (and an encoded index), identifies a processing scheme (e.g., an interpolation filter selected by the encoder) for the portion of the video sequence based on one or more syntax elements and the index (which may be contained in the bitstream in an encoded form), the processing scheme is for taking samples for motion compensation predictions (e.g., samples, e.g., fractional samples) at integer and / or fractional positions within the portion of the video sequence, the decoder is configured to assign different processing schemes to a given encoded index value depending on one or more syntax elements (e.g., depending on the value of one or more syntax elements), and is configured to apply the identified processing scheme to the portion of the video sequence.
[0052] The decoder is based on the idea of identifying a processing scheme for a portion based on syntax elements and indices, but a given encoded index value representing an index value may be associated with various processing schemes. Identification of a processing scheme can be achieved by assigning a processing scheme to an encoded index value according to the syntax elements. Therefore, the decoder can use the information provided by the syntax elements to identify a processing scheme from the different processing schemes that may be associated with an encoded index value. Since a given encoded index value can be assigned to several processing schemes, the information density of the bitstream may be particularly high, and the decoder can still identify the processing scheme.
[0053] According to one embodiment, in the decoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0054] According to one embodiment, in the decoder, the filters in the set of filters are separably applicable to parts of the video sequence (for example, one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or the y-direction).
[0055] According to one embodiment, the decoder is configured to select (e.g., apply) a filter (e.g., an interpolation filter) in accordance with a fractional part (or multiple fractional parts) of a motion vector (e.g., in accordance with one or more fractional parts of a motion vector) (e.g., by selecting a first filter if the endpoint of the motion vector is at a full sample position or at a full sample position in the direction under consideration, and by selecting a second filter different from the first filter if the endpoint of the motion vector is between full samples or between full samples in the direction under consideration).
[0056] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction (e.g., x) and for filtering in a second direction (e.g., y) depending on the fractional part (or multiple fractional parts) of the motion vector.
[0057] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction and for filtering in a second direction, such that stronger low-pass filtering is applied to filtering in directions where one, more, or all fractional parts (or multiple fractional parts) of the motion vector are on the full sample (e.g., when the exact motion vector(s) of the half-sample point to the full sample position(s)) (compared to filtering applied in directions where the fractional parts(s) of the motion vector are between full samples).
[0058] According to one embodiment, the decoder is configured to select a filter (e.g., an interpolation filter) signaled by a syntax element (e.g., an interpolation filter index if_idx) or a predetermined filter, depending on one or more fractional parts (or more fractional parts) of a motion vector (e.g., to apply, e.g., individually for each filtering direction) (for example, the filter signaled by the syntax element is selected for filtering in the first direction, for example, when considering the position coordinates of a first direction, e.g., when one or more fractional parts (or more fractional parts) of the motion vector are between full samples, and the predetermined filter is selected for filtering in a second direction different from the first direction, for example, when considering the position coordinates of a second direction, e.g., when all fractional parts (or more fractional parts) of the motion vector are on full samples).
[0059] According to one embodiment, the decoder is configured to select (for example, individually for each filtering direction) a first filter (e.g., interpolation filter) signaled by a first syntax element (e.g., interpolation filter index if_idx) or a second filter signaled by a filter syntax element, depending on one or more fractional parts (or fractional parts) of a motion vector (for example, the first filter signaled by the filter syntax element is selected for filtering in the first direction, for example, when considering position coordinates in the first direction, for example, when one or more fractional parts (or fractional parts) of the motion vector are between full samples, and the second filter signaled by the filter syntax element is selected for filtering in a second direction different from the first direction, for example, when considering position coordinates in the second direction, for example, when all fractional parts (or fractional parts) of the motion vector are on full samples).
[0060] According to one embodiment, the decoder is configured to take over filter selection from adjacent blocks in merge mode (for example, selectively depending on whether the motion vector is identical to the motion vector of the adjacent block).
[0061] According to one embodiment, in the decoder, the selected (e.g., applied) processing method defines (e.g., spatial) filtering of the lumens and / or (e.g., spatial) filtering of the chromens. According to one embodiment, the decoder is configured to select (e.g., apply) a filter depending on the prediction mode.
[0062] According to one embodiment, the decoder is configured to adapt mapping rules that map encoded index values to instructions (e.g., index values) indicating a processing method in accordance with one or more syntax elements.
[0063] According to one embodiment, the decoder is configured to adapt the number of bins used to identify an encoded index indicating a selected processing method, depending on one or more syntax elements.
[0064] According to one embodiment, the decoder is configured to determine a set of acceptable processing schemes in terms of one or more syntax elements identified based on an index.
[0065] According to one embodiment, the decoder is configured to determine whether a single processing method is acceptable for a portion of a video sequence, and to selectively omit the identification of an encoded index if it is found that only a single processing method (e.g., a spatial filter or an interpolation filter) is acceptable for the portion of the video sequence.
[0066] According to one embodiment, the decoder is configured to select a mapping rule that maps encoded index values to instructions (e.g., index values) indicating a processing method, such that the number of bins representing encoded index values matches the number of processing methods acceptable for a portion of the video sequence.
[0067] According to one embodiment, the decoder is configured to select (e.g., apply) different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or "picture").
[0068] According to one embodiment, the decoder is configured to decode an index indicating a selected processing scheme (e.g., depending on one or more syntax elements) such that a given encoded index value represents (e.g., associated with) a different processing scheme depending on the motion vector precision (which can be signaled, for example, using syntax elements in a bitstream, and can be, for example, one of the following values: QPEL, HPEL, FPEL, 4PEL).
[0069] According to one embodiment, one or more syntax elements include at least one of the following: motion vector accuracy (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translation inter, affine inter, translation merge, affine merge, combined inter / intra), availability of coded residual signal (e.g., coded block flag), spectral characteristics of coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edges from co-located prediction units PU within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, deblocking filter determination and boundary intensity), motion vector length (e.g., enable only long motion vectors or additional smoothing filters for specific directions), or adaptive motion vector resolution mode.
[0070] According to one embodiment, the decoder is configured to decode an index indicating a selected processing scheme (for example, depending on one or more syntax elements) such that a given encoded index value represents (for example, associated with) a different processing scheme depending on the fractional part (or multiple fractional parts) of the motion vector (for example, the different encoding schemes used depend on whether all motion vectors of the part under consideration point to integer sample positions, e.g., lumar sample positions).
[0071] According to one embodiment, the decoder is configured to selectively determine a set of processing schemes (e.g., "encoding of multiple possibilities") for motion vector precision (e.g., subsample motion vector precision, e.g., HPEL) between maximum motion vector precision (e.g., subsample motion vector resolution finer than the aforementioned motion vector precision, e.g., QPEL) and minimum motion vector precision (e.g., 4PEL), or for motion vector precision (e.g., HPEL) between maximum motion vector precision (e.g., QPEL) and full sample motion vector precision (e.g., FPEL) that identifies the processing scheme to be applied.
[0072] According to one embodiment, the decoder is configured to decode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering), a second FIR filtering, and a third FIR.
[0073] According to one embodiment, the decoder is configured to decode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering) and a second FIR filtering.
[0074] According to one embodiment, the decoder is configured to apply a processing method (e.g., an interpolation filter) having different characteristics (e.g., a stronger low-pass characteristic versus a weaker low-pass characteristic).
[0075] According to one embodiment, the decoder is configured to select a mapping rule depending on the motion vector accuracy and the fractional sample position or fractional part (or multiple fractional parts) of the motion vector. According to one embodiment, the decoder is configured to determine a set of available processing methods (e.g., interpolation filters) depending on the motion vector accuracy.
[0076] According to one embodiment, the decoder is configured to apply one of the following motion vector resolutions: quarter-sample, half-sample, full-sample, and four-sample (for example, an index element or interpolation filter describing the processing method is selectively included in the bitstream when half-sample motion vector resolution is selected, and the index element describing the processing method is omitted in other cases, for example). According to one embodiment, the decoder includes one or more FIR filters in its processing method.
[0077] Another aspect of the present disclosure relates to a decoder for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction error), wherein the decoder provides output video content based on an encoded representation of a video sequence, the decoder takes a bitstream as the encoded representation of the video sequence (e.g., receive, download, open from file), identifies one or more syntax elements from the bitstream relating to a portion of the video sequence, and, depending on one or more syntax elements (e.g., motion vector precision syntax elements), provides an encoded index to the bitstream indicating a processing method (e.g., interpolation filter, e.g., selected by the encoder) for the portion of the video sequence (e.g., block), the entropy used in the bitstream. The encoding scheme is determined, and the processing scheme is configured to take samples for motion compensation prediction (e.g., samples, e.g., fractional samples, relating to the video sequence) at integer and / or fractional positions within a portion of the video sequence, and if the determined entropy encoding scheme specifies a number of bins greater than zero, the encoded index in the bitstream is identified, the index is decoded, and the processing scheme is identified (e.g., selected) based on one or more syntax elements, and if the number of bins specified by the determined encoding scheme is greater than zero, based on the decoded index (or, only if the number of determined bins is zero, the processing scheme is identified based on one or more syntax elements), and the identified processing scheme is configured to be applied to the portion of the video sequence.
[0078] The decoder is based on the idea of identifying the processing method for a part by determining the entropy coding scheme based on the syntax elements. From the determined entropy coding scheme, the decoder can infer how to identify the processing method. If the number of bins defined by the entropy coding scheme is greater than zero, the decoder can infer the information encoded in the bitstream by the entropy coding scheme regardless of the number of bins. In other words, the decoder can use the information that the number of bins is 0 to identify the processing method. Furthermore, since the decoder can use the syntax elements and indices to identify the processing method, the decoder can identify the processing method even if a given encoded index value can be associated with various processing methods. Thus, the decoder can identify the processing method even if the number of bins in the bitstream for transmitting information about the processing method is small or even zero. According to one embodiment, in a decoder, the entropy coding scheme includes a binarization scheme used to provide an encoded index. According to one embodiment, in a decoder, the binarization scheme includes the number of bins used to provide an encoded index.
[0079] According to one embodiment, in the decoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0080] According to one embodiment, in the decoder, the filters in the set of filters are separably applicable to parts of the video sequence (for example, one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or the y-direction).
[0081] According to one embodiment, the decoder is configured to select (e.g., apply) a filter (e.g., an interpolation filter) in accordance with a fractional part (or multiple fractional parts) of a motion vector (e.g., in accordance with one or more fractional parts of a motion vector) (e.g., by selecting a first filter if the endpoint of the motion vector is at a full sample position or at a full sample position in the direction under consideration, and by selecting a second filter different from the first filter if the endpoint of the motion vector is between full samples or between full samples in the direction under consideration).
[0082] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction (e.g., x) and for filtering in a second direction (e.g., y) depending on the fractional part (or multiple fractional parts) of the motion vector.
[0083] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction and for filtering in a second direction, such that stronger low-pass filtering is applied to filtering in directions where one, more, or all fractional parts (or multiple fractional parts) of the motion vector are on the full sample (e.g., when the exact motion vector(s) of the half-sample point to the full sample position(s)) (compared to filtering applied in directions where the fractional parts(s) of the motion vector are between full samples).
[0084] According to one embodiment, the decoder is configured to select a filter (e.g., an interpolation filter) signaled by a syntax element (e.g., an interpolation filter index if_idx) or a predetermined filter, depending on one or more fractional parts (or more fractional parts) of a motion vector (e.g., to apply, e.g., individually for each filtering direction) (for example, the filter signaled by the syntax element is selected for filtering in the first direction, for example, when considering the position coordinates of a first direction, e.g., when one or more fractional parts (or more fractional parts) of the motion vector are between full samples, and the predetermined filter is selected for filtering in a second direction different from the first direction, for example, when considering the position coordinates of a second direction, e.g., when all fractional parts (or more fractional parts) of the motion vector are on full samples).
[0085] According to one embodiment, the decoder is configured to select (for example, individually for each filtering direction) a first filter (e.g., interpolation filter) signaled by a first syntax element (e.g., interpolation filter index if_idx) or a second filter signaled by a filter syntax element, depending on one or more fractional parts (or fractional parts) of a motion vector (for example, the first filter signaled by the filter syntax element is selected for filtering in the first direction, for example, when considering position coordinates in the first direction, for example, when one or more fractional parts (or fractional parts) of the motion vector are between full samples, and the second filter signaled by the filter syntax element is selected for filtering in a second direction different from the first direction, for example, when considering position coordinates in the second direction, for example, when all fractional parts (or fractional parts) of the motion vector are on full samples).
[0086] According to one embodiment, the decoder is configured to take over filter selection from adjacent blocks in merge mode (for example, selectively depending on whether the motion vector is identical to the motion vector of the adjacent block).
[0087] According to one embodiment, in the decoder, the selected (e.g., applied) processing method defines (e.g., spatial) filtering of the lumens and / or (e.g., spatial) filtering of the chromens. According to one embodiment, the decoder is configured to select (e.g., apply) a filter depending on the prediction mode.
[0088] According to one embodiment, the decoder is configured to adapt mapping rules that map encoded index values to instructions (e.g., index values) indicating a processing method in accordance with one or more syntax elements.
[0089] According to one embodiment, the decoder is configured to adapt the number of bins used to identify an encoded index indicating a selected processing method, depending on one or more syntax elements.
[0090] According to one embodiment, the decoder is configured to determine a set of acceptable processing schemes in terms of one or more syntax elements identified based on an index.
[0091] According to one embodiment, the decoder is configured to determine whether a single processing method is acceptable for a portion of a video sequence, and to selectively omit the identification of an encoded index if it is found that only a single processing method (e.g., a spatial filter or an interpolation filter) is acceptable for the portion of the video sequence.
[0092] According to one embodiment, the decoder is configured to select a mapping rule that maps encoded index values to instructions (e.g., index values) indicating a processing method, such that the number of bins representing encoded index values matches the number of processing methods acceptable for a portion of the video sequence.
[0093] According to one embodiment, the decoder is configured to select (e.g., apply) different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or "picture").
[0094] According to one embodiment, the decoder is configured to decode an index indicating a selected processing scheme (e.g., depending on one or more syntax elements) such that a given encoded index value represents (e.g., associated with) a different processing scheme depending on the motion vector precision (which can be signaled, for example, using syntax elements in a bitstream, and can be, for example, one of the following values: QPEL, HPEL, FPEL, 4PEL).
[0095] According to one embodiment, the decoder includes at least one of the following syntax elements: motion vector accuracy (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translation inter, affine inter, translation merge, affine merge, combined inter / intra), availability of coded residual signal (e.g., coded block flag), spectral characteristics of coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edges from co-located prediction units PU within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, deblocking filter determination and boundary intensity), motion vector length (e.g., enable only long motion vectors or additional smoothing filters for specific directions), or adaptive motion vector resolution mode.
[0096] According to one embodiment, the decoder is configured to decode an index indicating a selected processing scheme (for example, depending on one or more syntax elements) such that a given encoded index value represents (for example, associated with) a different processing scheme depending on the fractional part (or multiple fractional parts) of the motion vector (for example, the different encoding schemes used depend on whether all motion vectors of the part under consideration point to integer sample positions, e.g., lumar sample positions).
[0097] According to one embodiment, the decoder is configured to selectively determine a set of processing schemes (e.g., "encoding of multiple possibilities") for motion vector precision (e.g., subsample motion vector precision, e.g., HPEL) between maximum motion vector precision (e.g., subsample motion vector resolution finer than the aforementioned motion vector precision, e.g., QPEL) and minimum motion vector precision (e.g., 4PEL), or for motion vector precision (e.g., HPEL) between maximum motion vector precision (e.g., QPEL) and full sample motion vector precision (e.g., FPEL) that identifies the processing scheme to be applied.
[0098] According to one embodiment, the decoder is configured to decode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering), a second FIR filtering, and a third FIR.
[0099] According to one embodiment, the decoder is configured to decode an index (processing method index, if_idx) for selecting between a first FIR filtering (e.g., HEVC filtering) (e.g., HEVC 8-tap filtering) and a second FIR filtering.
[0100] According to one embodiment, the decoder is configured to apply a processing method (e.g., an interpolation filter) having different characteristics (e.g., a stronger low-pass characteristic versus a weaker low-pass characteristic).
[0101] According to one embodiment, the decoder is configured to select a mapping rule depending on the motion vector accuracy and the fractional sample position or fractional part (or multiple fractional parts) of the motion vector. According to one embodiment, the decoder is configured to determine a set of available processing methods (e.g., interpolation filters) depending on the motion vector accuracy.
[0102] According to one embodiment, the decoder is configured to apply one of the following motion vector resolutions: quarter-sample, half-sample, full-sample, and four-sample (for example, an index element or interpolation filter describing the processing method is selectively included in the bitstream when half-sample motion vector resolution is selected, and the index element describing the processing method is omitted in other cases, for example). According to one embodiment, the processing method comprises one or more FIR filters.
[0103] Another aspect of the present disclosure relates to a method for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the method provides an encoded representation of a video sequence based on input video content (e.g., a method performed by an encoder), the method determines one or more syntax elements (e.g., motion vector precision syntax elements) associated with a portion of the video sequence (e.g., a block), and a processing scheme (e.g., applied by a decoder) to the portion of the video sequence (e.g., The process involves selecting an interpolation filter, obtaining samples (e.g., samples, e.g., fractional samples) for motion compensation prediction at integer and / or fractional positions within a portion of the video sequence, encoding an index indicating the selected process (dependent on one or more syntax elements) such that a given encoded index value represents a different process (e.g., related) depending on the characteristics described by one or more syntax elements, and providing (e.g., sending, uploading, saving as a file) a bitstream containing one or more syntax elements and the encoded index as an encoded representation of the video sequence.
[0104] Another aspect of the present disclosure relates to a method for hybrid video coding (e.g., video coding with prediction and transformation coding of prediction error), wherein the method provides an encoded representation of a video sequence based on input video content (e.g., a method performed by an encoder), the method determines one or more syntax elements (e.g., motion vector precision syntax elements) associated with a portion of the video sequence (e.g., a block), and selects a processing method (e.g., interpolation filter) to be applied to the portion of the video sequence (e.g., by a decoder) based on a characteristic described by one or more syntax elements (e.g., decoder settings signaled by one or more syntax elements, or characteristics of a portion of the video sequence signaled by one or more syntax elements), wherein the processing method applies integers and / or minutes within the portion of the video sequence. This includes obtaining samples for motion compensation prediction at several positions (e.g., samples, e.g., fractional samples, related to a video sequence), selecting an entropy coding scheme used to provide an encoded index indicating a selected processing scheme depending on a characteristic described by one or more syntax elements (e.g., depending on one or more syntax elements) (e.g., the number of bins described by the entropy coding scheme may be equal to zero if, in terms of one or more syntax elements considered, only one processing scheme is acceptable), and providing a bitstream as an encoded representation of the video sequence that includes one or more syntax elements and an encoded index (e.g., if the selected number of bins is greater than zero) (e.g., send, upload, save as file).
[0105] Another aspect of the present disclosure relates to a method for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction error), wherein the method provides output video content based on an encoded representation of a video sequence (e.g., a method performed by a decoder), the method includes obtaining a bitstream as an encoded representation of the video sequence (e.g., receiving, downloading, opening from a file), identifying from the bitstream one or more syntax elements (e.g., motion vector precision syntax elements) related to a portion of the video sequence (e.g., a block) (and an encoded index), and one or more syntax elements The method includes identifying a processing scheme (e.g., an interpolation filter selected by an encoder) for a portion of a video sequence based on a subscript and an index (which may be contained in a bitstream in an encoded form), the processing scheme being for obtaining samples for motion compensation predictions (e.g., samples, e.g., fractional samples, relating to the video sequence), and assigning different processing schemes to a given encoded index value depending on one or more syntax elements (e.g., depending on the value of one or more syntax elements), and applying the identified processing scheme to the portion of the video sequence.
[0106] Another aspect of the present disclosure relates to a method for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction error), wherein the method provides output video content based on an encoded representation of a video sequence (e.g., a method performed by a decoder), the method includes obtaining a bitstream as an encoded representation of the video sequence (e.g., receiving, downloading, opening from a file), identifying from the bitstream one or more syntax elements (e.g., motion vector precision syntax elements) related to a portion of the video sequence (e.g., a block), and, depending on the one or more syntax elements, using the bitstream to provide an encoded index indicating a processing scheme (e.g., interpolation filter, e.g., selected by an encoder) for the portion of the video sequence. The process involves determining an entropy encoding scheme, the process being for obtaining samples for motion compensation predictions (e.g., samples, e.g., fractional samples, relating to the video sequence) at integer and / or fractional positions within a portion of the video sequence, and if the determined entropy encoding scheme specifies a number of bins greater than zero, identifying the encoded index in the bitstream; decoding the index and identifying (e.g., selecting) a process based on one or more syntax elements, and if the number of bins specified by the determined entropy encoding scheme is greater than zero, based on the decoded index (or identifying a process based on one or more syntax elements only if the number of bins determined is zero); and applying the identified process to the portion of the video sequence.
[0107] Another aspect of the present disclosure relates to a video bitstream for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the video bitstream describes a video sequence (e.g., in the form of an encoded representation of the video sequence), the video bitstream includes one or more syntax elements (e.g., motion vector precision syntax elements) relating to a portion of the video sequence (e.g., a block), the video bitstream includes a processing scheme index bitstream element (e.g., if_idx) describing an interpolation filter used by a video decoder, the presence of the processing scheme index bitstream element (e.g., describing a filtering or interpolation filter performed by a video decoder) varies depending on one or more other syntax elements (e.g., the processing scheme index bitstream element is present for a first value of another syntax element or for a first combination of values of several other syntax elements, for example, the processing scheme index bitstream element is omitted for a second value of another syntax element or for a second combination of values of several other syntax elements).
[0108] Another aspect of the present disclosure relates to a video bitstream for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the video bitstream describes a video sequence (e.g., in the form of an encoded representation of the video sequence), the video bitstream includes one or more syntax elements (e.g., motion vector precision syntax elements) relating to a portion of the video sequence (e.g., a block), the video bitstream includes a processing scheme index bitstream element (e.g., describing an interpolation filter used by a video decoder, e.g., if_idx), the number of bins in the processing scheme index bitstream element varies depending on one or more syntax elements relating to a portion of the video sequence (e.g., the number of bins in the processing scheme index bitstream element is such that it conforms to the number of acceptable processing schemes for a portion of the video sequence, e.g., the number of bins in the processing scheme index bitstream element is equal to zero, e.g., the processing scheme index bitstream element is omitted if only one processing scheme is acceptable for a portion of the video sequence).
[0109] Another aspect of the present disclosure relates to a video bitstream for hybrid video coding (e.g., video coding with prediction and conversion coding of prediction errors), wherein the video bitstream describes a video sequence (e.g., in the form of an encoded representation of the video sequence), the video bitstream includes one or more syntax elements (e.g., motion vector precision syntax elements) relating to a portion of the video sequence (e.g., a block), the video bitstream includes a processing scheme index bitstream element (e.g., if_idx) describing an interpolation filter used by a video decoder, the meaning of the codeword (e.g., encoded index value) of the processing scheme index bitstream element varies depending on one or more syntax elements relating to a portion of the video sequence (e.g., different processing schemes are assigned to a given encoded index value depending on one or more syntax elements relating to a portion of the video sequence).
[0110] Another aspect of the present disclosure relates to a hybrid video coding decoder, the decoder provides output video content based on an encoded representation of a video sequence, the decoder takes a bitstream as the encoded representation of the video sequence, identifies one or more syntax elements from the bitstream relating to a portion of the video sequence, selects a processing scheme for the portion of the video sequence according to the one or more syntax elements, the processing scheme is entirely specified by the characteristics of the portion and fractional parts of the motion vector (the characteristics are described by one or more syntax elements), the processing scheme is configured to take samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and applies the identified processing scheme to the portion of the video sequence.
[0111] Since a decoder can select a processing method based on one or more syntax elements, it can select a processing method for a portion without requiring a syntax element specifically dedicated to the processing method. Therefore, the decoder can select a processing method even if the processing method is not explicitly indicated in the bitstream, resulting in a particularly small bitstream. When the decoder uses one or more syntax elements to obtain a processing method, it can select a processing method at the granularity provided by the one or more syntax elements. Thus, the decoder can efficiently use information already provided in the bitstream to select a processing method at a fine granularity without requiring additional bits in the bitstream. A fractional part of the motion vector can characterize a processing method used to obtain a sample particularly accurately, since the motion vector may be used for motion compensation prediction. Therefore, a processing method can be selected quickly and reliably based on a fractional part of the motion vector.
[0112] According to one embodiment, the characteristics of the part specifying the processing method relate to one, two, or all of the following: motion vector accuracy, fractional position of the sample (acquired by the processing method), and prediction mode used for the part. Since the processing method may be for acquiring samples for motion-compensated prediction, the MV accuracy can provide a precise hint about the appropriate processing method to use for the part. Since different processing methods can be used for full-sample, half-sample, and quarter-sample positions, the fractional position of the sample can limit the number of acceptable processing methods or even specify a processing method. One, several, or all of the motion vector accuracy, fractional position of the sample, and prediction mode may be indicated by one or more syntax elements for decoding a portion of the video sequence. Thus, the decoder main infers the processing method from the already available information.
[0113] According to one embodiment, the part is a prediction unit (PU) or a coding unit (CU), and the decoder is configured to individually select a processing method for different parts of the video sequence. By individually selecting a processing method for different prediction units or coding units, it becomes possible to adapt the processing method to the local characteristics of the video sequence, particularly with fine granularity. This precise adaptation can enable decoding of parts, particularly from small-sized bitstreams, and / or particularly fast decoding. Since the processing method is selected depending on one or more syntax elements, it is possible to determine the processing method individually for different prediction units or coding units without increasing the size of the bitstream. According to one embodiment, the decoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the decoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0114] Another aspect of the present disclosure relates to a hybrid video coding decoder, the decoder provides output video content based on an encoded representation of a video sequence, the decoder takes a bitstream as the encoded representation of the video sequence, identifies one or more syntax elements from the bitstream relating to a portion of the video sequence, selects a processing scheme for the portion of the video sequence depending on the motion vector accuracy, the processing scheme takes samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and is configured to apply the identified processing scheme to the portion of the video sequence.
[0115] According to one embodiment, the decoder is configured to select a processing method depending on the motion vector accuracy and the fractional position of the sample (obtained by the processing method). According to one embodiment, the decoder is configured to further select a processing method depending on the prediction mode used for the part.
[0116] According to one embodiment, the part is a prediction unit (PU) or a coding unit (CU), and the decoder is configured to individually select a processing method for different parts of the video sequence. By individually selecting a processing method for different prediction units or coding units, it becomes possible to adapt the processing method to the local characteristics of the video sequence, particularly with fine granularity. This precise adaptation can enable decoding of parts, particularly from small-sized bitstreams, and / or particularly fast decoding. Since the processing method is selected depending on one or more syntax elements, it is possible to determine the processing method individually for different prediction units or coding units without increasing the size of the bitstream. According to one embodiment, the decoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the decoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0117] Another aspect of the present disclosure relates to a hybrid video coding decoder, the decoder provides output video content based on an encoded representation of a video sequence, the decoder takes a bitstream as the encoded representation of the video sequence, identifies one or more syntax elements from the bitstream related to a portion of the video sequence, determines a processing scheme for the portion of the video sequence depending on the one or more syntax elements using PU granularity or CU granularity, the processing scheme is configured to take samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and applies the identified processing scheme to the portion of the video sequence. According to one embodiment, the decoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the decoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0118] According to one embodiment, in the decoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0119] According to one embodiment, in the decoder, the filters in the set of filters are separably applicable to parts of the video sequence (for example, one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or the y-direction).
[0120] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction (e.g., x) and for filtering in a second direction (e.g., y) depending on the fractional part (or multiple fractional parts) of the motion vector.
[0121] According to one embodiment, the decoder is configured to select (e.g., apply) different filters (e.g., interpolation filters) for filtering in a first direction and for filtering in a second direction, such that stronger low-pass filtering is applied to filtering in directions where one, more, or all fractional parts (or multiple fractional parts) of the motion vector are on the full sample (e.g., when the exact motion vector(s) of the half-sample point to the full sample position(s)) (compared to filtering applied in directions where the fractional parts(s) of the motion vector are between full samples).
[0122] According to one embodiment, the decoder is configured to take over filter selection from adjacent blocks in merge mode (for example, selectively depending on whether the motion vector is identical to the motion vector of the adjacent block).
[0123] According to one embodiment, in the decoder, the selected (e.g., applied) processing method defines (e.g., spatial) filtering of the lumens and / or (e.g., spatial) filtering of the chromens.
[0124] According to one embodiment, the decoder is configured to select (e.g., apply) different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or "picture").
[0125] According to one embodiment, one or more syntax elements include at least one of the following: motion vector accuracy (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translation inter, affine inter, translation merge, affine merge, combined inter / intra), availability of coded residual signal (e.g., coded block flag), spectral characteristics of coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edges from co-located prediction units PU within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, deblocking filter determination and boundary intensity), motion vector length (e.g., enable only long motion vectors or additional smoothing filters for specific directions), or adaptive motion vector resolution mode.
[0126] According to one embodiment, the decoder is configured to selectively determine a set of processing schemes (e.g., "encoding of multiple possibilities") for motion vector precision (e.g., subsample motion vector precision, e.g., HPEL) between maximum motion vector precision (e.g., subsample motion vector resolution finer than the aforementioned motion vector precision, e.g., QPEL) and minimum motion vector precision (e.g., 4PEL), or for motion vector precision (e.g., HPEL) between maximum motion vector precision (e.g., QPEL) and full sample motion vector precision (e.g., FPEL) that identifies the processing scheme to be applied.
[0127] According to one embodiment, the decoder is configured to apply a processing method (e.g., an interpolation filter) having different characteristics (e.g., a stronger low-pass characteristic versus a weaker low-pass characteristic).
[0128] According to one embodiment, the decoder is configured to apply one of the following: a quarter-sample motion vector resolution, a half-sample motion vector resolution, a full-sample motion vector resolution, or a four-sample motion vector resolution. According to one embodiment, the decoder includes one or more FIR filters in its processing method.
[0129] Another aspect of the present disclosure relates to a hybrid video coding encoder, wherein the encoder provides an encoded representation of a video sequence based on input video content, the encoder determines one or more syntax elements relating to a portion of the video sequence, selects a processing scheme to be applied to the portion of the video sequence based on the characteristics described by the one or more syntax elements, the processing scheme is entirely specified by the characteristics of the portion and fractional parts of the motion vector, the processing scheme is for taking samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and is configured to provide a bitstream containing one or more syntax elements as the encoded representation of the video sequence.
[0130] According to one embodiment, the characteristics described by one or more syntax elements relate to motion vector accuracy, fractional position of a sample (obtained by the processing scheme), and one, two, or all of the prediction modes used for that portion.
[0131] According to one embodiment, the component is a prediction unit (PU) or a coding unit (CU), and the encoder is configured to individually select a processing method for different parts of a video sequence. According to one embodiment, the encoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the encoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0132] Another aspect of the present disclosure relates to a hybrid video coding encoder, wherein the encoder provides an encoded representation of a video sequence based on input video content, the encoder determines one or more syntax elements related to a portion of the video sequence, selects a processing scheme to be applied to the portion of the video sequence based on the motion vector precision described by the one or more syntax elements, the processing scheme is fully specified by the characteristics of the portion and a fractional part of the motion vector, and is configured to provide a bitstream containing one or more syntax elements as an encoded representation of the video sequence.
[0133] According to one embodiment, the encoder is configured to select a processing method based on motion vector accuracy and the fractional position of the sample (obtained by the processing method). According to one embodiment, the encoder is configured to further select a processing method depending on the prediction mode used for a portion.
[0134] According to one embodiment, the component is a prediction unit (PU) or a coding unit (CU), and the encoder is configured to individually select a processing method for different parts of a video sequence. According to one embodiment, the encoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the encoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0135] Another aspect of the present disclosure relates to a hybrid video coding encoder, the encoder providing an encoded representation of a video sequence based on input video content, the encoder determining one or more syntax elements related to a portion of the video sequence, selecting a processing scheme to apply to the portion of the video sequence based on the motion vector precision described by the one or more syntax elements using PU granularity or CU granularity, the processing scheme being fully specified by the characteristics of the portion and a fractional part of the motion vector, and the encoder providing a bitstream containing the one or more syntax elements as an encoded representation of the video sequence. According to one embodiment, the encoder is configured to use different processing methods for the same fractional sample position in different parts of a video sequence. According to one embodiment, the encoder is configured to switch processing methods at the PU or CU level to adapt to local image characteristics. According to one embodiment, none of the one or more syntax elements are explicitly dedicated to indicating a processing method.
[0136] According to one embodiment, the encoder is configured to determine a set of acceptable processing methods in terms of characteristics described by one or more syntax elements, and to select a processing method from the determined set.
[0137] According to one embodiment, in an encoder, the processing method is a filter (e.g., an interpolation filter) or a set of filters (e.g., an interpolation filter) (e.g., parameterized by the fractional part of the motion vector).
[0138] According to one embodiment, in the encoder, the filters in the set of filters are separably applicable to parts of the video sequence (for example, one-dimensional filters that can be sequentially applied in the spatial direction, e.g., the x-direction and / or the y-direction).
[0139] According to one embodiment, the encoder is configured to select different processing methods (e.g., different interpolation filters) for different blocks within a video frame or picture (e.g., different prediction units PU and / or coding units CU and / or coding tree units CTU) (different interpolation filters for the same fractional sample position in different parts of a video sequence or different parts of a single video frame or picture).
[0140] According to one embodiment, the encoder includes at least one of the following syntax elements: motion vector accuracy (e.g., quarter sample, half sample, full sample), fractional sample position, block size, block shape, number of prediction hypotheses (e.g., one hypothesis for single prediction, two hypotheses for dual prediction), prediction mode (e.g., translation inter, affine inter, translation merge, affine merge, combined inter / intra), availability of coded residual signal (e.g., coded block flag), spectral characteristics of coded residual signal, reference picture signal (e.g., reference block defined by motion data, block edges from co-located prediction unit PU within the reference block, high or low frequency characteristics of the reference signal), loop filter data (e.g., edge offset or band offset classification from sample adaptive offset filter SAO, deblocking filter determination and boundary intensity), motion vector length (e.g., enable only long motion vectors or additional smoothing filters for specific directions), or adaptive motion vector resolution mode.
[0141] According to one embodiment, the encoder is configured to selectively determine a set of processing methods (e.g., "encoding of multiple possibilities") for motion vector precision between maximum motion vector precision (e.g., subsample motion vector resolution finer than the aforementioned motion vector precision, e.g., QPEL) and minimum motion vector precision (e.g., 4PEL) (e.g., subsample motion vector precision, e.g., HPEL), or for motion vector precision between maximum motion vector precision (e.g., QPEL) and full sample motion vector precision (e.g., FPEL) (e.g., HPEL), and to select a processing method from the determined set.
[0142] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing output video content based on an encoded representation of a video sequence, the method comprising: obtaining a bitstream as an encoded representation of a video sequence; identifying one or more syntax elements from the bitstream relating to a portion of the video sequence; selecting a processing scheme for the portion of the video sequence in accordance with the one or more syntax elements, the processing scheme being entirely specified by the characteristics of the portion and fractional parts of the motion vector, the processing scheme being for obtaining samples for motion compensation predictions at integer and / or fractional positions within the portion of the video sequence, and applying the identified processing scheme to the portion of the video sequence.
[0143] According to one embodiment, the encoder is configured to select between a quarter-sample motion vector resolution, a half-sample motion vector resolution, a full-sample motion vector resolution, and a four-sample motion vector resolution.
[0144] According to one embodiment, the encoder comprises one or more FIR filters. The FIR filters are particularly stable filters for processing video sequences.
[0145] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing output video content based on an encoded representation of a video sequence, the method comprising: obtaining a bitstream as an encoded representation of a video sequence; identifying one or more syntax elements from the bitstream relating to a portion of the video sequence; selecting a processing scheme for the portion of the video sequence, depending on the motion vector precision, the processing scheme for obtaining samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence; and applying the identified processing scheme to the portion of the video sequence.
[0146] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing output video content based on an encoded representation of a video sequence, the method comprising: obtaining a bitstream as an encoded representation of a video sequence; identifying one or more syntax elements from the bitstream relating to a portion of the video sequence; determining a processing scheme for the portion of the video sequence in accordance with the one or more syntax elements, using PU granularity or CU granularity, the processing scheme for obtaining samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence; and applying the identified processing scheme to the portion of the video sequence.
[0147] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing an encoded representation of a video sequence based on input video content, the method comprising determining one or more syntax elements relating to a portion of a video sequence, selecting a processing scheme to be applied to the portion of the video sequence based on the characteristics described by the one or more syntax elements, the processing scheme being entirely specified by the characteristics of the portion and fractional parts of a motion vector, the processing scheme being for taking samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and providing a bitstream containing one or more syntax elements as an encoded representation of the video sequence.
[0148] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing an encoded representation of a video sequence based on input video content, the method comprising determining one or more syntax elements related to a portion of the video sequence, selecting a processing scheme to be applied to the portion of the video sequence based on motion vector precision described by the one or more syntax elements, the processing scheme for taking samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence, and providing a bitstream containing one or more syntax elements as an encoded representation of the video sequence.
[0149] Another aspect of the present disclosure relates to a hybrid video coding method, the method for providing an encoded representation of a video sequence based on input video content, the method comprising: determining one or more syntax elements related to a portion of the video sequence; selecting a processing scheme to be applied to the portion of the video sequence based on a property described by the one or more syntax elements, using PU granularity or CU granularity, the processing scheme for taking samples for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence; and providing a bitstream containing one or more syntax elements as an encoded representation of the video sequence.
[0150] Another aspect of this disclosure relates to a computer program for performing one of the described methods for hybrid video coding when the computer program is running on a computer.
[0151] The decoders described herein offer similar functionality and advantages to those described for encoders, and vice versa. The functionality and advantages of encoders referencing encoder features apply equally to the corresponding features of decoders, with decoding being replaced by encoding, and vice versa. For example, if encoder features enable encoding or efficient encoding of information, particularly in small bitstreams, the corresponding features of decoders may enable the deriving of information from small bitstreams or efficient decoding, respectively. In other words, any of the features, functions, and details disclosed herein with respect to encoders may be optionally included in decoders, individually or in combination, and vice versa.
[0152] The methods described for hybrid video coding rely on the same concepts as the encoders and decoders described above and provide equivalent or equivalent functionality and benefits. The methods may optionally be combined with (or supplemented with) any of the features, functions, and details described herein with respect to the corresponding encoders and decoders. The methods for hybrid video coding can optionally be combined with the mentioned features, functions, and details individually or in any combination thereof. [Brief explanation of the drawing]
[0153] Embodiments of the present disclosure are described herein with reference to the accompanying drawings and figures. [Figure 1] This is a schematic diagram of an encoder that can implement the disclosed concept. [Figure 2] This is a schematic diagram of a decoder that can implement the disclosed concept. [Figure 3] This figure shows the signals used by an encoder or decoder according to one embodiment. [Figure 4] This figure shows fractional sample positions that can be used for motion compensation prediction according to one embodiment. [Figure 5] This figure shows an encoder according to one embodiment. [Figure 6] This figure shows an encoder according to another embodiment. [Figure 7] This figure shows a decoder according to one embodiment. [Figure 8] This figure shows a decoder according to another embodiment. [Figure 9] This is a flowchart of a hybrid video coding method according to one embodiment. [Figure 10] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 11] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 12]A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 13] This figure shows an encoder according to another embodiment. [Figure 14] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 15] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 16] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 17] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 18] A flowchart of a hybrid video coding method according to another embodiment is shown. [Figure 19] A flowchart of a hybrid video coding method according to another embodiment is shown. [Modes for carrying out the invention]
[0154] Different embodiments and aspects of the invention will be described below. Further embodiments are defined by the appended claims.
[0155] It should be noted that any embodiment defined by the claims may be supplemented by any of the details (features and functions) described herein. Furthermore, the embodiments described herein may be used individually and may be supplemented by any of the features included in the claims.
[0156] Furthermore, it should be noted that the individual embodiments described herein can be used individually or in combination. Therefore, details can be added to each individual embodiment without adding details to another embodiment.
[0157] It should also be noted that this disclosure explicitly or implicitly describes features available in video encoders (devices for providing an encoded representation of an input video signal) and video decoders (devices for providing a decoded representation of a video signal based on an encoded representation). Therefore, any of the features described herein may be used in the context of video encoders and video decoders.
[0158] Furthermore, the features and functions disclosed herein in relation to the methods may also be used in apparatus (configured to perform such functions). Additionally, any features and functions disclosed herein in relation to an apparatus may also be used in a corresponding manner. In other words, the methods disclosed herein may be complemented by any of the features and functions described in relation to an apparatus.
[0159] Furthermore, any of the features and functions described herein can be implemented in hardware, software, or a combination of hardware and software, as described in other sections such as "Alternative Implementations."
[0160] Furthermore, any of the features and syntax elements described herein can be optionally introduced into the video bitstream, individually or in combination.
[0161] The following description details embodiments, but it should be understood that these embodiments provide many applicable concepts that can be incorporated into a wide variety of video encoders and video decoders. The specific embodiments described are merely examples of specific ways of implementing and using the concepts and do not limit the scope of the embodiments. In the following description of embodiments, the same or similar elements or elements having the same function are given the same reference numeral, or identified by the same name, and repeated descriptions of elements that are given the same reference number, or identified by the same name are generally omitted. Thus, the descriptions provided for elements that have the same or similar reference number, or are identified by the same name, are interchangeable or applicable to each other in different embodiments. The following description provides several details to provide a more complete description of the embodiments of the disclosure. However, it will be apparent to those skilled in the art that other embodiments can be carried out without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail, so as not to obscure the examples described herein. Also, features of different embodiments described herein can be combined with each other unless otherwise noted.
[0162] 1. Examples of coding frameworks The following description of the drawings begins with a presentation of a description of a block-based predictive codec encoder and decoder for encoding video pictures, in order to form an example of a coding framework in which embodiments of the present invention can be incorporated. Each encoder and decoder is described with respect to Figures 1 to 3. Hereafter, embodiments of the concept of the present invention are presented, along with a description of how such concepts can be incorporated into the encoders and decoders of Figures 1 and 2, respectively. However, embodiments described in subsequent Figures 4 and beyond may also be used to form encoders and decoders that do not operate according to the underlying coding framework of the encoders and decoders of Figures 1 and 2.
[0163] Figure 1 illustrates an apparatus for predictively encoding picture 12 into a data stream 14 using transform-based residual coding as an example. The apparatus or encoder is indicated by reference numeral 10. Figure 2 illustrates a corresponding decoder 20, i.e., an apparatus 20 configured to predictively decode picture 12' from the data stream 14 using transform-based residual decoding, where an apostrophe is used to indicate that picture 12' reconstructed by decoder 20 deviates from picture 12 initially encoded by apparatus 10 with respect to coding loss introduced by quantization of the predictive residual signal. While Figures 1 and 2 use transform-based predictive residual coding as an example, embodiments of the present application are not limited to this type of predictive residual coding. This also applies to other details described with respect to Figures 1 and 2, as outlined below.
[0164] The encoder 10 is configured to perform a spatial-spectral transformation on the predicted residual signal and encode the resulting predicted residual signal into the data stream 14. Similarly, the decoder 20 is configured to decode the predicted residual signal from the data stream 14 and perform a spectral-space transformation on the resulting predicted residual signal.
[0165] Internally, the encoder 10 may include a predictive residual signal generator 22 that generates a predictive residual 24 to measure the deviation of the original signal, i.e., the predicted signal 26 from the picture 12. The predictive residual signal generator 22 may be, for example, a subtractor that subtracts the predicted signal from the original signal, i.e., from the picture 12. The encoder 10 then further includes a converter 28 that performs a spatial-spectral transform on the predictive residual signal 24 to obtain a spectral domain predictive residual signal 24' which is quantized by a quantizer 32 also included in the encoder 10. The quantized predictive residual signal 24'' is encoded into a bitstream 14. For this purpose, the encoder 10 may optionally include an entropy coder 34 that entropy codes the predictive residual signal to be transformed and quantized into a datastream 14. The predictive signal 26 is generated by the prediction stage 36 of the encoder 10 based on the predictive residual signal 24'' which is encoded into the datastream 14 and decodeable from the datastream 14. For this purpose, the prediction stage 36 may internally include an inverse quantizer 38 that inversely quantizes the prediction residual signal 24'' to obtain a spectral domain prediction residual signal 24'''' corresponding to the signal 24' other than the quantization loss, as shown in Figure 1, and an inverse converter 40 that inversely transforms the latter prediction residual signal 24'''', i.e., spectral-spatial, to obtain a prediction residual signal 24'''' corresponding to the original prediction residual signal 24 other than the quantization loss. The coupler 42 of the prediction stage 36 then recombines the prediction signal 26 and the prediction residual signal 24'''' by addition or other means to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 can correspond to signal 12'. Next, the prediction module 44 of the prediction stage 36 generates a prediction signal 26 based on signal 46, for example, using spatial prediction, i.e., in-picture prediction, and / or temporal prediction, i.e., inter-picture prediction.
[0166] Similarly, as shown in Figure 2, the decoder 20 may be internally composed of components corresponding to the prediction stage 36 and interconnected in a manner corresponding to the prediction stage. In particular, the entropy decoder 50 of the decoder 20 can entropy decode the spectral domain prediction residual signal 24'' quantized from the data stream, in which the inverse quantizer 52, inverse converter 54, coupler 56, and prediction module 58 are interconnected and cooperate in the manner described above with respect to the module of the prediction stage 36 to recover the reconstructed signal based on the prediction residual signal 24'', and as a result, as shown in Figure 2, the output of the coupler 56 yields the reconstructed signal, i.e., picture 12'.
[0167] Although not specifically described above, it is readily apparent that encoder 10 can set several coding parameters, including prediction mode, motion parameters, etc., according to several optimization schemes, such as several rate and distortion-related criteria, i.e., methods for optimizing coding cost. For example, encoder 10 and decoder 20 and corresponding modules 44 and 58 can each support different prediction modes, such as intra-coding mode and inter-coding mode. The granularity at which the encoder and decoder switch between these prediction mode types may correspond to the subdivision of picture 12 and 12' into coded segments or coded blocks, respectively. At the level of these coded segments, for example, the picture may be subdivided into intra-coded blocks and inter-coded blocks. Intra-coded blocks are predicted based on the spatially already coded / decoded neighborhood of each block, as outlined in more detail below. Several intracoding modes exist and may be selected for each intracoding segment, including directional or angular intracoding modes in which each segment is filled into each intracoding segment by extrapolating neighboring sample values along a specific direction inherent to each directional intracoding mode. The intracoding modes may also include, for example, one or more further modes, such as a DC coding mode in which the prediction of each intracoding block assigns DC values to all samples within each intracoding segment, and / or a planar intracoding mode in which the prediction of each block is approximated or determined to be a spatial distribution of sample values described by a two-dimensional linear function over the sample positions of each intracoding block, having a tilt drive and offset of a plane defined by a two-dimensional linear function, based on adjacent samples. In contrast, the intercoding blocks may be predicted, for example, in time.In the case of an intercoded block, motion vectors can be signaled within the data stream, and these motion vectors indicate the spatial displacement of the previously coded portion of the video to which picture 12 belongs, and the previously coded / decoded picture is sampled to obtain the prediction signal for each intercoded block. This means that in addition to the residual signal coding included in the data stream 14, such as the entropy coding transformation coefficient level representing the quantized spectral domain prediction residual signal 24'', the data stream 14 can also encode several prediction parameters for the block, such as coding mode parameters for assigning coding modes to different blocks, motion parameters for the intercoded segments, and parameters for controlling and signaling the subdivision of picture 12 and 12' into their respective segments. The decoder 20 uses these parameters to subdivide the picture in the same way that the encoder did, assign the same prediction mode to the segments, perform the same predictions, and obtain the same prediction signal.
[0168] Figure 3 shows the relationship between, on the one hand, the reconstructed signal, i.e., the reconstructed picture 12', and on the other hand, the combination of the predicted residual signal 24'''' and the predicted signal 26, which are signaled in the data stream 14. As already mentioned above, the combination may be additive. In Figure 3, the predicted signal 26 is shown as the picture area subdivided into intra-coded blocks, shown exemplary with line shading, and inter-coded blocks, shown exemplary without line shading. The subdivision may be a regular subdivision of the picture area into rows and columns of square or non-square blocks, or any subdivision such as a multi-tree subdivision of picture 12 into multiple leaf blocks of varying sizes from a tree root block, such as a quad-tree subdivision, a mixture of which is shown in Figure 3, where the picture area is first subdivided into rows and columns of a tree root block, and then further subdivided into one or more leaf blocks according to a recursive multi-tree subdivision. For example, pictures 12, 12' can be subdivided into coded blocks 80, 82, which can represent coded tree blocks, and can be further subdivided into smaller coded blocks 83.
[0169] In this case as well, the data stream 14 may have an intra-coding mode coded for the intra-coding block 80, which assigns one of several supported intra-coding modes to each intra-coding block 80. For the inter-coding block 82, the data stream 14 may have one or more coded motion parameters. Generally speaking, the inter-coding block 82 is not limited to being coded in time. Alternatively, the inter-coding block 82 may be any block predicted from a previously coded portion beyond the current picture 12 itself, such as a previously coded picture of the video to which the picture 12 belongs, or, if the encoder and decoder are scalable encoders and decoders, a picture of another view or a hierarchically lower layer.
[0170] The predicted residual signal 24'''' in Figure 3 is also shown as a subdivision of the picture region into block 84. These blocks may be called transformation blocks to distinguish them from coded blocks 80 and 82. In practice, Figure 3 shows that the encoder 10 and decoder 20 may use two different subdivisions of picture 12 and picture 12' into blocks, namely one subdivision into coded blocks 80 and 82, and the other subdivision into transformation block 84. Both subdivisions may be the same, i.e., each coded block 80 and 82 may simultaneously form a transformation block 84, but Figure 3 shows, for example, that the subdivision into transformation block 84 forms an extension of the subdivision into coded blocks 80 and 82 such that any boundary between the two blocks 80 and 82 covers the boundary between the two blocks 84, or that each block 80 and 82 coincides with one of the transformation blocks 84 or coincides with a cluster of transformation blocks 84. However, these subdivisions may also be determined or selected independently of each other, so that the transformation block 84 can alternatively cross the block boundary between blocks 80 and 82. Therefore, as far as the subdivision to the transformation block 84 is concerned, the same description as that presented for the subdivision to blocks 80 and 82 is true: namely, block 84 may be the result of regular subdivision of the picture area into blocks (with or without row and column arrangement), the result of recursive multi-tree subdivision of the picture area, or a combination thereof, or any other type of block formation. Note that blocks 80, 82, and 84 are not limited to squares, rectangles, or any other shape.
[0171] Figure 3 further shows that the combination of the prediction signal 26 and the prediction residual signal 24'''' directly yields the reconstructed signal 12'. However, it should be noted that, according to an alternative embodiment, multiple prediction signals 26 can be combined with the prediction residual signals 24'''' to form the picture 12'.
[0172] As outlined above, Figures 1 to 3 are presented as examples that can implement the concepts of the present invention, which will be further described below, in order to form specific examples of encoders and decoders according to this application. To that extent, the encoder and decoder in Figures 1 and 2 can represent possible implementations of encoders and decoders as described herein. However, Figures 1 and 2 are merely examples. Nevertheless, an encoder according to an embodiment of this application can perform block-based coding of picture 12 using a concept outlined in more detail below. Similarly, a decoder according to an embodiment of this application can perform block-based decoding of picture 12' from data stream 14 using a coding concept outlined in more detail below, but may differ from, for example, the decoder 20 in Figure 2 in that it subdivides the same picture 12' into blocks in a different way than described with respect to Figure 3, and / or in that it does not derive a predicted residual from data stream 14 in the spatial domain, although this is in the transformation domain.
[0173] 2. Fractional sample shown in Figure 4 Embodiments of this disclosure may be part of the concept of motion-compensated prediction, which may be an example of inter-picture prediction, as can be applied by prediction modules 44, 58. MCP may be performed with fractional sample precision. For example, a sample may refer to a pixel or position in a picture or video sequence, a full sample position may refer to a signaled position in the picture or video sequence, and a fractional sample position may refer to a position between full sample positions.
[0174] Figure 4 shows integer samples 401 (uppercase shaded blocks) and fractional sample positions 411 (lowercase unshaded blocks) of quarter-sample lumer interpolation. The samples shown in Figure 4 may form or be part of a portion of a video sequence. Furthermore, fractional sample positions 411 at quarter-sample resolution are shown as unshaded blocks and in lowercase. For example, portion 400 may represent a lumer-coded block or be part of a lumer-coded block. Thus, Figure 4 can represent a scheme for quarter-sample lumer interpolation. However, a lumer-coded block is merely one example of a coded block that may be used in the context of MCP. MCP may also be performed using chroma samples, or other schemes of color planes, such as samples of a video sequence using three color planes. For example, fractional sample positions 411 with quarter-sample resolution include half-sample positions 415 and quarter-sample positions 417.
[0175] The fractional sample position 411 can be obtained from the full sample 401 by an interpolation filter. Different schemes for the interpolation filter may be possible. A scheme for the interpolation filter can provide a set of interpolation filters, each of which represents a rule for obtaining one or more specific fractional sample positions.
[0176] As an example of an interpolation filter, the method used in HEVC to obtain 15 fractional lumen sample positions, for example, the unshaded fractional sample position of block 412, is explained with reference to Figure 4. The HEVC method for the interpolation filter utilizes the FIR filters listed in Table 1.
[0177] 2.1 Interpolation Filters in HEVC In HEVC, which uses motion vector accuracy of one-quarter lumar samples, there are 15 fractional lumar sample positions (as shown in Figure 4, for example) that are computed using various one-dimensional FIR filters. However, for each fractional sample position, the interpolation filter is fixed.
[0178]
Table 1
[0179] This method includes one symmetric 8-tap filter h[i] for the half-sample (HPEL) position, and position b 0,0 is calculated by horizontally filtering from integer sample position A -3,0 to A 4,0 and position h 0,0 is calculated by vertically filtering from integer sample position A 0,-3 to A 0,4 and position j 0,0 is calculated by first calculating from HPEL position b 0,-3 to b 0,4 through horizontal filtering as described above, and then vertically filtering these using the same 8-tap filter. b 0,0 =Σ i=-3..4 A i,0 h[i], h 0,0 =Σ j=-3..4 A 0,j h[j], j 0,0 =Σ j=-3..4 b 0,j h[j]
[0180] One asymmetric 7-tap filter with coefficients q[i] for the quarter-sample (QPEL) position: Again, positions with only one fractional sample component are calculated by 1D filtering of the nearest integer sample position in the vertical or horizontal direction. a0,0 = Σi=-3..3 Ai,0 q[i], c0,0 = Σi=-2..4 Ai,0 q[1-i] d0,0 = Σj=-3..3 A0,j q[j], n0,0 = Σj=-2..4 A0,j q[1-j]
[0181] All other QPEL positions are calculated by vertical filtering of the nearest fractional sample positions. For positions with a vertical HPEL component, an HPEL 8-tap filter is used.
[0182] j0,0=Σj=-3..4 a0,jh[j] or i0,0=Σj=-3..4 a0,jh[j],k0,0=Σj=-3..4 c0,jh[j] For positions with a vertical QPEL component, a QPEL 7-tap filter is used. e0,0=Σj=-3..4 a0,jh[j],f0,0=Σj=-3..4 b0,jh[j],g0,0=Σj=-3..4 c0,jh[j]
[0183] p0,0=Σj=-2..4 a0,jq[1-j],q0,0=Σj=-2..4 b0,jq[1-j],r0,0=Σj=-2..4 c0,jq[1-j] By applying this method, fractional sample positions can be obtained as shown in Figure 4.
[0184] 3. Encoders according to Figures 5, 6, and 13 Figure 5 shows an encoder 500 for hybrid video coding according to one embodiment. The encoder 500 is configured to provide an encoded representation 514 of a video sequence based on input video content 512. The encoder 500 is configured to determine one or more syntax elements 520 related to a portion 585 of the video sequence, for example by a syntax determination module 521. Furthermore, the encoder 500 is configured to select a processing scheme 540 to be applied to the portion 585 of the video sequence based on the characteristics described by one or more syntax elements 520. For example, the encoder 500 includes a processing scheme selector 541 for selecting the processing scheme 540. The processing scheme 540 is for obtaining samples, e.g., samples 401, 411, 415, 417, for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence. The encoder is configured, for example, to encode an index 550 by an index encoder 561, where the index 550 indicates a selected processing scheme 540 such that a given encoded index value represents a different processing scheme depending on the characteristics described by one or more syntax elements 520. Furthermore, the encoder 500 is configured to provide a bitstream containing one or more syntax elements 520 and an encoded index 560 as an encoded representation 514 of a video sequence. For example, the encoder 500 includes a coder 590, such as an entropy coder, for providing the encoded representation 514. The index encoder 561 may be part of or separate from the coder 590.
[0185] Figure 6 shows an encoder 600 for hybrid video coding according to a further embodiment. The encoder 600 is configured to provide an encoded representation 614 of a video sequence based on input video content 512. The encoder 600 is configured to determine one or more syntax elements 520 related to a portion 585 of the video sequence, for example by a syntax determination module 521. Furthermore, the encoder 500 is configured to select a processing scheme 540 to be applied to the portion 585 of the video sequence based on the characteristics described by one or more syntax elements 520. For example, the encoder 500 includes a processing scheme selector 541 for selecting the processing scheme 540. The processing scheme 540 is for obtaining samples, e.g., samples 401, 411, 415, for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence. The encoder 600 is configured to select an entropy coding scheme 662 used to provide an encoded index 660 indicating the selected processing scheme 540, depending on the characteristics described by one or more syntax elements 520. Furthermore, the encoder 600 is configured to provide a bitstream as an encoded representation 614 of the video sequence, which includes one or more syntax elements 520 and an encoded index 660. The encoded index value can represent the value of the encoded index 660.
[0186] Figure 13 shows an encoder 1300 for hybrid video coding according to a further embodiment. The encoder 1300 is configured to provide an encoded representation 1314 of a video sequence based on input video content 512. The encoder 1300 is configured to determine one or more syntax elements 520 related to a portion 585 of the video sequence.
[0187] According to one embodiment, the encoder 1300 is configured to select a processing method 540 to be applied to a portion 585 of a video sequence based on a characteristic described by one or more syntax elements 520, the processing method being entirely specified by the characteristics of the portion (585) and fractional parts of the motion vector, and the processing method 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 585 of the video sequence.
[0188] According to an alternative embodiment, the encoder 1300 is configured to select a processing method 540 to be applied to a portion 585 of a video sequence based on the motion vector accuracy described by one or more syntax elements 520.
[0189] According to another alternative embodiment, the encoder 1300 is configured to select a processing method 540 to be applied to a portion 585 of a video sequence based on a characteristic described by one or more syntax elements 520, at a PU granularity or CU granularity.
[0190] Furthermore, the encoder 1300 is configured to provide a bitstream containing one or more syntax elements 520 as an encoded representation 1314 of the video sequence.
[0191] According to one embodiment, the encoder 1300 is configured to select a different processing method than the standard processing method for the HPEL position of the acquired sample if one or more determined syntax elements 520 exhibit MV accuracy of HEPL and the prediction mode is one of translation-inter, translation-merge, and combined inter / intra.
[0192] The following description applies equally to encoder 500 in Figure 5, encoder 600 in Figure 6, and encoder 1300 in Figure 13, unless otherwise specified.
[0193] For example, encoders 500, 600, and 1300 receive a video sequence as input video content and provide an encoded representation 514, 614, and 1314 of the video sequence, which may contain multiple video frames. The syntax determination module 521 can obtain a portion 585 of the video sequence, for example, by dividing the video sequence or its video frames, or by obtaining a portion 585 from a module for dividing a video sequence. The portion 585 may be the result of segmenting the video sequence into coded tree units or coded units or coded tree blocks or coded blocks, for example, as described with respect to Figure 3. For example, the portion 585 may arise from segmenting the video sequence used for inter-picture or intra-picture prediction by prediction modules 44 and 58 in Figure 1 or 2. The syntax determination module 521 can further determine one or more syntax elements 520 based on the portion 585 of the video sequence and optionally further portions.
[0194] Encoders 500, 600, and 1300 optionally include a processing module 544 for applying processing scheme 540 to section 585 to obtain a prediction signal, as indicated by the dashed arrow. For example, the prediction signal may be used to obtain an encoded representation 514, 614, and 1314 of a video sequence. For example, encoding of input video content 512 to an encoded representation 514, 614, and 1314 can be performed according to the scheme described with respect to encoder 10 in Figure 1, where picture 12 may represent video content 512 and bitstream 14 may include an encoded representation 514, 614, and 1314. For example, processing module 544 may include or correspond to prediction module 44, and as a result, processing scheme 540 may represent the input of prediction module 44, where processing scheme 540 for MCP can be used, for example, to provide a prediction signal 26 using AMVR as introduced in the introduction of the description.
[0195] According to one embodiment, the processing method 540 is a filter or a set of filters. For example, one or more filters may be FIR filters, such as interpolation filters.
[0196] For example, the processing method 540 may be an interpolation filter such as the 8-tap and 7-tap filters described in Section 2.1, or a similar interpolation filter, which may be used, for example, to obtain fractional sample positions for MCP. In embodiments of the present invention, two additional 1D 6-tap low-pass FIR filters may be supported and applied instead of the HEVC 8-tap filter (with coefficients h[i] listed in Table 2). Below, an interpolation filter may function as an example of a processing method.
[0197] For example, the processing method 540 may have spectral characteristics such as low-pass or high-pass characteristics, and as a result, the encoders 500, 600, and 1300 may be configured to select between processing methods having different characteristics. According to one embodiment, the encoders 500 and 600 are configured to select different processing methods for different blocks within a video frame or picture.
[0198] For example, a video frame in a video sequence can be divided into multiple blocks or parts by encoders 500, 600, and 1300, for example, as shown by blocks 80, 82, and 83 in Figure 3. For example, encoders 500, 600, and 1300 can obtain a prediction signal 26 for each of parts 80, 82, and 83. Encoders 500, 600, and 1300 can individually select a processing method 540 for different parts of the video frame. In this regard, a part or block of a video frame can refer to a regularly sized block or first-level block resulting from the first recursion of the video frame division, such as a coded tree block or coded tree unit, for example, coded block 80 or 82. Encoders 500, 600, and 1300 can also select different processing methods for different subparts of such a part or first-level block or coded tree unit, for example, coded block 83. Encoders 500, 600, and 1300 can even select different processing methods for different parts or blocks of a coded block or coded unit.
[0199] Therefore, encoders 500, 600, and 1300 can fit interpolation filters at a fine granularity of picture or video frame in a video sequence, enabling fitting to local characteristics within a single slice, tile, or picture, i.e., fitting at a granularity smaller than the size of the slice, tile, or picture.
[0200] Therefore, encoders 500, 600, and 1300 can determine information regarding the size or shape of portion 585, such as block size or block shape, and provide this information as one or more syntax elements 520. Encoders 500, 600, and 1300 can also determine or obtain further parameters that may be used in the context of in-picture or inter-picture prediction models. For example, encoders 500, 600, and 1300 can determine the MV precision to be applied to portion 585, for example, by a processing module 544 to obtain a prediction signal for encoding portion 585.
[0201] Therefore, one or more syntax elements 520 may include at least one of the following: motion vector accuracy, fractional sample position, block size, block shape, number of prediction hypotheses, prediction mode, availability of coded residual signal, spectral characteristics of coded residual signal, reference picture signal, loop filter data, motion vector length, or adaptive motion vector resolution mode.
[0202] For example, encoders 500, 600, and 1300 can estimate a processing method 540 from already determined characteristics or processing steps, which may be indicated by one or more syntax elements 520. For example, in the case of HEVC or VVC merge modes, the interpolation filter to be used may be inherited from a selected merge candidate. In another embodiment, it may be possible to override the interpolation filter inherited from a merge candidate by explicit signaling.
[0203] According to one embodiment, encoders 500, 600, 1300 are configured to determine a set of acceptable processing schemes in terms of characteristics described by one or more syntax elements, and to select a processing scheme 540 from the determined set. For example, a processing scheme selector 541 can infer from one or more syntax elements 520 which processing scheme entities are compatible with or can be favorably applied to part 585 in order to obtain efficient encoding of a video sequence, or in particular part 585.
[0204] In other words, depending on one or more characteristics of the portion 585 indicated by one or more syntax elements 520, one or more processing methods are available, acceptable, or selectable for the portion 585. Thus, only a single processing method may be acceptable for the portion 585. In this case, encoders 500, 600 may select only the acceptable processing methods as processing method 540. Since the processing method can be uniquely identified in this case from one or more syntax elements 520, encoders 500, 600 may skip providing the encoded indices 550, 560 in the encoded representations 514, 614 to avoid redundant information and signaling overhead in the bitstream. If multiple processing methods are acceptable for the portion 585, the encoder may determine the acceptable processing methods as a set of acceptable processing methods for the portion 585 and select processing method 540 from this set of processing methods. In the latter case, the encoders 500, 600 can encode the encoded indices 560, 660 within the encoded representations 514, 614 to indicate the selected processing method 540. The embodiments described in Section 6 may be examples of such embodiments.
[0205] Therefore, in some cases, the processing method or interpolation filter can be adapted to different parts of the video frame without the need to explicitly signal the applied processing method or interpolation filter. Thus, additional bits may not be required in the bitstream containing the encoded representations 514, 516.
[0206] According to embodiments that may correspond to encoder 1300, the processing method is uniquely determined in all cases based on a characteristic described by one or more syntax elements, and as a result, explicit indication of the processing method 540 in the encoded representation 1314 having an encoded index dedicated to the processing method 540 is not required. Therefore, according to embodiments, encoder 1300 may not need to decide whether to encode an encoded index or not, or may not need to select an encoding method.
[0207] According to one embodiment, encoders 500, 600, and 1300 are configured to determine a set of available processing methods depending on the motion vector accuracy. For example, the available (or acceptable) set of interpolation filters may be determined by the motion vector accuracy selected for a block (which is signaled using a dedicated syntax element).
[0208] For example, one or more syntax elements 520 may indicate a selected MV precision for a portion 585 by using, for example, a scheme for signaling MV precision as described in Section 6.1.2 with respect to Table 3.
[0209] According to one embodiment, the encoders 500, 600, and 1300 are configured to select between a quarter-sample motion vector resolution, a half-sample motion vector resolution, a full-sample motion vector resolution, and a four-sample motion vector resolution.
[0210] According to one embodiment, the encoders 500, 600, and 1300 are configured to selectively determine a set of processing methods for motion vector accuracy between the maximum motion vector accuracy and the minimum motion vector accuracy, or between the maximum motion vector accuracy and the full sample motion vector accuracy, and to select processing method 540 from the determined set.
[0211] For example, in one embodiment, multiple sets of interpolation filters can be supported at half the accuracy of lumar samples, while a default set of interpolation filters can be used for all other possible motion vector accuracies.
[0212] According to one embodiment of encoders 500 and 600, when the AMVR mode exhibits HPEL accuracy (amvr mode=2), the interpolation filter can be explicitly signaled by one syntax element (if_idx) per CU. If if_idx is equal to 0, it indicates that the filter coefficients h[i] of the HEVC interpolation filter are used to generate the HPEL position.
[0213] If if_idx is equal to 1, it indicates that the filter coefficients h[i] of the flat-top filter (see table below) are used to generate the HPEL position. If if_idx is equal to 2, it indicates that the filter coefficients h[i] of the Gaussian filter (see table below) are used to generate the HPEL position.
[0214] When a CU is coded in merge mode and merge candidates are generated by copying motion vectors from adjacent blocks without modifying the motion vectors, the index if_idx is also copied. If if_idx is greater than 0, the HPEL position is filtered using one of two non-HEVC filters (flat-top or Gaussian).
[0215] [Table 2]
[0216] To obtain the encoded indices 560 and 660, encoders 500 and 600 can apply mapping rules that map indices representing index 550 or a selected processing scheme 540 to the encoded indices 560 and 660. For example, the mapping rules may correspond to a binarization scheme. Encoders 500 and 600 can adapt the encoded indices 560 and 660 according to several acceptable processing schemes such that the encoded indices 560 and 660 require a low number of bits.
[0217] For example, encoders 500 and 600 can select a processing method, and therefore a mapping rule, depending on the motion vector accuracy and / or the fractional sample position or fractional part of the motion vector.
[0218] The encoders 500, 600 can further adapt the number of bins used to provide the encoded indices 560, 660 depending on the characteristics described by one or more syntax elements 520, for example, by adapting a mapping rule.
[0219] Therefore, encoders 500 and 600 select a mapping rule that maps the index values to the encoded index values such that the number of bins representing the encoded index values matches the number of acceptable processing methods for the portion 585 of the video sequence.
[0220] For example, the number of bins used to provide the encoded index 560 may depend on the number of processing methods allowed in terms of characteristics. For example, if only one processing method is allowed, the number of bins may be 0. Table 2 shows three examples of allowed processing methods and up to three bins for encoded indices 560 and 660, along with the corresponding mapping rules for mapping index values to encoded index values.
[0221] For example, Table 2 shows one embodiment of a mapping rule for mapping if_idx to an encoded index value. The mapping rule may differ for different MV prisms.
[0222] In the embodiments described in Table 2, the interpolation filter may be modified only for HPEL positions. If the motion vector points to an integer lumen sample position, no interpolation filter is used in all cases. If the motion vector points to a reference picture position representing a full sample x (or y) position and a half sample y (or x) position, the selected half-HPEL filter is applied only vertically (or horizontally). Furthermore, only the filtering of the lumen component is affected; for the chroma component, the conventional chroma interpolation filter is used in all cases.
[0223] The if_idx index can correspond to index 550, for example, as shown in Table 2, and the binarized value of if_idx can correspond to the encoded indices 560 and 660. Thus, the index encoder 561 can derive the encoded index 560 from index 550 based on the binarization scheme, for example, the binarization scheme shown in the last column of Table 2. In the context of encoder 600, the binarization scheme shown in the last column of Table 2 may correspond to the entropy coding scheme 662, or to a part of the entropy coding scheme 662 selected by the coding scheme selector 661, and based on this, the coder 690 can provide an encoded index 660 that can correspond to, for example, the binarized value shown in Table 2, or its encoded representation. For example, a characteristic that allows encoder 600 to select the entropy coding scheme 662 may correspond to MV precision. Thus, Table 2 shows an example where one or more of the syntax elements 520 indicate that the MV precision of part 585 is HPEL.
[0224] According to one embodiment, encoders 500, 600 are configured to determine whether a single processing scheme 540 is acceptable for a portion 585 of a video sequence, and to selectively omit the inclusion of the encoded index in response to the finding that only a single processing scheme 540 is acceptable for a portion 585 of a video sequence. For example, encoder 500 may omit the inclusion of the encoded index 560 in the encoded representation 514. Encoder 600 may select, for example, an entropy coding scheme 662 that indicates that the encoded index 660 is not included in the encoded representation 614, or that the number of bins or bits used for the encoded index 660 in the bitstream is 0.
[0225] According to one embodiment, adaptive filter selection is extended to chroma components. For example, encoders 500, 600 can select a processing method 540 according to the characteristics of the chroma components. The chroma components may be parts of a chroma-coding block that can be associated with a chroma-coding block. For example, part 585 may be part of a chroma-coding block and / or part of a chroma-coding block. 4. Decoder according to Figures 7 and 8
[0226] Figure 7 shows a decoder 700 for hybrid video coding according to one embodiment. The decoder 700 is configured to provide output video content 712 based on an encoded representation 714 of a video sequence. The encoded representation 714 can correspond to encoded representations 514, 614, 1314 provided by encoders 500, 600, 1300. The decoder 700 is configured to acquire a bitstream as an encoded representation 714 of a video sequence and to identify from the bitstream one or more syntax elements 520 related to a portion 785 of a video sequence that can correspond to a portion 585 of the input video content 512.
[0227] According to one embodiment, the decoder 700 includes a syntax identifier 721 for identifying one or more syntax elements 520 which may be encoded in a representation 714 encoded by the encoders 500, 600. Furthermore, the decoder 700 is configured to identify a processing scheme 540 for a portion of a video sequence 785 based on one or more syntax elements 520 and an index 550, the processing scheme 540 for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer 401 and / or fractional positions 411, 415, 417 within the portion of the video sequence 785, and the decoder 700 is configured to assign different processing schemes to a given encoded index value depending on one or more syntax elements 520.
[0228] For example, the processing method selector 741 can determine an index 550 from an encoded representation 714 based on one or more syntax elements 520, and then use one or more syntax elements 520 and / or index 550 to determine a processing method 540. Depending on one or more syntax elements 520, the processing method selector 741 may also infer that the encoded representation 714 and one or more syntax elements 520 do not contain index 550, and may estimate index 550 from one or more syntax elements 520.
[0229] According to an alternative embodiment, the decoder 700 is configured to select a processing method 540 for a portion 785 of a video sequence in accordance with one or more syntax elements 520, the processing method being entirely specified by the characteristics of the portion 785 and fractional parts of the motion vector, and the processing method 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence. According to another alternative embodiment, the decoder 700 is configured to select a processing method 540 for a portion of the video sequence 785 depending on the motion vector accuracy.
[0230] According to another alternative embodiment, the decoder 700 is configured to determine a processing method 540 for a portion 785 of a video sequence at PU granularity or CU granularity, and depending on one or more syntax elements 520. Furthermore, the decoder 700 is configured to apply the identified processing method 540 to a portion 785 of the video sequence, for example, by a processing module 744.
[0231] Figure 8 shows a decoder 800 for hybrid video coding according to another embodiment. The decoder 800 is configured to provide output video content 812 based on an encoded representation 714 of a video sequence. The encoded representation 714 may correspond to encoded representations 514, 614 provided by encoders 500, 600. The decoder 800 is configured to acquire a bitstream as an encoded representation 714 of a video sequence and to identify from the bitstream one or more syntax elements 520 related to a portion 785 of a video sequence that may correspond to a portion 585 of the input video content 512. For example, the decoder 700 includes a syntax identifier 821 for identifying one or more syntax elements 520 that may be encoded in the representation 714 encoded by encoders 500, 600. Furthermore, the decoder 800 is configured to determine an entropy coding scheme 662 used in the bitstream to provide an encoded index indicating a processing scheme 540 for a portion of the video sequence 785, depending on one or more syntax elements 520, the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion of the video sequence 785. For example, the decoder 800 includes a coding scheme selector 861, which may correspond to the coding scheme selector 661 of the encoder 600, and determines the coding scheme 662 based on one or more syntax elements 520. The decoder 800 is further configured to identify an encoded index in the bitstream if the determined entropy coding scheme 662 specifies a number of bins greater than zero, and to decode the index by, for example, a processing scheme selector 841. The decoder 800 is further configured to identify a processing scheme 540 based on one or more syntax elements 520 and the decoded index, if the number of bins defined by the determined coding scheme is greater than zero. For example, a processing scheme selector 841 may identify a processing scheme 540.Furthermore, the decoder 800 is configured to apply the identified processing method 540 to a portion 785 of the video sequence, which can, for example, be associated with a processing module 744.
[0232] Decoders 700, 800 can obtain output video content 712, 812 from encoded representations 514, 614, 1314 provided by encoders 500, 600 based on input video content 512, so that output video content 712, 812 can correspond to input video content 512, excluding encoding losses such as quantization losses. Thus, the description of encoders 500, 600 and their features may be equally applicable to decoders 700, 800, although some processes may be performed in reverse. For example, features of encoders 500, 600 referring to encoding may correspond to features of decoders 700, 800 referring to decoding.
[0233] For example, decoders 700, 800 can read or decode one or more syntax elements 520 from the encoded representations 714, 814, 1314, as provided by encoders 500, 600, 1300. Thus, encoders 500, 600, 1300 and decoders 700, 800, 1300 can apply equivalent assumptions or knowledge to determine a processing scheme 540 from one or more syntax elements 520.
[0234] For example, the encoded indices 560, 660 may indicate information that the decoders 700, 800 cannot infer from one or more syntax elements 520, and as a result, some embodiments of the decoders 700, 800 can determine a processing method 540 based on the encoded indices 560, 660. For example, some embodiments of the decoders 700, 800 can decode the encoded indices 560, 660 from the encoded representations 714, 814 to obtain an index, for example, index 550, if the decoders 700, 800 infer from one or more syntax elements 520 that several processing methods are acceptable. However, it should be noted that not all embodiments of the decoder 700 are configured to infer whether several processing methods are acceptable. Some embodiments of the decoder 700 may be configured to infer a processing method 540 from one or more syntax elements of the encoded representation 1314, for example, in all cases or independently of the indications of one or more syntax elements 520. Therefore, the embodiment of the decoder 700 can infer the processing method 540 from one or more syntax elements 520 without explicitly decoding an index dedicated to the processing method 540.
[0235] Therefore, further details of the features, functions, advantages, and examples of one or more syntax elements 520, processing schemes 540, parts 585, indices 550, optionally encoded indices 560, 660, and encoders 500, 600, 1300 can be applied equivalently to decoders 700, 800, although the descriptions of the elements are not explicitly repeated.
[0236] For example, decoders 700, 800 may be implemented in decoder 20 in Figure 2. For example, processing modules 744, 844 may include or correspond to prediction module 58, and as a result, processing scheme 540 may represent an input for prediction module 58, and processing scheme 540 for MCP can be used to provide prediction signal 59. The relationship between encoders 500, 600 and decoders 700, 800 may correspond to the relationship between encoder 10 and decoder 20, as described in Section 1.
[0237] According to one embodiment, the decoders 700, 800 are configured to determine whether a single processing method 540 is acceptable for a portion 785 of a video sequence, and to selectively omit the identification of an encoded index in response to the finding that only a single processing method 540 is acceptable for a portion 785 of a video sequence.
[0238] For example, the decoder 700 can determine, based on one or more syntax elements 520, whether a single processing method 540 is acceptable for part 785. Depending on the result, the decoder 700 can identify or decode encoded indices that can correspond to encoded indices 560, 660, or the decoder 700 can omit or skip the identification of encoded indices and select a single acceptable processing method for part 785.
[0239] In a further example, the decoder 800 can estimate from the entropy coding scheme 662 whether to decode an encoded index from the encoded representation 814. If an encoded index is to be decoded, the decoder 800 can decode an encoded index that may correspond to the encoded index 560, 660 to determine an index that may correspond to index 550, and then determine a processing scheme 540 based on the index. If there is no encoded index to be decoded, the decoder 800 can omit or skip the decoding of an encoded index and determine a processing scheme 540 based on one or more syntax elements 520.
[0240] According to one embodiment, decoders 700, 800 are configured to select a filter according to the fractional portion of the motion vector. For example, the motion vector may be represented by two components that refer to two different directions. For example, the motion vector may be represented by a horizontal component and a vertical component. Horizontal and vertical can represent two orthogonal directions in the plane of the picture or video frame. Thus, the fractional portion of the motion vector can refer to the component of the MV having fractional resolution.
[0241] According to one embodiment, decoders 700, 800 are configured to select different filters for filtering in a first direction and filtering in a second direction, depending on the fractional portion of the motion vector. For example, decoders 700, 800 can select different filters according to a first component of MV representing the MV component in the first direction and a second component of MV representing the MV component in the second direction.
[0242] The encoded representations 514, 614, 714, and 1314 may be portions of the video bitstream provided by the encoders 500, 600, and 1300, the video bitstream containing encoded representations of multiple portions of the video sequence of the input video content.
[0243] According to one embodiment, a video bitstream for hybrid video coding, the video bitstream describing a video sequence, includes one or more syntax elements 520 associated with a portion 585 of the video sequence. The video bitstream includes a scheme 540 index bitstream element for at least some of the portions of the video sequence, the presence of which the scheme 540 index bitstream element varies depending on one or more other syntax elements 520. For example, the video bitstream may correspond to or include encoded representations 514, 614, 714.
[0244] According to one embodiment, a video bitstream for hybrid video coding, the video bitstream describing a video sequence, includes one or more syntax elements 520 associated with a portion 585 of the video sequence. The video bitstream includes index bitstream elements of a processing scheme 540 for at least some of the portions of the video sequence, the number of bins in the index bitstream elements of the processing scheme 540 varies depending on the one or more syntax elements 520 associated with the portion 585 of the video sequence. For example, the video bitstream may correspond to or include encoded representations 514, 614, 714.
[0245] According to one embodiment, a video bitstream for hybrid video coding, the video bitstream describing a video sequence, includes one or more syntax elements 520 associated with a portion 585 of the video sequence. The video bitstream includes index bitstream elements of a processing scheme 540 for at least some of the portions of the video sequence, the meaning of the codewords of the index bitstream elements of the processing scheme 540 changes depending on one or more syntax elements 520 associated with a portion 585 of the video sequence. For example, the video bitstream may correspond to or include encoded representations 514, 614, 714. 5. Method for hybrid video coding according to Figures 9, 10, 11, and 12.
[0246] Figure 9 shows a flowchart of a method 5000 for hybrid video coding according to one embodiment. The method 5000 is for providing an encoded representation 514 of a video sequence based on input video content 512, the method 5000 includes step 5001 determining one or more syntax elements 520 related to a portion 585 of the video sequence. Step 5002 includes selecting a processing scheme 540 to be applied to the portion 585 of the video sequence based on characteristics described on one or more syntax elements 520, the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 585 of the video sequence. Step 5003 includes encoding an index indicating the selected processing scheme 540 such that a given encoded index value represents a different processing scheme depending on the characteristics described by one or more syntax elements 520. Step 5004 includes providing a bitstream as an encoded representation of the video sequence, the bitstream including one or more syntax elements 520 and the encoded index.
[0247] Figure 10 shows a flowchart of a method 6000 for hybrid video coding according to one embodiment. The method 6000 provides an encoded representation of a video sequence based on input video content, the method 6000 includes step 6001 determining one or more syntax elements 520 related to a portion 585 of the video sequence. Step 6002 includes selecting a processing scheme 540 to be applied to the portion 585 of the video sequence based on characteristics described on one or more syntax elements 520, the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 585 of the video sequence. Step 6003 includes selecting an entropy coding scheme used to provide an encoded index indicating the selected processing scheme 540, depending on the characteristics described by one or more syntax elements 520. Step 6004 includes providing a bitstream as an encoded representation of the video sequence, which includes one or more syntax elements 520 and an encoded index.
[0248] Figure 11 shows a flowchart of a method 7000 for hybrid video coding according to one embodiment. The method 7000 is for providing output video content based on an encoded representation of a video sequence, the method 7000 includes step 7001 of obtaining a bitstream as an encoded representation of the video sequence. Step 7002 includes identifying one or more syntax elements 520 related to a portion 785 of the video sequence from the bitstream. Step 7003 includes identifying a processing scheme 540 for the portion 785 of the video sequence based on one or more syntax elements 520 and an index, the processing scheme 540 for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence, the method assigns different processing schemes to a given encoded index value depending on one or more syntax elements 520. Step 7004 includes applying the identified processing scheme 540 to the portion 785 of the video sequence.
[0249] Figure 12 shows a flowchart of Method 8000 for Hybrid Video Encoding According to One Embodiment. Method 8000 is for providing output video content based on an encoded representation of a video sequence, the method comprising step 8001 obtaining a bitstream as an encoded representation of the video sequence. Step 8002 comprises identifying one or more syntax elements 520 from the bitstream that relate to a portion 785 of the video sequence. Step 8003 comprises determining an entropy encoding scheme used in the bitstream to provide an encoded index indicating a processing scheme 540 for the portion 785 of the video sequence, depending on the one or more syntax elements 520, the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence. Step 8004 comprises identifying an encoded index in the bitstream and decoding the index if the determined entropy encoding scheme specifies a number of bins greater than zero. Step 8005 includes identifying a processing scheme 540 based on one or more syntax elements 520 and a decoded index, if the number of bins defined by the determined entropy coding scheme is greater than zero. Step 8006 includes applying the identified processing scheme 540 to a portion 785 of the video sequence.
[0250] Figure 14 shows a flowchart of a method 7100 for hybrid video coding according to one embodiment. The method 7100 is for providing output video content 712 based on an encoded representation 1314 of a video sequence, and the method 7100 includes: 7101 obtaining a bitstream as an encoded representation 714 of a video sequence; 7102 identifying one or more syntax elements 520 related to a portion 785 of a video sequence from the bitstream; 7103 selecting a processing scheme 540 for the portion 785 of the video sequence in accordance with one or more syntax elements 520, the processing scheme being entirely specified by the characteristics of the portion 785 and fractional parts of the motion vector, the processing scheme 540 being for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence; and 7104 applying the identified processing scheme 540 to the portion 785 of the video sequence.
[0251] Figure 15 shows a flowchart of a method 7200 for hybrid video coding according to one embodiment. The method 7200 is for providing output video content 712 based on an encoded representation 1314 of a video sequence, and the method 7200 includes: 7201 obtaining a bitstream as an encoded representation 714 of a video sequence; 7202 identifying one or more syntax elements 520 related to a portion 785 of a video sequence from the bitstream; 7203 selecting a processing scheme 540 for the portion 785 of the video sequence depending on the motion vector accuracy, wherein the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence; and 7204 applying the identified processing scheme 540 to the portion 785 of the video sequence.
[0252] Figure 16 shows a flowchart of a method 7300 for hybrid video coding according to one embodiment. The method 7300 is for providing output video content 712 based on an encoded representation 1314 of a video sequence, and the method 7300 comprises: 7301 obtaining a bitstream as the encoded representation 714 of the video sequence; 7302 identifying one or more syntax elements 520 related to a portion 785 of the video sequence from the bitstream; and 7303 determining a processing scheme 540 for the portion 785 of the video sequence, depending on one or more syntax elements 520, using PU granularity or CU granularity, wherein the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions within the portion 785 of the video sequence. This includes applying the identified processing method 540 to a portion 785 of a video sequence 7304.
[0253] Figure 17 shows a flowchart of a method 1310 for hybrid video coding according to one embodiment. The method 1310 is for providing an encoded representation 1314 of a video sequence based on input video content 512, and the method 1310 includes determining one or more syntax elements 520 related to a portion 585 of the video sequence 1311, selecting a processing scheme 540 to be applied to the portion 585 of the video sequence based on the characteristics described by the one or more syntax elements 520 1312, the processing scheme being entirely specified by the characteristics of the portion 585 and fractional parts of the motion vector, the processing scheme 540 being for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions in the portion 585 of the video sequence 1313, and providing a bitstream containing one or more syntax elements 520 as the encoded representation 1314 of the video sequence.
[0254] Figure 18 shows a flowchart of a method 1320 for hybrid video coding according to one embodiment. The method 1320 is for providing an encoded representation 1314 of a video sequence based on input video content 512, and the method 1320 includes determining one or more syntax elements 520 related to a portion 585 of the video sequence 1321, selecting a processing scheme 540 to be applied to the portion 585 of the video sequence based on the motion vector precision described by the one or more syntax elements 520 1322, the processing scheme 540 for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions in the portion 785 of the video sequence 1323, and providing a bitstream containing one or more syntax elements 520 as the encoded representation 1314 of the video sequence 1323.
[0255] Figure 19 shows a flowchart of a method 1330 for hybrid video coding according to one embodiment. The method 1330 is for providing an encoded representation 1314 of a video sequence based on input video content 512, and the method 1330 includes determining one or more syntax elements 520 related to a portion 585 of the video sequence 1331, selecting a processing scheme 540 to be applied to the portion 585 of the video sequence based on the characteristics described by the one or more syntax elements 520, using PU granularity or CU granularity 1332, the processing scheme 540 is for obtaining samples 401, 411, 415, 417 for motion compensation prediction at integer and / or fractional positions in the portion 785 of the video sequence 1333, and providing a bitstream containing the one or more syntax elements 520 as the encoded representation 1314 of the video sequence 1333. 6. Further Embodiments
[0256] In many, or in some cases all, current video coding standards, including the current VVC draft [5], the interpolation filter is fixed for each fractional sample position. (As used herein, "fractional sample position" refers to the relative position between four surrounding FPEL or integer positions.) Embodiments of the present invention enable the use of different interpolation filters (e.g., processing methods) for the same fractional sample position in different parts (e.g., portions) of a video sequence. The aforementioned technique for switching interpolation filters (see the "Background Art of the Invention" section) has the disadvantages of enabling switching between different interpolation filters only at a coarse granularity (i.e., slice level), and, in some cases (with respect to AIF and its variants), of having the selection be explicitly signaled by transmitting individual FIR filter coefficients.
[0257] By enabling the switching of interpolation filters at a fine granularity (e.g., PU / CU, or CTU level), better adaptation can be achieved according to local image characteristics (e.g., in addition to high-precision HEVC interpolation filters, it may be beneficial to use filters with different characteristics, such as filters with stronger low-pass characteristics for attenuating high-frequency noise components). However, explicit signaling of the interpolation filter used for each PU may be too high from the perspective of the required bitrate. Therefore, embodiments of the present invention avoid this signaling overhead by restricting the interpolation filters available for a portion or region of the video sequence (e.g., PU / CU, CTU, slice, frame, etc.) according to the characteristics of a given portion or region of the video sequence known to the decoder (e.g., because they are indicated by one or more syntax elements 520 that may be part of the encoded representations 514, 614, 1314). If only one interpolation filter remains available, no further signaling is required. If multiple remain available, only which of the available interpolation filters is used may be transmitted. (It should be noted that herein and hereinafter, "interpolation filter" means a broader meaning of any signal processing applied to obtain a reference sample at a given, possibly fractional sample position, and may also include filtering at FPEL positions.)
[0258] According to embodiments of the present invention, a change in the interpolation filter used may only be possible if the current block has certain characteristics (given by the values of other decoded syntax elements). In this case, two cases can be distinguished as follows.
[0259] Case 1: For a given property of a block (determined by other syntax elements), only one specific set of interpolation filters is supported. Here, the term “set of interpolation filters” refers to the fact that the filters actually used also depend on the subsample position (as in the case of conventional video coding with subsample precision). The property then determines the set of interpolation filters to be used, and the fractional part of the motion vector determines the actual filters (or horizontal and vertical filters if filtering is specified by separable filters). No additional syntax elements are sent in this case, as the interpolation filters are entirely specified by the properties of the block and the fractional part of the motion vector.
[0260] Case 2: Multiple sets of interpolation filters are supported for a given property of a block. In this case, the set of interpolation filters can be selected by coding the corresponding syntax element (e.g., an index to a list of supported sets). The interpolation filter used (from the selected set) can again be determined by the fractional part of the motion vector.
[0261] It should be noted that both cases may be supported by the same picture. This means that for blocks having the first characteristic, a fixed set of interpolation filters can be used (this set may be the same as or deviate from the conventional set of interpolation filters). Also, for blocks having the second characteristic, multiple sets of interpolation filters may be supported, and the selected filter is determined by the submitted filter index (or any other means). Furthermore, the block may have a third characteristic, in which case one or more of other sets of interpolation filters may be used. These characteristics of a portion or region of a video sequence (a portion of a video sequence) that can be used to limit the available interpolation filters include:
[0262] -MV precision (e.g., QPEL, HPEL, FPEL, etc.), i.e., different interpolation filter sets may be supported for different motion vector precisions. - Fractional sample position - Block size, or more generally, block shape - Number of prediction hypotheses (i.e., 1 for single prediction, 2 for dual prediction, etc.), e.g., additional smoothing filters for single prediction only - Predictive modes (e.g., translated interconnect, affine interconnect, translated merge, affine merge, combined interconnect / intra) - Availability of coded residual signals (e.g., coded block flags in HEVC) - Spectral characteristics of coded residual signals -Reference picture signal (a reference block defined by motion data), for example, • Block edges from co-located PUs within a reference block • High-frequency or low-frequency characteristics of the reference signal - Loop filter data, for example, • Sample-Adaptive Offset Filter (SAO) Edge offset or bandwidth offset classification from • Determination of deblocking filters and boundary strength - Allows for additional smoothing filters only for long MVs or specific directions, e.g., MV length. Two or more combinations of these characteristics (e.g., MV accuracy and fractional sample position) are also possible (see above).
[0263] In one embodiment of the present invention, available interpolation filters may be directly linked to MV precision. In particular, additional or different interpolation filters (with respect to HEVC or VVC interpolation filters) may be available only for blocks encoded using a specific MV precision (e.g., HPEL) and applied, for example, only to HPEL locations. In this case, MV precision signaling can be performed in a manner similar to AMVR for VVC (see Sec. 6.1.2). If only one interpolation filter (i.e., replacing the corresponding HEVC or VVC interpolation filter) is available, no further signaling is required. If two or more interpolation filters are available, the indication of the interpolation filter used may be signaled only for blocks using a given MV precision, but the signaling remains unchanged for all other blocks.
[0264] In one embodiment of the present invention, in the case of HEVC or VVC merge mode, the interpolation filter used may be inherited from the selected merge candidate. In another embodiment, it may be possible to override the interpolation filter inherited from the merge candidate by explicit signaling. 6.1 Further Embodiments of the Invention
[0265] In one embodiment of the present invention, the set of available interpolation filters may be determined by the motion vector precision selected for a block (signaled using a dedicated syntax element). For half-luma sample precision, multiple sets of interpolation filters may be supported, but for all other possible motion vector precisions, a default set of interpolation filters is used. 6.1.1 Interpolation Filter Set
[0266] In embodiments of the present invention, two additional 1D 6-tap low-pass FIR filters may be supported, which may be applied instead of the HEVC 8-tap filter (having the coefficients h[i] described above). 6.1.2 Signal Transmission
[0267] An extension of the AMVR CU syntax with discarded monobinary binarization from the current VVC draft by HPEL decomposition energy, the extension adding one CABAC context to an additional bin: - The index exists only if at least one MVD is greater than 0 - When the index is equal to 0, it indicates QPEL accuracy and the HEVC interpolation filter
[0268] - When the index is equal to 1, it indicates that all MVDs have HPEL MV accuracy and must be shifted to QPEL accuracy before being added to the MVP rounded to HPEL accuracy. When the index is equal to 2 or 3, it indicates FPEL or 4PEL accuracy respectively, and the interpolation filter is not applied.
[0269]
Table 3
[0270] The context of the first bin for switching between QPEL and other accuracies can be selected from among three contexts according to the values of the AMVR modes of the adjacent blocks to the left and above the current block. If both adjacent blocks have the QPEL AMVR mode, the context is 0; otherwise, if only one adjacent block does not use the QPEL AMVR mode, the context is 1; otherwise (if neither adjacent block uses the QPEL AMVR mode), the context is 2.
[0271] When the AMVR mode indicates HPEL accuracy (amvr mode = 2), the interpolation filter can be explicitly signaled by one syntax element (if_idx) per CU. When if_idx is equal to 0, it indicates that the filter coefficients h[i] of the HEVC interpolation filter are used to generate HPEL positions.
[0272] If if_idx is equal to 1, it indicates that the filter coefficients h[i] of the flat-top filter (see Table 2 in Section 3) are used to generate the HPEL position.
[0273] If if_idx is equal to 2, it indicates that the filter coefficients h[i] of the Gaussian filter (see Table 2 in Section 3) are used to generate the HPEL position.
[0274] If the CU is coded in merge mode and merge candidates are generated by copying motion vectors from adjacent blocks without modifying the motion vectors, the index if_idx is also copied. If if_idx is greater than 0, the HPEL position is filtered using one of two non-HEVC filters (flat-top or Gaussian).
[0275] In this particular embodiment, the interpolation filter may be modified only for HPEL positions. If the motion vector points to an integer lumen sample position, no interpolation filter is used in all cases. If the motion vector points to a reference picture position representing a full sample x (or y) position and a half sample y (or x) position, the selected half-HPEL filter is applied only vertically (or horizontally). Furthermore, only the filtering of the lumen component is affected, and for the chroma component, the conventional chroma interpolation filter is used in all cases. Possible modifications of the described embodiments include the following aspects:
[0276] The set of interpolation filters for HPEL accuracy also includes a modified filter for the full sample position. That is, if the accurate motion vector of the half-sample points to the full sample position in the x (or y) direction, a stronger low-pass filtering in the horizontal (or vertical) direction may also be applied.
[0277] If motion vector accuracy for all lumens or 4 lumens is selected, filtering of the reference sample may also be applied. In this case, either a single filter or multiple filters can be supported. Adaptive filter selection is extended to chroma components. The selected filter further depends on whether single or dual prediction is used. 7. Alternative implementations:
[0278] Some embodiments are described in the context of apparatus, but these embodiments also represent a description of the corresponding method, and it is clear that the block or apparatus corresponds to a method step or a feature of a method step. Similarly, embodiments described in the context of a method step also represent a description of an item or feature of the corresponding block or corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0279] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Embodiments can be implemented using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which have electronically readable control signals stored therein and cooperate with (or are capable of cooperating with) a computer system that is programmable for each method to be performed. Thus, the digital storage media may be computer-readable.
[0280] Some embodiments of the present invention include a data carrier having an electronically readable control signal, such that one of the methods described herein is performed in cooperation with a programmable computer system.
[0281] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier. Other embodiments include a computer program for performing one of the methods described herein, which is stored on a machine-readable carrier.
[0282] In other words, embodiments of the method of the present invention are computer programs having program code for performing one of the methods of the present invention when the computer program is executed on a computer.
[0283] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) recorded therein, which includes a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-temporary.
[0284] Therefore, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.
[0285] Further embodiments include processing means configured to perform or applied to perform one of the methods described herein, such as a computer or a programmable logic device. Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0286] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0287] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessing unit to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0288] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer. The apparatus described herein, or any component of the apparatus described herein, may be implemented at least partially in hardware and / or software.
[0289] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0290] Any method described herein, or any component of the apparatus described herein, may be performed at least in part by hardware and / or software.
[0291] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended that the invention be limited only by the imminent claims and not by the description of the embodiments herein and by the specific details presented herein.
[0292] References [1] Yuri Vatis, Bernd Edler, Dieu Thanh Nguyen, Joern Ostermann, “Two-dimensional non-separable Adaptive Wiener Interpolation Filter for H.264 / AVC”, ITU-T SG 16 / Q.6 Doc. VCEG-Z17, Busan, South Korea, April 2005.
[0293] [2] S. Wittmann, T. Wedi, “Separable Adaptive Interpolation Filter,” document COM16-C219, ITU-T, June 2007.
[0294] [3] M. Karczewicz, Y. Ye, and P. Chen, “Switched interpolation filter with offset,” document COM16-C463, ITU-T, Apr. 2008.
[0295] [4] Y. Ye and M. Karczewicz, “Enhanced Adaptive Interpolation Filter,” document COM16-C464, ITU-T, Apr. 2008.
[0296] [5] B. Bross, J. Chen, S. Liu (editors), “Versatile Video Coding (Draft 4),” document JVET-M1001, Jan. 2019.
Claims
1. Encoding a syntax element comprising an instruction for Adaptive Motion Vector Resolution (AMVR) mode and an instruction for whether the interpretation mode for the coded block of the picture is translational or affine, To obtain a sample corresponding to the half-sample position of the reference picture of the coded block, the selection of an FIR interpolation filter from a plurality of finite impulse response (FIR) interpolation filters, in accordance with the instruction that the indicated AMVR mode and the interprediction mode are translational or affine, wherein the plurality of FIR interpolation filters comprises a first interpolation filter having six non-zero coefficients, an interpolation filter having six different non-zero coefficients from the first interpolation filter, and a second interpolation filter having eight non-zero coefficients. The selected FIR interpolation filter is applied to the sample of the reference picture to obtain the sample corresponding to the half-sample position, Using the sample corresponding to the half-sample position, predict the sample of the coded block, A video encoder configured to perform the following actions.
2. The video encoder according to claim 1, wherein the second interpolation filter has coefficients of -1, 4, -11, 40, 40, -11, 4, and -1.
3. Decoding syntax elements from a video bitstream, wherein the syntax elements include an instruction for Adaptive Motion Vector Resolution (AMVR) mode and an instruction for whether the interpretation mode for the coded blocks of a picture is parallel or affine. To obtain a sample corresponding to the half-sample position of the reference picture of the coded block, the selection of FIR interpolation filters from a plurality of finite impulse response (FIR) interpolation filters for different coded units of the picture, depending on the instruction that the indicated AMVR mode and the interprediction mode are translational or affine, wherein the plurality of FIR interpolation filters comprises a first interpolation filter having six non-zero coefficients and a second interpolation filter having eight non-zero coefficients. The selected FIR interpolation filter is applied to the sample of the reference picture to obtain the sample corresponding to the half-sample position, Using the sample corresponding to the half-sample position, predict the sample of the coded block, A video decoder configured to perform the following actions.
4. The video decoder according to claim 3, wherein the second interpolation filter has coefficients of -1, 4, -11, 40, 40, -11, 4, -1.
5. The video decoder according to claim 3, wherein the plurality of FIR interpolation filters comprises a plurality of filters having six non-zero coefficients.
6. The video decoder according to claim 3, wherein none of the syntax elements are explicitly dedicated to indicating the selected FIR interpolation filter.
7. A video encoding method, Encoding a syntax element comprising an instruction for Adaptive Motion Vector Resolution (AMVR) mode and an instruction for whether the interpretation mode for the coded block of the picture is translational or affine, To obtain a sample corresponding to the half-sample position of the reference picture of the coded block, the selection of an FIR interpolation filter from a plurality of finite impulse response (FIR) interpolation filters, in accordance with the instruction that the indicated AMVR mode and the interprediction mode are translational or affine, wherein the plurality of FIR interpolation filters comprises a first interpolation filter having six non-zero coefficients, an interpolation filter having six different non-zero coefficients from the first interpolation filter, and a second interpolation filter having eight non-zero coefficients. The selected FIR interpolation filter is applied to the sample of the reference picture to obtain the sample corresponding to the half-sample position, Using the sample corresponding to the half-sample position, predict the sample of the coded block, A method for providing this.
8. The method according to claim 7, wherein the second interpolation filter has coefficients of -1, 4, -11, 40, 40, -11, 4, and -1.
9. A method for video decoding, Decoding syntax elements from a video bitstream, wherein the syntax elements include an instruction for Adaptive Motion Vector Resolution (AMVR) mode and an instruction for whether the interpretation mode for the coded blocks of a picture is parallel or affine. To obtain a sample corresponding to the half-sample position of the reference picture of the coded block, the selection of FIR interpolation filters from a plurality of finite impulse response (FIR) interpolation filters for different coded units of the picture, depending on the instruction that the indicated AMVR mode and the interprediction mode are translational or affine, wherein the plurality of FIR interpolation filters comprises a first interpolation filter having six non-zero coefficients and a second interpolation filter having eight non-zero coefficients. The selected FIR interpolation filter is applied to the sample of the reference picture to obtain the sample corresponding to the half-sample position, Using the sample corresponding to the half-sample position, predict the sample of the coded block, A method for providing this.
10. The method according to claim 9, wherein the second interpolation filter has coefficients of -1, 4, -11, 40, 40, -11, 4, and -1.
11. The method according to claim 9, wherein the plurality of FIR interpolation filters comprises a plurality of filters having six non-zero coefficients.
12. The method according to claim 9, wherein none of the syntax elements are explicitly dedicated to indicating the selected FIR interpolation filter.
Citation Information
Patent Citations
JPP7547590B