Video encoding and decoding method and device
By independently selecting the interpolation filter in the video decoder based on the expansion ratio and motion vector resolution of the current picture and the reference picture, the problems of artifacts and prediction errors in the motion compensation process in the prior art are solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202080088794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-12-21
AI Technical Summary
Existing video encoding standards use phase differences of different interpolation filters during motion compensation lead to artifacts and prediction errors, increasing the code rate to achieve the same video quality.
By independently selecting the interpolation filter in the video decoder, an appropriate interpolation filter is selected according to the extended ratio and motion vector resolution between the current picture and the reference picture to reduce artifacts and improve the effect of motion compensation.
Effectively reduce artifacts in motion compensation process, improve the quality of inter-frame prediction, and reduce the code rate to achieve the same video quality.
Smart Images

Figure CN114788273B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to video coding concepts and, in particular, to interpolation filters for motion compensation. Background Art
[0002] Current video coding standards such as Versatile Video Coding (VVC) allow switching the interpolation filter used for motion compensation depending on the motion vector (MV) resolution, which can be signaled at the block level. In the case where the MV (or MV difference MVD) is coded at a specific resolution (e.g., half-sample accuracy), different interpolation filters can be used to interpolate certain fractional sample positions.
[0003] Another new feature is reference picture resampling, which allows referencing previously coded pictures in motion compensated inter-picture prediction with a different resolution / size than the current picture. To do this, the referenced picture area is resampled to a block with the same size as the current block. This may result in a situation where several fractional positions are obtained by using different phases of the interpolation filter.
[0004] For example, when a 16x16 block references a picture with a quarter size in each dimension, the corresponding 4x4 block in the referenced picture needs to be upsampled to 16x16, which may involve different interpolation filters for specific fractional positions / phases. For example, when the MV is signaled with the precision associated with a smoothing interpolation filter, the filter is applied to the phase associated with the smoothing filter in the reference picture upsample, while sharpening interpolation filters may be applied to other phases.
[0005] This mixing produces visible artifacts and therefore results in a worse motion compensated inter predictor, which in turn increases the prediction error and the bit rate required to decode the prediction residual to achieve equal quality. Summary of the invention
[0006] The present application seeks to obtain a more efficient video decoding concept supporting reference picture resampling. This object is achieved by the subject matter of the independent claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The preferred embodiments of the present application are described below with reference to the accompanying drawings, in which:
[0008] Figure 1 An apparatus for predictively encoding a picture into a data stream is shown;
[0009] Figure 2 An apparatus for predictively decoding a picture from a data stream is shown;
[0010] Figure 3 The relationship between the reconstructed signal and the combination of the prediction residual signal and the prediction signal is explained;
[0011] Figure 4a A sample of the current and reference images illustrating spatial coverage;
[0012] Figure 4b Samples of a reference picture and a current picture are illustrated having the same resolution and motion vectors defined at half-pixel resolution;
[0013] Figure 4c A video decoder according to an embodiment of the present invention is described;
[0014] Figure 4d The selection of interpolation filters for different types of samples (ie samples with different phases) in the reference picture is described (here exemplarily using Fig. 9 ); there are two options for the selection of the half-pixel resolution position, between which the embodiment is advantageously selected; in addition, for an example sample (here a quarter-pixel sample), the application of the selected interpolation filter is discussed; the filter selection between the two half-pixel interpolation filters can be performed separately for the horizontal direction and the vertical direction, so that one half-pixel interpolation filter can be used to vertically interpolate the reference picture for the half-pixel sample, and another half-pixel interpolation filter can be used horizontally for the difference at the half-pixel position, or a global selection is performed for both directions depending on whether the resolutions of the current picture and the reference picture coincide vertically and horizontally, so as to use one or the other half-pixel interpolation filter horizontally and vertically;
[0015] Figure 5 A video encoder according to an embodiment of the present invention is described;
[0016] Figure 6 An example of reference picture resampling rate adaptation for video conferencing with varying throughput is shown;
[0017] Figure 7 An example of reference picture resampling rate adaptation for DASH and Open GOP resolution switching is shown;
[0018] Figure 8 An example of a third picture showing performing RoI zooming in a portion of a second picture;
[0019] Fig. 9 shows an example of a smoothing filter for adaptive motion vector resolution, where the motion vector difference is at half sample resolution, wherein the figure also shows additional interpolation filters collected in a table to be used for samples at phases other than half pixel phase; and
[0020] Fig.10 An example for signaling in the bitstream using syntax amvr_flag equal to 1 and amvr_precision_idx equal to 0 is shown. DETAILED DESCRIPTION
[0021] The following description of the drawings begins with the presentation of a description of an encoder and a decoder of a block-based prediction codec for decoding pictures of a video in order to form an example of a decoding framework into which embodiments of the present invention may be built. Figures 1 to 3 The corresponding encoders and decoders are described. Thereafter, a description of embodiments of the concepts of the invention is given together with information on how these concepts can be implemented into Figure 1 and Figure 2 The description of the encoder and decoder is presented together, although the embodiments described in the following Figure 4 and below can also be used to form a structure not based on Figure 1 and Figure 2 An encoder and decoder operating on a coding framework based on encoders and decoders of the present invention (such as intra-coded blocks without competing with inter-coded blocks in one picture, and / or such as no transform-based residual decoding, etc.).
[0022] Figure 1 An apparatus is shown for predictively decoding a picture 12 into a data stream 14, exemplarily using transform-based residual coding. The apparatus or encoder is denoted by reference numeral 10. Figure 2 A corresponding decoder 20 is shown, i.e. an apparatus 20 configured to predictively decode a picture 12' from a data stream 14 also using transform-based residual decoding, wherein a prime sign has been used to indicate that the picture 12' reconstructed by the decoder 20 deviates from the picture 12 originally encoded by the apparatus 10 with respect to decoding losses introduced by quantization of the prediction residual signal. Figure 1 and Figure 2 Transform-based prediction residual decoding is used illustratively, although embodiments of the present application are not limited to this type of prediction residual decoding. Figure 1 and Figure 2 The same is true for other details of the description, as will be outlined below.
[0023] The encoder 10 is configured to perform a spatial-to-spectral transform on a prediction residual signal and encode the prediction residual signal obtained thereby into a data stream 14. Likewise, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and perform a spectral-to-spatial transform on the prediction residual signal obtained thereby.
[0024] Internally, the encoder 10 may include a prediction residual signal former 22, which generates a prediction residual 24 in order to measure the deviation of the prediction signal 26 from the original signal, i.e. from the picture 12. The prediction residual signal former 22 may, for example, be a subtractor that subtracts the prediction signal from the original signal (i.e. from the picture 12). The encoder 10 then also includes a transformer 28, which performs a spatial to spectral transform on the prediction residual signal 24 to obtain a spectral domain prediction residual signal 24', which is then quantized by a quantizer 32 also included by the encoder 10. The quantized prediction residual signal 24" is thus decoded into the bitstream 14. To this end, the encoder 10 may optionally include an entropy decoder 34, which entropy decodes the prediction residual signal transformed and quantized into the data stream 14. The prediction signal 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24 encoded into the data stream 14 and decodable from the data stream 14. To this end, as Figure 1 As shown in , the prediction stage 36 may internally include a dequantizer 38, which dequantizes the prediction residual signal 24" to obtain a spectral domain prediction residual signal 24"' corresponding to the signal 24' excluding the quantization loss, followed by an inverse transformer 40, which inversely transforms the latter prediction residual signal 24"' (i.e., spectrum to spatial transformation) to obtain a prediction residual signal 24"" corresponding to the original prediction residual signal 24 excluding the quantization loss. Then, the combiner 42 of the prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24"" such as by addition to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 may correspond to the signal 12'. Then, the prediction module 44 of the prediction stage 36 generates the prediction signal 26 based on the signal 46 by using, for example, spatial prediction (i.e., intra-picture prediction) and / or temporal prediction (i.e., inter-picture prediction).
[0025] Likewise, if Figure 2 As shown in , the decoder 20 may be internally composed of components corresponding to the prediction stage 36 and interconnected in a manner corresponding to the prediction stage 36. In particular, the entropy decoder 50 of the decoder 20 may entropy decode the quantized spectral domain prediction residual signal 24" from the data stream, and then the dequantizer 52, the inverse transformer 54, the combiner 56 and the prediction module 58, which are interconnected and cooperate in the manner described above with respect to the modules of the prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24", so as to be Figure 2 As shown in , the output of combiner 56 results in a reconstructed signal, picture 12 ′.
[0026] Although not specifically described above, it is readily apparent that the encoder 10 may set some decoding parameters, including, for example, prediction modes, motion parameters, etc., according to some optimization schemes, such as, for example, in a manner that optimizes certain rate and distortion related criteria (i.e., decoding costs). For example, the encoder 10 and the decoder 20 and the corresponding modules 44, 58 may support different prediction modes, such as intra-coding mode and inter-coding mode, respectively. The granularity at which the encoder and the decoder switch between these prediction mode types may correspond to subdividing the pictures 12 and 12' into decoding segments or decoding blocks, respectively. For example, in units of these decoding segments, the pictures may be subdivided into blocks that are intra-coded and blocks that are inter-coded. As outlined in more detail below, intra-coded blocks are predicted based on the spatial, already decoded / decoded neighborhood of the corresponding blocks. There may be several intra-coding modes and selected for respective intra-coding segments including directional or angular intra-coding modes, according to which respective segments are filled by extrapolating sample values of a neighborhood into the respective intra-coding segment along a specific direction of the respective directional intra-coding mode. The intra-coding modes may, for example, also include one or more other modes, such as a DC coding mode according to which the prediction of the respective intra-coding block assigns a DC value to all samples within the respective intra-coding segment, and / or a plane intra-coding mode according to which the prediction of the respective block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function at the sample positions of the respective intra-coding block, with a driving tilt and offset of a plane defined by the two-dimensional linear function based on neighboring samples. In contrast, inter-coding blocks may, for example, be predicted in time. For an inter-coded block, a motion vector may be signaled within the data stream, the motion vector indicating a spatial displacement of a portion of a previously coded picture of the video to which the picture 12 belongs, at which portion the previously coded / decoded picture was sampled in order to obtain a prediction signal for the corresponding inter-coded block. This means that, in addition to the residual signal coding comprised by the data stream 14 (such as the entropy coded transform coefficient levels representing the quantized spectral domain prediction residual signal 24″), the data stream 14 may also have encoded therein: coding mode parameters for assigning a coding mode to individual blocks; prediction parameters for some of the blocks (such as motion parameters for inter-coded segments) and optionally other parameters (such as parameters for controlling and signaling the subdivision of the pictures 12 and 12′ respectively into segments). The decoder 20 uses these parameters to subdivide the picture in the same way as the encoder did, to assign the same prediction modes to the segments and to perform the same predictions resulting in the same prediction signals.
[0027] Figure 3The relationship between the reconstruction signal, ie the reconstructed picture 12', on the one hand, and the combination of the prediction residual signal 24"" and the prediction signal 26 as signaled in the data stream 14, on the other hand, is illustrated. As already indicated above, the combination may be an addition. The prediction signal 26 is Figure 3 1 is illustrated as subdividing the picture region into intra-coded blocks indicated illustratively using shading and inter-coded blocks indicated illustratively using non-shading. The subdivision may be any subdivision, such as a regular subdivision of the picture region into rows and columns of square blocks or non-square blocks, or a subdivision of the picture 12 from a multi-tree root block into a plurality of leaf blocks of different sizes, such as a quadtree subdivision, etc., wherein in Figure 3 Their hybrid is illustrated in , where the picture region is first subdivided into rows and columns of root blocks and then further subdivided into one or more leaf blocks according to a recursive multitree subdivision.
[0028] Likewise, for an intra-coded block 80, the data stream 14 may have an intra-coded mode decoded therein that assigns one of several supported intra-coded modes to the corresponding intra-coded block 80. For an inter-coded block 82, the data stream 14 may have one or more motion parameters encoded therein. In general, the inter-coded block 82 is not limited to being decoded in time. Alternatively, the inter-coded block 82 may be any block predicted from a previously decoded portion other than the current picture 12 itself, such as a previously decoded picture of the video to which the picture 12 belongs, or another view or hierarchically lower-level picture in the case where the encoder and decoder are scalable encoders and decoders, respectively.
[0029] Figure 3 The prediction residual signal 24″″ in is also illustrated as subdividing the picture area into blocks 84. These blocks may be referred to as transform blocks to distinguish them from the decoded blocks 80 and 82. In practice, Figure 3 It is illustrated that the encoder 10 and the decoder 20 may use two different subdivisions of the pictures 12 and 12', respectively, into blocks, namely one subdivision into coding blocks 80 and 82, respectively, and another subdivision into transform blocks 84, respectively. The two subdivisions may be identical, i.e., each decoding block 80 and 82 may simultaneously form a transform block 84, but Figure 3The case is described in which, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into decoding blocks 80, 82, such that any boundary between two blocks of blocks 80 and 82 overlaps a boundary between two blocks 84, or alternatively, each block 80, 82 either coincides with one of the transform blocks 84 or with a cluster of transform blocks 84. However, the subdivision can also be determined or selected independently of each other, so that the transform block 84 can alternatively cross a block boundary between blocks 80, 82. As far as the subdivision into transform blocks 84 is concerned, similar statements are true as those made about the subdivision into blocks 80, 82, i.e., the block 84 can be the result of a regular subdivision of a picture region into blocks (with or without arrangement into rows and columns), the result of a recursive multitree subdivision of a picture region, or a combination thereof or any other type of blockation. By the way, it is noted that the blocks 80, 82 and 84 are not limited to quadratic, rectangular or any other shape.
[0030] Figure 3 It is further illustrated that the combination of the prediction signal 26 and the prediction residual signal 24"" directly results in the reconstructed signal 12'. However, it should be noted that according to alternative embodiments more than one prediction signal 26 may be combined with the prediction residual signal 24"" to result in the picture 12'.
[0031] exist Figure 3 In the embodiment described below, the transform block 84 shall have the following meaning. The transformer 28 and the inverse transformer 54 perform their transforms in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping of transforms so as to decode the prediction residual signal directly in the spatial domain for some of the transform blocks 84. However, according to the embodiments described below, the encoder 10 and the decoder 20 are configured in such a way that they support several transforms. For example, the transforms supported by the encoder 10 and the decoder 20 may include:
[0032] o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
[0033] o DST-IV, where DST stands for Discrete Sine Transform
[0034] o DCT-IV
[0035] o DST-VII
[0036] oIdentity Transformation (IT)
[0037] Naturally, while the transformer 28 will support all forward transformed versions of these transforms, the decoder 20 or inverse transformer 54 will support their corresponding backward or inverse versions:
[0038] oInverse DCT-II (or Inverse DCT-III)
[0039] oInverse DST-IV
[0040] oInverse DCT-IV
[0041] oInverse DST-VIIo Identity Transformation (IT)
[0042] It should be noted that the set of supported transforms may include only one transform, such as a spectral to spatial or spatial to spectral transform.
[0043] As already outlined above, Figures 1 to 3 The examples have been presented in which the concepts of the present invention described further below may be implemented in order to form specific examples for encoders and decoders according to the present application. Figure 1 and 2 The encoder and decoder of may represent possible implementations of the encoder and decoder described below, respectively. However, Figure 1 and Figure 2 However, an encoder according to an embodiment of the present application may use the concepts outlined in more detail below to perform block-based encoding of picture 12 and Figure 1 The encoder differs from the one of FIG. 1 in that it does not support intra-frame prediction, or in that it uses a different Figure 3 Likewise, a decoder according to an embodiment of the present application may perform block-based decoding of a picture 12' from a data stream 14 using the decoding concepts further outlined below, but may be different from Figure 2 The decoder 20 of the embodiment of the present invention is different in that it does not support intra-frame prediction, or it uses a different method than that described in the embodiment of the present invention. Figure 3 The picture 12 ′ is divided into blocks in the manner described and / or the prediction residual is derived from the data stream 14 , for example not in the transform domain but in the spatial domain.
[0044] There are several applications that utilize resolution adaptation for several purposes, such as code rate adaptation for throughput variation or for Region of Interest (RoI) use cases.
[0045] The current VVC draft specifies a process commonly referred to as reference picture resampling, which allows for varying picture sizes within a video sequence during the RoI encoding process, such as Figures 6 to 8 To this end, the VVC draft specification includes the maximum picture size in the sequence parameter set (SPS), the actual picture size in the picture parameter set (PPS), and the scaling window offset in the PPS that allows the deriving of the scaling ratio to be used between the current picture and the reference picture (e.g. Figure 8 (red edge in the figure).
[0046] Having described possible implementations of encoder and decoder frameworks into which embodiments of the present application may be built, this specification first refers again to current VVC developments and provides motivation for the specifics of the embodiments outlined later.
[0047] In VVC, considering the expansion windows defined in PPSS for the current picture (PicOutputWidthL) and the reference picture (fRefWidth), the picture width is used to expand the ratio everywhere, as follows:
[0048] RefPicScale[i][j][0]=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL
[0049] RefPicScale[i][j][1]=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL
[0050] PicOutputWidth and PicOutputHeight are sometimes referred to as CurrPicScalWinWidth and CurrPicScalWinHeight below.
[0051] An expansion ratio <1, ie, a RefPicScale value <(1<<14), indicates that the current picture is larger than the reference picture, and an expansion ratio >1, ie, a RefPicScale value >(1<<14), indicates that the current picture is smaller than the reference picture.
[0052] The current VVC draft specifies 4 interpolation filters for motion compensation of up to 1 / 16 samples using fractional sample differences. The first interpolation filter was designed for conventional motion compensation when there is no reference picture RPR resampling and is not an affine mode. A second interpolation filter is designed for the case where an affine mode is used. The remaining two interpolation filters are used for downsampling with factors 1.5 and 2.
[0053] The expansion ratios allowed are from 1 / 8 (8x upsampling) to 2 (2x downsampling). Depending on whether affine mode is used and the expansion ratio, one of four filters is used. The conditions are as follows:
[0054] Use affine mode => for affine interpolation filter
[0055] ·Expansion ratio > 1.75 => use 2x downsampling interpolation filter
[0056] 1.25 < expansion ratio <= 1.75 => use 1.5x downsampling interpolation filter
[0057] · Spread ratio <= 1.25 => like conventional interpolation filter without RPR
[0058] For resolution changes where the current picture is larger than the reference picture, or for very small ratio values (expansion ratio <= 1.25x downsampling factor) when the current picture is smaller than the reference picture, a regular interpolation filter is used.
[0059] A conventional interpolation filter for a case where the affine mode is not used, or there is no RPR (extension ratio=1), or the extension ratio is less than or equal to 1.25 may apply a specific smoothing filter.
[0060] The 1 / 16 sampling regular interpolation filter is defined as an 8-tap filter in VVC. However, the VVC specification defines a special 6-tap smoothing filter for the following situations:
[0061] No intra-block copy (IBC) mode
[0062] • Signaling motion vector difference resolution at 1 / 2 luma sample resolution.
[0063] This 6-tap smoothing filter is used when adaptive motion vector resolution is used and the motion vector difference is at half sample resolution. Fig. 9 was copied in.
[0064] Given 1 / 16 fractional sampling precision in VVC, a fractional sampling position P=8 corresponds to a half sampling position (8 / 16=1 / 2). The variable hpelIfIdx equal to 1 indicates whether a 6-tap smoothing filter is used for the half sampling position (highlighted). When AmvrShift equals 3, hpelIfIdx is set to 1, indicating half-sampled MVD resolution if IBC mode is not used. This is signaled in the bitstream using the syntax amvr_flag equal to 1 and the syntax amvr_precision_idx equal to 0. See also Fig.10 .
[0065] When RPR is not used, in the above case, a smoothing filter is used to generate each sample in the reference block, because each sample in the block refers to the same fractional (half-sampled) interpolation position. However, when RPR is used, each sample may refer to a different fractional interpolation position.
[0066] Note in the text below that the n-sample difference (x″) in the current block L -x′ L or y″ L -y′ L ) is affected by the expansion ratio.
[0067] – For each luma sample position (x) in the predicted luma sample array predSamplesLX L =0..sbWidth-1+brdExtSize,y L =0..sbHeight-1+brdExtSize), the corresponding predicted brightness sample value predSamplesLX[x L ][y L ] is exported as follows:
[0068] –Assume (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position pointed to by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. Variable refxSb L ,refx L ,refySb L and refy L The export is as follows:
[0069] sbX L =(((xSb-scaling_win_left_offset)<<4)+refMvLX[0])·scalingRatio[0]
[0070] refx L =((Sign(refxSb L )·((Abs(refxSb L )+128)>>8)+x L ·((scalingRatio[0]+8)>>4))+fRefLeftOffset+32)>>6
[0071] sb L =(((ySb-scaling_win_top_offset)<<4)+refMvLX[1])·scalingRatio[1]
[0072] refy L =((Sign(refySb L)*((Abs(refySb L )+128)>>8)+yL*((scalingRatio[1]+8)>>4))+fRefTopOffset+32)>>6
[0073] Among them, scaling_win_left_offset can be calculated as SubWidthC x pps_scaling_win_left_offset, and scaling_win_top_offset can be calculated as SubHeightC x pps_scaling_win_top_offset.
[0074] For example, let's assume that the current image is 2 times larger than the reference image. Figure 4a , where samples of the reference picture and the current picture are shown spatially superimposed with crosses representing samples of the reference (d) picture and block circles representing samples of the current picture. The current 4x4 block of the current picture is shown. The motion vector of the block is indicated by an arrow. It is defined at half-pixel resolution relative to the pixel grid of the samples of the current picture. The position of the resulting block sample within the reference picture at which the reference picture is to be interpolated to produce a predictor for the block in the current picture is shown using white circles. The motion vector is exemplarily selected so that sample x' L = 0 (top left sample of 4x4 block) points to the half sample position, and sample x' L = 1 (samples to its right) will point to integer or full sample positions, both of which are indicated as half-pixels or full-pixels relative to the reference picture. Figure 4b As shown in , for reference pictures and current pictures with the same resolution and motion vectors defined at half-pixel resolution, this situation resulting in pixels being associated with different pixel positions (or in other words, different phases) does not occur. Similarly, if the expansion would be 4x instead of 2x, and sample x' L = 0 will point to the half-sampling position, then the sample x' L = 1 will point to the quarter sample position, and sample x' L = 2 will point to an integer sample location.
[0075] This means that within a single block, some samples will use the smoothing filter and some will not, which will lead to unpleasant visual effects and visible artifacts.
[0076] In one embodiment, the derivation of the variable hpelIfIdx in the motion compensation step is modified to include the extension ratio between the current picture and the reference picture, as follows:
[0077] If AmvrShift is equal to 3 and the expansion ratio==1, then hpelIfIdx=1, that is, RefPicScale is equal to 16384.
[0078] Figure 4c A video decoder according to the invention is shown. Video decoder 410 decodes video 450 from data stream 420 using motion compensated prediction. Motion prediction may be performed in motion prediction section 440 and is based on first motion vector 423 and second motion vector 424 transmitted in data stream 420.
[0079] The first motion vector 423 is transmitted in the data stream 420 at a first resolution of half the sampled resolution, and the second motion vector 423 is transmitted in the data stream 420 at a second resolution different from the first resolution.
[0080] Motion compensation is performed between a first picture 421 of equal picture resolution and a second picture 422 of different picture resolution, i.e., RPR is supported, or in other words, a motion vector can point from a current picture to a reference picture of the same resolution as the current picture, and the two then form the first picture, and a motion vector can point from a current picture to a reference picture of a different resolution than the current picture, and the two then form the second picture. Therefore, the picture size and / or resolution can vary and be signaled in the data stream. Motion compensation is performed using interpolation filters 431 and 432 to obtain subsample values within the reference picture (i.e., within the reference sample array).
[0081] The video decoder 410 selects an interpolation filter for a predetermined first motion vector from the first interpolation filter version 431 and the second interpolation filter version 432, for example, in the selection portion 430. The second interpolation filter version 432 has a higher edge preserving property than the first interpolation filter version 431. As will be shown in more detail below, the selection may be specific to samples of a specific phase, or differently stated, at a specific sub-pixel position such as a half-pixel position.
[0082] The selection of the interpolation filter depends on whether the current picture to which the predetermined first motion vector belongs is equal to the reference sample array associated with the predetermined first motion vector in picture resolution. The selection and the check for equality can be performed separately for the dimensions (i.e. the horizontal and / or vertical dimensions). Additionally or alternatively, the selection can also depend on constraint information 425 transmitted in the data stream, as will be outlined in more detail below.
[0083] exist Figure 4c The dependencies of constraint information 425 are not shown.
[0084] In addition, the encoder and the decoder can obtain the full sample values in the reference sample array for the predetermined first motion vector without using an interpolation filter. Figure 4a and 4b In , if the motion vector is transmitted in the data stream at a half-pixel resolution such as using AmvrShift=3, the motion vector shown therein is the "first motion vector". The positions to which the samples of the prediction block are shifted according to their respective motion vectors are shown by circles in Figures 4 and 4b. Those circles falling on the crosses are "full sample values". They can be determined directly from the collocated samples of the reference picture (crosses) without any interpolation. That is, the sample values of the samples of the reference picture on which the shifted positions of the samples of the inter-frame prediction block fall directly are directly used as predictors for the samples of the inter-frame prediction block whose shifted positions fall above them. Naturally, the same method can also be applied to the shifted sample positions of the inter-frame prediction blocks having a second motion vector (i.e., a motion vector transmitted in the data stream at a resolution other than half-pixel).
[0085] Furthermore, the decoder may obtain non-half-sampled subsample values by using another interpolation filter (eg, using a filter with a higher edge-preserving property than the first interpolation filter version). Figure 4a and 4b "Non-half-sampled subsample values" are those circles that neither fall on any reference picture sample nor are they placed between two horizontally, vertically or diagonally adjacent samples of the reference picture, i.e., they neither fall on any cross nor are they placed in the middle between two horizontally, vertically or diagonally adjacent crosses. For this purpose, an interpolation filter with higher edge-preserving properties is used. See Figure 4d , where it is shown for quarter-pixel positions: the shift position of the top and second left sample of the block is a quarter-pixel position. It is from the reference picture sample ( Figure 4d The interpolation filter to be used to determine the value of the shifted sample (i.e., the topmost and second white circle from the left) is therefore Fig. 9 The table includes the filter coefficients to be applied to the samples of the reference picture, and the sample positions to be interpolated are located between the samples of the reference picture. Figure 4dThe entry in which the interpolation filter is defined is highlighted, and some of the reference picture samples weighted according to the filter are illustrated to produce interpolated quarter-pixel samples. Note that horizontal interpolation can be applied first to obtain sample values at sub-pixel positions between samples of the reference picture, and then vertical interpolation can be performed using these interpolated intermediate samples to obtain the actual required sub-pixel samples if the required sub-pixel samples are offset from the reference picture samples both vertically and horizontally (or vice versa, i.e., vertically then horizontally) with sub-pixel accuracy. Similarly, the selection between the two half-pixel sample position interpolation filter versions can be done separately for the horizontal and vertical directions, or globally for both directions, relying on the equality of the picture resolution in both dimensions.
[0086] As described above, the selection can be performed separately for horizontal interpolation and vertical interpolation. Figure 4d In , the choice is explained at two entries regarding half-pixel position 8 / 16: which filter to use depends on hpelifidx. The setting of the latter variable depends, for example, on whether the reference picture and the current picture have the same resolution. The latter check for equality can be performed for a and y separately, as explained below by using the terms hpelHorIfIdx and hpelVerIfIdx. In particular, if the current picture and the reference sample array are not equal in horizontal picture resolution, a second interpolation filter (i.e., a filter with higher edge-preserving properties) can be selected for horizontal interpolation. This is the filter defined in the row where hpelifidx=0 in the table. Similarly, for example, if the current picture and the reference sample array are not equal in vertical picture resolution, a second interpolation filter (i.e., a filter with higher edge-preserving properties) can be selected for vertical interpolation. And, if the current picture and the reference sample array are not equal in horizontal and vertical picture resolution, a second interpolation filter (i.e., a filter with higher edge-preserving properties) can be selected for horizontal and vertical interpolation. For any direction where the second interpolation filter is not used, the first interpolation filter is used, ie, the one in the row of the table where hpelifidx=1.
[0087] The choice of which of the two half-pixel position interpolation filters to use can naturally be interpreted as being performed on all motion vectors, not just those that are half-pixel motion vectors. In this broad sense, the choice between the two also depends on whether the motion vector has half-sampled resolution. If so, the choice depends on the equality of resolution between the reference picture and the current picture, and if not, the choice is inevitably made using the second interpolation filter with higher edge-preserving properties.
[0088] As is clear from the above description, a decoder may use an alphabet of one or more syntax elements in a data stream in order to determine the resolution at which a certain motion vector is transmitted in the data stream. For example, an adaptive motion vector resolution may be indicated by amvr_flag, and if amvr_flag is set, thereby being consistent with a deviation from some default motion vector resolution, and an adaptive motion vector resolution precision may be indicated by an index amvr_precision_idx. This syntax is decoded by a decoder to derive the resolution at which a motion vector for a particular inter-prediction block is transmitted in a data stream, and the syntax is decoded accordingly by an encoder to indicate the resolution of the motion vector.
[0089] The decoder and encoder may exclude half-sampled resolution from the set of signalable settings of motion vector resolutions. They may map the alphabet of the one or more syntax elements onto a first set of vector resolutions that does not include half-sampled resolution if one of the following conditions is met (otherwise, mapping may be performed onto a second set of vector resolutions that includes half-sampled resolution):
[0090] • Constraint information indicates that a filter version with a lower edge preserving property is disabled for the current picture, which may be indicated in a picture or slice header, for example, and may be provided by a constraint equal to 1 for ph_disable_hpel_smoothing_filter or sh_disable_hpel_smoothing_filter.
[0091] The current picture to which the predetermined first motion vector belongs and the reference sample array associated with the predetermined first motion vector differ in at least one dimension of picture resolution.
[0092] The constraint information indicates that resampling of the reference sample array is enabled. This may be indicated, for example, on sequence level in the sequence parameter set SPS. An example of such an indication is sps_ref_picture_resample_enable_flag equal to 1. Thus, when resampling of the reference sample array is enabled, a filter version with a higher edge preserving property is used.
[0093] If none of the above conditions are met, the decoder will map the alphabet onto a second set of vector resolutions comprising half-sampled resolutions.
[0094] It is also noted that the data stream may include information whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
[0095] Furthermore, as mentioned above, the current picture may be equal in picture resolution to the reference sample array, in particular in the horizontal and vertical dimensions.
[0096] And the reference sample array can be a region, a sub-picture or a picture.
[0097] The decoder may also derive the constraint information from the data stream one of per picture sequence, per picture, or per slice.
[0098] Figure 5 A video encoder according to the present invention is described. The same principles apply to the decoder. In summary, the video encoder 510 encodes the video 550 into a data stream 520 using motion compensation prediction. The motion prediction can be performed in a motion prediction section 540. The encoder 510 indicates by transmitting a first motion vector 523 and a second motion vector 524 in the data stream 520.
[0099] The first motion vector 523 is transmitted in the data stream 520 at a first resolution of half the sampled resolution, and the second motion vector 523 is transmitted in the data stream 520 at a second resolution different from the first resolution.
[0100] Motion compensation is performed between a first picture 521 having an equal picture resolution and a second picture 522 having a different picture resolution using interpolation filters 531 and 532 to obtain subsample values within a reference picture (ie, within a reference array).
[0101] The video encoder 510 selects an interpolation filter for a predetermined first motion vector from the first interpolation filter version 531 and the second interpolation filter version 532 in the selection portion 530. The second interpolation filter version 532 has a higher edge preservation characteristic than the first interpolation filter version 531.
[0102] The selection of the interpolation filter depends on whether the current picture to which the predetermined first motion vector belongs is equal in horizontal and / or vertical dimension to the reference sample array associated with the predetermined first motion vector in picture resolution. Additionally or alternatively, the selection may also depend on constraint information 525 to be transmitted in the data stream.
[0103] As mentioned before, the same principles that can be embodied by a decoder can also be embodied by an encoder.
[0104] Therefore, the encoder can also obtain full sample values within the reference sample array for a predetermined first motion vector without using an interpolation filter.
[0105] Furthermore, the encoder may obtain non-half-sampled sub-sampled values by using another interpolation filter (eg, using a filter with a higher edge preserving property than a version of the first interpolation filter).
[0106] As described above, the selection may be performed separately for horizontal interpolation and vertical interpolation.
[0107] In particular, if the current picture and the reference sample array are not equal in horizontal picture resolution, the second interpolation filter (ie, a filter with a higher edge-preserving property) may be selected for horizontal interpolation.
[0108] Likewise, for example, if the current picture and the reference sample array are not equal in vertical picture resolution, a second interpolation filter (ie, a filter with a higher edge-preserving property) may be selected for vertical interpolation.
[0109] Furthermore, if the current picture and the reference sample array are not equal in horizontal and vertical picture resolution, a second interpolation filter (ie, a filter with a higher edge-preserving property) may be selected for horizontal and vertical interpolation.
[0110] The selection may also be performed in dependence on whether the predetermined first motion vector has a half-sample resolution or not.
[0111] Furthermore, for selecting the resolution of motion vectors, the encoder may refrain from using half-sample resolution for one or more vectors if the current picture is equal in picture resolution to the reference sample array in the horizontal and / or vertical dimensions.
[0112] To select, the encoder may map an alphabet of one or more syntax elements in the data stream that indicate a resolution of a predetermined first motion vector. For example, the adaptive motion vector resolution may be indicated by amvr_flag, the adaptive motion vector resolution precision may be indicated by amvr_precision_idx, and the encoder may map the alphabet to a first set of vector resolutions that does not include half-sampled resolutions if one of the following conditions is met:
[0113] • Constraint information indicates that a filter version with a lower edge preserving property is disabled for the current picture, which may be indicated in a picture or slice header, for example, and may be provided by a constraint equal to 1 for ph_disable_hpel_smoothing_filter or sh_disable_hpel_smoothing_filter.
[0114] The current picture to which the predetermined first motion vector belongs and the reference sample array associated with the predetermined first motion vector differ in at least one dimension of picture resolution.
[0115] The constraint information indicates that resampling of the reference sample array is enabled. This may be indicated, for example, on sequence level in the sequence parameter set SPS. An example of such an indication is sps_ref_picture_resample_enable_flag equal to 1. Thus, when resampling of the reference sample array is enabled, a filter version with a higher edge preserving property is used.
[0116] If none of the above conditions are met, the encoder may map the alphabet to a second set of vector resolutions including half the sampled resolution.
[0117] It is also noted that the data stream may include information whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
[0118] Furthermore, as mentioned above, the current picture may be equal in picture resolution to the reference sample array, in particular in the horizontal and vertical dimensions.
[0119] And the reference sample array can be a region, a sub-picture or a picture.
[0120] The encoder may also derive constraint information from the data stream one of per picture sequence, per picture, or per slice.
[0121] Finally, when a program having software code portions for employing the above principles is run on a processing device, the above principles can also be implemented with a computer program product comprising the program. In addition, the computer program product can also be embodied in a computer-readable medium on which software code portions are stored.
[0122] The principles listed above and below may also be embodied as data streams generated by an encoder or by an encoder, as described in this document.
[0123] Let us return to the description of the embodiment for modifying the current VVC draft. Figure 5 The dependencies of constraint information are not shown.
[0124] In one embodiment, the derivation of the variable hpelIfIdx in the motion compensation step is modified to incorporate an enable flag for reference picture signaling at the sequence level in the SPS, and the smoothing filter coefficients are used only when reference picture resampling is disabled, as follows:
[0125] If AmvrShift is equal to 3 and sps_ref_picture_resample_enable_flag==0, then hpelIfIdx=1, ie, reference picture resampling is disabled.
[0126] In another embodiment, a control syntax flag is added to the picture or slice header to indicate whether the smoothing filter is disabled for the current picture. hpelIfIdx is then derived as follows:
[0127] If AmvrShift is equal to 3 and the control flag is equal to 0 (eg, ph_disable_hpel_smoothing_filter or sh_disable_hpel_smoothing_filter), then hpelIfIdx=1.
[0128] In another embodiment, the derivation of the variable AmvrShift is modified to include information about the reference picture resampling, such as and an avoid value equal to 3 when:
[0129] ph_disable_hpel_smoothing_filter or sh_disable_hpel_smoothing_filter is equal to 1, or
[0130] scaling ratio != 1, i.e., RefPicScale is not equal to 16384, or
[0131] ·sps_ref_picture_resample_enable_flag==1
[0132] In another embodiment, when RPR is used for reference pictures, that is, if the sizes of the current picture and the reference picture are not equal, or the expansion ratio derived from the expansion window is not equal to 1, that is, RefPicScale is not equal to 16384, then one of the bitstream constraints is that AmvrShift is not equal to 3.
[0133] Therefore, the horizontal and vertical half-sampling interpolation filter indices hpelHorIfIdx and hpelVerIfIdx are derived as follows:
[0134] hpelHorIfIdx=(scalingRatio[0]==16384)? hpelIfIdx:0
[0135] hpelVerIfIdx=(scalingRatio[1]==16384)? hpelIfIdx:0
[0136] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent the description of the corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, the aspects described in the context of a method step also represent the description of the corresponding block or item or feature of the corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0137] The data stream of the present invention may be stored on a digital storage medium, or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0138] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium (e.g., a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon that cooperate (or can cooperate with) a programmable computer system to cause the corresponding method to be performed. Thus, the digital storage medium can be computer readable.
[0139] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0140] Generally speaking, the embodiments of the present invention can be realized as a computer program product with a program code, the program code being operative so as to perform one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine-readable carrier.
[0141] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0142] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0143] A further embodiment of the method of the invention is therefore a data carrier (or a digital storage medium or a computer-readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.
[0144] Therefore, another embodiment of the method of the present invention is a data stream or a signal sequence, which represents a computer program for implementing one of the methods described herein. The data stream or signal sequence can be configured, for example, to transfer via a data communication connection (e.g., via the Internet).
[0145] Further embodiments comprise a processing means, for example a computer or a programmable logic device, configured to or adapted to carry out one of the methods described herein.
[0146] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0147] Further embodiments according to the invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transferring the computer program to the receiver.
[0148] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to implement some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to implement one of the methods described herein. In general, the method is preferably implemented by any hardware device.
[0149] The devices described herein may be implemented using hardware devices or using computers or using a combination of hardware devices and computers.
[0150] The apparatus described herein or any component of an apparatus described herein may be implemented at least partially in hardware and / or software.
[0151] The methods described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0152] Any component of the methods described herein or the apparatus described herein may be performed at least in part by hardware and / or software.
[0153] The above embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to other persons skilled in the art. Therefore, the present invention is intended to be limited only by the scope of the patent claims to be filed, and not by the specific details presented in the manner of describing and explaining the embodiments herein.
Claims
1. A video decoder (410) comprising a processor, the processor being configured to decoding video (450) from the data stream (420) using motion compensated prediction to predict a current picture from a reference picture based on motion vector differences (423) at half sample resolution transmitted in the data stream (420), wherein the first reference picture (421) has the same picture resolution as each current picture predicted using the first reference picture, and the second reference picture (422) has a different picture resolution than each current picture predicted using the second reference picture, and Using interpolation filters (431, 432) to obtain sub-sample values within a reference sample array formed by each reference picture, An interpolation filter is selected to obtain half sample values within a reference sample array formed from a reference picture using a motion vector, the interpolation filter being selected from a first interpolation filter version (431) and a second interpolation filter version (432), wherein the first interpolation filter version (431) is a 6-tap filter with coefficients 3, 9, 20, 20, 9, 3 and the second interpolation filter version (432) is an 8-tap filter with coefficients -1, 4, -11, 40, 40, -11, 4, -1 in the following manner: When the picture resolution of the reference picture in the horizontal dimension is equal to the current picture, selecting the first interpolation filter version, and when the picture resolution of the reference picture in the horizontal dimension is not equal to the current picture, selecting the second interpolation filter version; and / or When the picture resolution of the reference picture in the vertical dimension is equal to the current picture, the first interpolation filter version is selected, and when the picture resolution of the reference picture in the vertical dimension is not equal to the current picture, the second interpolation filter version is selected.
2. The video decoder according to claim 1, further configured to Full sample values within the reference sample array are obtained without using the interpolation filter.
3. The video decoder according to claim 1, further configured to Non-half-sampled sub-sample values within the reference sample array are obtained by using a further interpolation filter version having a higher edge preserving property than the first interpolation filter version.
4. The video decoder according to claim 1, further configured to The selection is performed separately for horizontal interpolation and vertical interpolation.
5. The video decoder according to claim 1, in: If the current picture is not equal in picture resolution to the reference picture in a horizontal dimension, selecting the second interpolation filter version for horizontal interpolation, If the current picture is not equal in picture resolution to the reference picture in the vertical dimension, selecting the second interpolation filter version for vertical interpolation, and If the current picture is not equal in picture resolution to the reference picture in horizontal and vertical dimensions, the second interpolation filter version is selected for horizontal and vertical interpolation.
6. The video decoder of claim 1, wherein the second interpolation filter version has a higher edge preserving property than the first interpolation filter version.
7. The video decoder according to claim 1, in, The data stream comprises information indicating whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
8. The video decoder according to claim 1, in, The reference sample array is a region, a sub-picture or a picture.
9. A video encoder (510) comprising a processor, the processor being configured to encoding the video (550) into a data stream (520) using motion compensated prediction to predict a current picture from a reference picture, indicating the motion vector difference (523) at half sample resolution by transmission in the data stream (520), wherein the first reference picture (521) has the same picture resolution as each current picture predicted using the first reference picture, and the second reference picture (522) has a different picture resolution than each current picture predicted using the second reference picture, and Using interpolation filters (531, 532) to obtain sub-sample values within a reference sample array formed by each reference picture, An interpolation filter is selected to obtain half sample values within a reference sample array formed from a reference picture using a motion vector, the interpolation filter being selected from a first interpolation filter version (531) and a second interpolation filter version (532), wherein the first interpolation filter version (531) is a 6-tap filter with coefficients 3, 9, 20, 20, 9, 3 and the second interpolation filter version (532) is an 8-tap filter with coefficients -1, 4, -11, 40, 40, -11, 4, -1 in the following manner: When the picture resolution of the reference picture in the horizontal dimension is equal to the current picture, selecting the first interpolation filter version, and when the picture resolution of the reference picture in the horizontal dimension is not equal to the current picture, selecting the second interpolation filter version; and / or When the picture resolution of the reference picture in the vertical dimension is equal to the current picture, the first interpolation filter version is selected, and when the picture resolution of the reference picture in the vertical dimension is not equal to the current picture, the second interpolation filter version is selected.
10. The video encoder according to claim 9, further configured to Full sample values within the reference sample array are obtained without using the interpolation filter.
11. The video encoder according to claim 9, further configured to Non-half-sampled sub-sample values within the reference sample array are obtained by using a further interpolation filter version having a higher edge preserving property than the first interpolation filter version.
12. The video encoder according to claim 9, further configured to The selection is performed separately for horizontal interpolation and vertical interpolation.
13. The video encoder of claim 9, wherein If the current picture is not equal in picture resolution to the reference picture in a horizontal dimension, selecting the second interpolation filter version for horizontal interpolation, If the current picture is not equal in picture resolution to the reference picture in the vertical dimension, selecting the second interpolation filter version for vertical interpolation, and If the current picture is not equal in picture resolution to the reference picture in horizontal and vertical dimensions, the second interpolation filter version is selected for horizontal and vertical interpolation.
14. The video encoder of claim 9, wherein the second interpolation filter version has a higher edge preserving property than the first interpolation filter version.
15. The video encoder according to claim 9, in, The data stream comprises information indicating whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
16. The video encoder according to claim 9, in, The reference sample array is a region, a sub-picture or a picture.
17. A method for decoding a video, comprising: Decoding video from the data stream using motion compensated prediction to predict the current picture from a reference picture based on motion vector differences at half sample resolution transmitted in the data stream, wherein a first reference picture has the same picture resolution as each current picture predicted by the first reference picture, and a second reference picture has a different picture resolution than each current picture predicted by the second reference picture, and an interpolation filter is used to obtain subsample values within a reference sample array formed by each reference picture, An interpolation filter is selected to obtain half sample values within a reference sample array formed from a reference picture using a motion vector, the interpolation filter being selected from a first interpolation filter version and a second interpolation filter version, wherein the first interpolation filter version is a 6-tap filter with coefficients 3, 9, 20, 20, 9, 3 and the second interpolation filter version is an 8-tap filter with coefficients -1, 4, -11, 40, 40, -11, 4, -1 in the following manner: When the picture resolution of the reference picture in the horizontal dimension is equal to the current picture, selecting the first interpolation filter version, and when the picture resolution of the reference picture in the horizontal dimension is not equal to the current picture, selecting the second interpolation filter version; and / or When the picture resolution of the reference picture in the vertical dimension is equal to the current picture, the first interpolation filter version is selected, and when the picture resolution of the reference picture in the vertical dimension is not equal to the current picture, the second interpolation filter version is selected.
18. The method for decoding a video according to claim 17, further comprising Full sample values within the reference sample array are obtained without using the interpolation filter.
19. The method for decoding a video according to claim 17, further comprising The non-half-sampled sub-sample values are obtained by using a further interpolation filter having a higher edge preserving property than the first interpolation filter version.
20. The method for decoding a video according to any one of claims 17 to 19, further comprising: The selection is performed separately for horizontal interpolation and vertical interpolation.
21. The method for decoding a video according to claim 17 in, If the current picture is not equal to the reference picture in the horizontal dimension in picture resolution, selecting the second interpolation filter for horizontal interpolation, wherein if the current picture is not equal to the reference picture in the vertical dimension in picture resolution, selecting the second interpolation filter for vertical interpolation, and If the current picture is not equal in picture resolution to the reference picture in horizontal and vertical dimensions, the second interpolation filter is selected for horizontal and vertical interpolation.
22. The method for decoding video of claim 17, wherein the second interpolation filter version has a higher edge preserving property than the first interpolation filter version.
23. The method for decoding a video according to claim 17 The data stream includes information indicating whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
24. The method for decoding a video according to claim 17 in, The reference sample array is a region, a sub-picture or a picture.
25. Methods for encoding video, include: Encode video into a data stream using motion compensated prediction to predict the current picture from a reference picture. Indicates the motion vector differences at half sample resolution by transmitting them in the data stream, wherein a first reference picture has the same picture resolution as each current picture predicted by the first reference picture, and a second reference picture has a different picture resolution than each current picture predicted by the second reference picture, and an interpolation filter is used to obtain subsample values within a reference sample array formed by each reference picture, An interpolation filter is selected to obtain half sample values within a reference sample array formed from a reference picture using a motion vector, the interpolation filter being selected from a first interpolation filter version and a second interpolation filter version, wherein the first interpolation filter version is a 6-tap filter with coefficients 3, 9, 20, 20, 9, 3 and the second interpolation filter version is an 8-tap filter with coefficients -1, 4, -11, 40, 40, -11, 4, -1 in the following manner: When the picture resolution of the reference picture in the horizontal dimension is equal to the current picture, selecting the first interpolation filter version, and when the picture resolution of the reference picture in the horizontal dimension is not equal to the current picture, selecting the second interpolation filter version; and / or When the picture resolution of the reference picture in the vertical dimension is equal to the current picture, the first interpolation filter version is selected, and when the picture resolution of the reference picture in the vertical dimension is not equal to the current picture, the second interpolation filter version is selected.
26. The method for encoding video according to claim 25, further comprising Full sample values within the reference sample array are obtained without using the interpolation filter.
27. The method for encoding video according to claim 25, further comprising The non-half-sampled sub-sample values are obtained by using a further interpolation filter having a higher edge preserving property than the first interpolation filter version.
28. The method for encoding video according to claim 25, further comprising The selection is performed separately for horizontal interpolation and vertical interpolation.
29. The method for encoding a video according to claim 25, in: If the current picture is not equal to the reference picture in the horizontal dimension in picture resolution, selecting the second interpolation filter for horizontal interpolation, If the current picture is not equal in picture resolution to the reference picture in a vertical dimension, selecting the second interpolation filter for vertical interpolation, and If the current picture is not equal in picture resolution to the reference picture in horizontal and vertical dimensions, the second interpolation filter is selected for horizontal and vertical interpolation.
30. The method for encoding video of claim 25, wherein the second interpolation filter version has a higher edge preserving property than the first interpolation filter version.
31. The method for encoding a video according to claim 25, The data stream includes information indicating whether temporally consecutive pictures have the same or different horizontal and / or vertical picture resolution dimensions.
32. The method for encoding a video according to claim 25, in, The reference sample array is a region, a sub-picture or a picture.
33. A computer program product comprising a computer program having software code portions for performing the steps of the method according to any one of claims 17 to 32 when the computer program is run on a processing device.
34. A computer readable medium having stored thereon a computer program comprising software code portions configured to perform the steps of the method according to any one of claims 17 to 32 when the computer program is run on a processing device.
Citation Information
Patent Citations
Method and apparatus for video coding with automatic motion information refinement
CN109417630A
Video coding with adaptive motion information refinement
CN109417631A