Method and apparatus for resampling reference images

By using a new low-pass interpolation filter and disabling RPR inter-frame prediction in VVC, the encoding/decoding efficiency and computational complexity issues in affine mode of VVC are addressed, thereby improving encoding/decoding performance and quality.

CN118890466BActive Publication Date: 2026-01-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410800824.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-24
Filing Date
2020-12-24
Publication Date
2026-01-06
Estimated Expiration
2040-12-24

AI Technical Summary

Technical Problem

In the existing video codec standard VVC, the reference image resampling design suffers from low encoding and decoding efficiency, high memory bandwidth, and high computational complexity in affine mode. This is especially true when the resolution of the reference image is higher than that of the current image, leading to aliasing artifacts and increased computational burden.

Method used

A new low-pass interpolation filter replaces the default 6-tap and 4-tap interpolation filters, downsampling is performed for affine modes, and RPR inter-frame prediction is disabled at certain CU sizes, while allowing dynamic changes in bit depth within the video sequence.

Benefits of technology

It improves the encoding and decoding efficiency of affine mode, reduces memory bandwidth and computational complexity, reduces aliasing artifacts, and achieves more flexible encoding and decoding performance and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118890466B_ABST
    Figure CN118890466B_ABST
Patent Text Reader

Abstract

Methods, apparatuses, and non-transitory computer-readable storage media for decoding a video signal are provided. A decoder obtains a reference picture I associated with a video block within the video signal. The decoder can further obtain reference samples I(i,j) of the video block from a reference block in the reference picture I. The decoder can also obtain a first down-sampling filter and a second down-sampling filter to generate luma and chroma inter prediction samples, respectively, of the video block. When the video block is coded by an affine mode, the decoder can further obtain a third down-sampling filter and a fourth down-sampling filter to generate luma and chroma inter prediction samples, respectively, of the video block. The decoder can also obtain inter prediction samples of the video block based on the third and fourth down-sampling filters being applied to the reference samples I(i,j).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to provisional application No. 62 / 953,471, filed on December 24, 2019, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] This disclosure relates to video encoding and compression. More specifically, this disclosure relates to methods and apparatus for reference image resampling techniques used in video encoding and decoding. Background Technology

[0004] Various video codec techniques can be used to compress video data. Video codecs are performed according to one or more video codec standards. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (H.265 / HEVC), High-Level Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically use prediction methods that utilize redundancy present in video images or sequences (e.g., inter-frame prediction, intra-frame prediction, etc.). A key goal of video codec techniques is to compress video data to a lower bitrate while avoiding or minimizing video quality degradation. Summary of the Invention

[0005] Examples of this disclosure provide methods and apparatus for resampling reference images.

[0006] According to a first aspect of this disclosure, a method for decoding a video signal is provided. The method may include a decoder obtaining a reference image I associated with a video block within the video signal. The decoder may also obtain reference samples I(i,j) of the video block from the reference block in the reference image I. i and j may represent the coordinates of a sample within the video block. When the video block is encoded and decoded in a non-affine inter-frame mode and the resolution of the reference image I is greater than the resolution of the current image, the decoder may further obtain a first downsampling filter and a second downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. When the video block is encoded and decoded in an affine mode and the resolution of the reference image is greater than the resolution of the current image, the decoder may also obtain a third downsampling filter and a fourth downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. The decoder may further obtain inter-frame prediction samples of the video block based on the application of the third and fourth downsampling filters to the reference sample I(i,j).

[0007] According to a second aspect of this disclosure, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable storage medium storing instructions executable by the one or more processors. The one or more processors may be configured to obtain a reference picture I associated with a video block within a video signal. The one or more processors may also be configured to obtain a reference sample I(i,j) of the video block from the reference block in the reference picture I. i and j may represent the coordinates of a sample within the video block. The one or more processors may be further configured to: when the video block is encoded / decoded in a non-affine inter-frame mode and the resolution of the reference picture I is greater than the resolution of the current picture, obtain a first downsampling filter and a second downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. The one or more processors may also be configured to: when the video block is encoded / decoded in an affine mode and the resolution of the reference picture is greater than the resolution of the current picture, obtain a third downsampling filter and a fourth downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. The one or more processors may be further configured to obtain inter-frame prediction samples of the video block based on the application of the third and fourth downsampling filters to the reference sample I(i,j).

[0008] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided in which instructions are stored. When executed by one or more processors of the apparatus, the instructions cause the apparatus to obtain a reference picture I associated with a video block within a video signal. The instructions also cause the apparatus to obtain a reference sample I(i,j) of the video block from a reference block in the reference picture I. i and j may represent the coordinates of a sample within the video block. The instructions further cause the apparatus to: when the video block is encoded / decoded in a non-affine inter-frame mode and the resolution of the reference picture I is greater than the resolution of the current picture, obtain a first downsampling filter and a second downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. The instructions further cause the apparatus to: when the video block is encoded / decoded in an affine mode and the resolution of the reference picture is greater than the resolution of the current picture, obtain a third downsampling filter and a fourth downsampling filter to generate luma and chroma inter-frame prediction samples of the video block, respectively. The instructions further cause the apparatus to obtain inter-frame prediction samples of the video block based on the application of the third and fourth downsampling filters to the reference sample I(i,j).

[0009] It should be understood that the foregoing general description and the following detailed description are merely examples and do not limit this disclosure. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate examples consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. Similar reference numerals denote corresponding parts.

[0011] Figure 1 A block diagram of an encoder according to an example of this disclosure.

[0012] Figure 2 A block diagram of an example decoder according to this disclosure.

[0013] Figure 3A This is a schematic diagram illustrating a block partition of a multi-type tree structure according to an example of this disclosure.

[0014] Figure 3B This is a schematic diagram illustrating a block partition of a multi-type tree structure according to an example of this disclosure.

[0015] Figure 3C This is a schematic diagram illustrating a block partition of a multi-type tree structure according to an example of this disclosure.

[0016] Figure 3D This is a schematic diagram illustrating a block partition of a multi-type tree structure according to an example of this disclosure.

[0017] Figure 3E This is a schematic diagram illustrating a block partition of a multi-type tree structure according to an example of this disclosure.

[0018] Figure 4A This is a schematic diagram of a 4-parameter affine model according to an example of this disclosure.

[0019] Figure 4B This is a schematic diagram of a 4-parameter affine model according to an example of this disclosure.

[0020] Figure 5 This is a schematic diagram of a 6-parameter affine model according to an example of this disclosure.

[0021] Figure 6 This is a schematic diagram illustrating an example of adaptive bit depth switching according to this disclosure.

[0022] Figure 7 This is an example of a method for decoding a video signal in accordance with this disclosure.

[0023] Figure 8 This is an example of a method for decoding a video signal in accordance with this disclosure.

[0024] Figure 9 This is a schematic diagram illustrating a computing environment coupled with a user interface, as an example of this disclosure. Detailed Implementation

[0025] Reference will now be made in detail to exemplary implementations illustrated in the accompanying drawings. The following description refers to the drawings, wherein, unless otherwise indicated, the same numbers in the different drawings denote the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with various aspects recounted in the appended claims and relating to this disclosure.

[0026] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. When used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the / said” are contemplated to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein is contemplated to represent and include any or all possible combinations of one or more of the associated listed items.

[0027] It should be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various types of information, such information should not be limited by these terms. These terms are merely used to distinguish one type of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information; and similarly, second information may be referred to as first information. As used herein, the term “if” may be understood, depending on the context, to mean “when,” “in the event of,” or “in response to a judgment.”

[0028] The first version of the HEVC standard was finalized in October 2013, offering approximately 50% bitrate savings or equivalent perceived quality compared to the previous generation video codec standard, H.264 / MPEG-AVC. While HEVC offered significant codec improvements over its predecessors, evidence suggested that even greater codec efficiency could be achieved using additional codec tools. Based on this, both VCEG and MPEG initiated exploration of new codec technologies for future video codec standardization. A Joint Video Exploration Team (JVET) was established in October 2015 by ITU-T VECG and ISO / IEC MPEG to begin significant research into advanced technologies capable of achieving substantial improvements in codec efficiency. JVET advocated for a reference software called the Joint Exploration Model (JEM) by integrating several additional codec tools on top of the HEVC test model (HM).

[0029] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) for video compression capabilities exceeding HEVC. In April 2018, at the 10th JVET meeting, 23 CfP responses were received and evaluated, demonstrating a compression efficiency gain of approximately 40% compared to HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard, named Universal Video Codec (VVC). That same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0030] Like HEVC, VVC is built on a block-based hybrid video codec framework.

[0031] Figure 1 A general schematic diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1 A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra-frame prediction 118, image buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, loop filter 122, entropy coding 138, and bitstream 144.

[0032] In encoder 100, video frames are divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.

[0033] The prediction residual, representing the difference between the current video block—a portion of video input 110—and its predictor—a portion of block predictor 140—is sent from adder 128 to transform 130. Then, the transform coefficients are sent from transform 130 to quantization 132 for entropy reduction. Next, the quantized coefficients are fed to entropy coding 138 to generate a compressed video bitstream. Figure 1 As shown, prediction-related information 142 from the intra / inter-frame mode decision 116, such as video block partitioning information, motion vectors (MV), reference picture indexes, and intra-frame prediction modes, is also fed through entropy coding 138 and saved into the compressed bitstream 144. The compressed bitstream 144 includes the video bitstream.

[0034] In encoder 100, a decoder-related circuitry is also required for pixel reconstruction for prediction purposes. First, the prediction residual is reconstructed via inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is then combined with block predictor 140 to generate unfiltered reconstructed pixels for the current video block.

[0035] Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video frame as the current video block to predict the current video block.

[0036] Timing prediction (also known as "inter-frame prediction") uses reconstructed pixels from already encoded and decoded video frames to predict the current video block. Timing prediction reduces the inherent temporal redundancy in the video signal. The timing prediction signal for a given codec unit (CU) or codec block is typically emitted via one or more MVs, which indicate the amount and direction of motion between the current CU and its timing reference. Furthermore, if multiple reference frames are supported, an additional reference frame index is sent, identifying which reference frame in the reference frame repository the timing prediction signal originates from.

[0037] Motion estimation 114 takes in the video input 110 and the signal from the image buffer 120, and outputs the motion estimation signal to motion compensation 112. Motion compensation 112 takes in the video input 110, the signal from the image buffer 120, and the motion estimation signal from motion estimation 114, and outputs the motion compensation signal to intra / inter-frame mode decision 116.

[0038] After performing spatial and / or temporal prediction, the intra / inter-frame mode decision 116 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Next, block predictor 140 is subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantization residual coefficients are inversely quantized using inverse quantization 134 and inversely transformed using inverse transform 136 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed CU signal. Further, before the reconstructed CU is placed in the reference image repository of image buffer 120 and used to encode and decode future video blocks, loop filtering 122, such as deblocking filters, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), can be applied to it. To form the output video bitstream 144, the encoding / decoding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are all sent to entropy coding unit 138 for further compression and packing to form the bitstream.

[0039] Figure 1A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called CU). In VVC, CUs can be up to 128x128 pixels. However, unlike HEVC, which is based solely on quadtree-based block partitioning, in VVC, a codec tree unit (CTU) is split into several CUs based on quadtrees / binaries / tritrees to accommodate varying local characteristics. Furthermore, the concept of multi-partition unit types in HEVC is removed; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transform, without further partitioning. In the multi-type tree structure, a CTU is first partitioned by a quadtree structure. Then, each quadtree leaf node can be further partitioned by binary and ternary tree structures.

[0040] like Figure 3A , Figure 3B , Figure 3C , Figure 3D and Figure 3E As shown, there are five types of splitting: quadruple splitting, horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

[0041] Figure 3A A schematic diagram illustrating block quadrilateral partitioning in a multi-type tree structure is shown in accordance with this disclosure.

[0042] Figure 3B A schematic diagram illustrating block vertical binary partitioning in a multi-type tree structure is shown in accordance with this disclosure.

[0043] Figure 3C A schematic diagram illustrating block-level binary partitioning in a multi-type tree structure is shown in accordance with this disclosure.

[0044] Figure 3D A schematic diagram illustrating block vertical ternary partitioning in a multi-type tree structure is shown in accordance with this disclosure.

[0045] Figure 3E A schematic diagram illustrating block-level ternary partitioning in a multi-type tree structure is shown in accordance with this disclosure.

[0046] exist Figure 1In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture or strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more MVs, which indicate the amount and direction of motion between the current CU and its temporal reference. Furthermore, if multiple reference pictures are supported, a reference picture index is additionally sent to identify which reference picture in the reference picture repository the temporal prediction signal comes from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Then, the prediction block is subtracted from the current video block; and the prediction residual is decorrelated using transform and quantization. The quantized residual coefficients are inversely quantized and inversely transformed to form the reconstructed residuals, which are then added back to the prediction block to form the reconstructed CU signal. Furthermore, before placing the reconstructed CU in the reference image repository and using it for encoding and decoding future video blocks, loop filters such as deblocking filters, SAO, and ALF can be applied to it. To form the output video bitstream, the encoding / decoding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit for further compression and packing to form the bitstream.

[0047] Figure 2 A general block diagram of a video decoder for VVC is shown. Specifically, Figure 2 A block diagram of a typical decoder 200 is shown. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter-frame mode selection 220, intra-frame prediction 222, memory 230, loop filter 228, motion compensation 224, image buffer 226, prediction-related information 234, and video output 232.

[0048] Decoder 200 is similar to residing in Figure 1The reconstruction-related part is located in the encoder 100. In the decoder 200, the incoming video bitstream 210 is first decoded by entropy decoding 212 to derive the quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by inverse quantization 214 and inverse transform 216 to obtain the reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter-frame mode selector 220 is configured to perform intra-frame prediction 222 or motion compensation 224 based on the decoded prediction information. The set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residuals from the inverse transform 216 and the prediction output generated by the block predictor mechanism using a summer 218.

[0049] The reconstructed blocks can be further passed through a loop filter 228 before being stored in an image buffer 226, which serves as a reference image storage. The reconstructed video in the image buffer 226 can be sent to drive a display device, as well as to predict future video blocks. In the case where the loop filter 228 is open, filtering operations are performed on these reconstructed pixels to produce the final reconstructed video output 232.

[0050] Figure 2 A general block diagram of a block-based video decoder is presented. First, the video bitstream is entropy-decoded at the entropy decoding unit. The encoding / decoding mode and prediction information are sent to the spatial prediction unit (if intra-frame encoding / decoding is performed) or the temporal prediction unit (if inter-frame encoding / decoding is performed) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. Next, the prediction blocks and the residual blocks are added together. The reconstructed blocks may be further subjected to loop filtering before being stored in the reference image repository. Then, the reconstructed video in the reference image repository is sent to drive the display device, along with the video blocks used to predict future blocks.

[0051] This disclosure focuses on improving and simplifying existing reference image resampling designs supported in VVC. Below, we briefly review current codec tools in VVC that are closely related to the techniques presented in this disclosure.

[0052] Affine mode

[0053] In HEVC, only the translational motion model is applied to motion compensation prediction. However, in the real world, many types of motion exist, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, affine motion compensation prediction is applied by emitting a flag for each inter-frame codec block to indicate whether a translational motion or an affine motion model is applied to the inter-frame prediction. In the current VVC design, for an affine codec block, two affine modes are supported: a 4-parameter affine mode and a 6-parameter affine mode.

[0054] The 4-parameter affine model has the following parameters: two parameters for translational motion in the horizontal and vertical directions, one parameter for scaling motion in these two directions, and one parameter for rotational motion. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To achieve more efficient affine parameter signaling, in VVC, those affine parameters are derived from two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block.

[0055] like Figure 4A and Figure 4B As shown, the affine motion field of the block is described by two control points MV(V0,V1).

[0056] Figure 4A An illustration of a 4-parameter affine model is shown. Figure 4B A diagram of a 4-parameter affine model is shown. Based on the motion of control points, the motion field (v) of an affine encoder-decoder block is... x ,v y It is described as:

[0057]

[0058] The 6-parameter affine model has the following parameters: two parameters for translational motion in the horizontal and vertical directions respectively; one parameter for scaling motion and one parameter for rotational motion in the horizontal direction; and one parameter for scaling motion and one parameter for rotational motion in the vertical direction. This 6-parameter affine motion model is encoded and decoded using three CPMVs.

[0059] Figure 5 A diagram of a 6-parameter affine model is shown. (As shown...) Figure 5 As shown, the three control points of a 6-parameter affine block are located at the top left, top right, and bottom left corners of the block. The motion at the top left control point is related to translation, the motion at the top right control point is related to horizontal rotation and scaling, and the motion at the bottom left control point is related to vertical rotation and scaling. Compared to a 4-parameter affine motion model, the horizontal rotation and scaling motions in a 6-parameter model may not be the same as those in the vertical direction. Assume (V0, V1, V2) are... Figure 5 The top-left, top-right, and bottom-left MVs of the current block are used to calculate the motion vector (v) of each sub-block using the three MVs at the control points. x ,v y Export as:

[0060]

[0061] In VVC, the CPMV of an affine codec block is stored in a separate buffer. The stored CPMV is used only to generate affine CPMV predictors for affine merge mode (i.e., inheriting affine CPMV from adjacent affine blocks) and affine explicit mode (i.e., emitting affine CPMV as a signal according to a prediction-based scheme). Sub-block MVs derived from the CPMV are used for motion compensation, MV prediction for translation MVs, and deblocking.

[0062] Similar to motion compensation in regular inter-frame blocks, the MV of each affine sub-block can point to a reference sample at a fractional sample location. In this case, an interpolation filtering process is required to generate reference samples for fractional pixel locations. To control the worst-case memory bandwidth requirements and the worst-case interpolation computation complexity, a set of 6-tap interpolation filters is used for motion compensation in affine sub-blocks. Tables 1 and 2 illustrate the interpolation filters used for motion compensation in regular inter-frame blocks and affine blocks, respectively. As can be seen, the 6-tap interpolation filter for affine mode is directly derived from the 8-tap filter by directly adding the two outermost filter coefficients on each side of the 8-tap filter used in regular inter-frame blocks into a single filter coefficient for the 6-tap filter. That is, filter coefficients P0 and P5 in Table 2 are equal to the sum of filter coefficients P0 and P1 and filter coefficients P6 and P7 in Table 1, respectively.

[0063] Table 1. Luminance interpolation filters used for regular inter-frame blocks.

[0064]

[0065] Table 2. Luminance interpolation filters for affine blocks.

[0066]

[0067] In addition, for motion compensation of chroma samples, the same 4-tap interpolation filter used for regular inter-frame blocks (as illustrated in Table 3) is used for affine blocks.

[0068] Table 3 shows the chroma interpolation filters used for inter-frame blocks (i.e., affine blocks and non-affine blocks).

[0069]

[0070] Reference image resampling

[0071] Unlike HEVC, the emerging VVC standard supports fast resolution switching within a bitstream of the same content. This capability is called Reference Picture Resampling (RPR) or Adaptive Resolution Switching (ARC). In real-time video applications, allowing resolution changes within a single encoded video sequence, without requiring the insertion of random access pictures or Intra-Random Access Point (IRAP) pictures (such as IDR pictures or CRA pictures), not only allows compressed video data to adapt to dynamic communication channel conditions but also avoids the surge in bandwidth consumption caused by the relatively large size of IDR or CRA pictures. Specifically, the following typical user scenarios can benefit from the RPR feature:

[0072] Rate adaptation in video calls and conferencing: To adapt the encoded video to changing network conditions, the encoder can adapt by encoding smaller resolution images when network conditions worsen to the point that available bandwidth becomes lower. Currently, changing image resolution can only be done after the IRAP image; this has several problems. A reasonably high-quality IRAP image is much larger than the inter-frame encoded image, and correspondingly more complex to decode: this is time-consuming and resource-intensive. This is problematic when the decoder requests a resolution change for loading reasons. It can also break low-latency buffering conditions, forcing audio resynchronization, and the end-to-end latency of the stream will increase, at least temporarily. This results in a poor user experience.

[0073] Active speaker changes in multi-party video conferencing: In multi-party video conferencing, it's common for the active speaker to be displayed with a larger video size than the other participants. When the active speaker changes, the image resolution for each participant also needs to be adjusted. When such active speaker changes occur frequently, the need for features like ARC (Automatic Redirection) becomes more important.

[0074] Fast Startup in Streaming: In streaming applications, it's common to buffer a significant amount of decoded images before launching the display. Launching a lower-resolution bitstream allows the application to have enough images in the buffer for a faster startup and display.

[0075] Adaptive Stream Switching in Streaming: The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature called `@mediaStreamStructureId`. This enables switching between different representations at random access points of an open GOP with an undecodeable leading picture, such as a CRA picture in HEVC with an associated RASL picture. When two different representations of the same video have different bitrates but the same spatial resolution, and they have the same `@mediaStreamStructureId` value, switching between these two representations at a CRA picture with an associated RASL picture can be performed, and the RASL picture associated with the switch at the CRA picture can be decoded with acceptable quality, thus enabling seamless switching. Using ARC, the `@mediaStreamStructureId` feature can also be used for switching between DASH representations with different spatial resolutions.

[0076] At the 15th JVET meeting, the VVC standard officially supported the RPR feature. The main aspects of the existing RPR design in VVC are summarized below:

[0077] RPR High-Order Signaling

[0078] According to the current RPR design, in the Sequence Parameter Set (SPS), the two syntax elements `pic_width_max_in_luma_samples` and `pic_height_max_in_luma_samples` are signaled to specify the maximum width and height of the codec image referencing the SPS. Therefore, when the image resolution changes, a new Picture Parameter Set (PPS) needs to be set when the relevant syntax elements `pic_width_in_luma_samples` and `pic_height_in_luma_samples` are signaled to specify different image resolutions referencing the PPS. Bitstream consistency exists; that is, the values ​​of `pic_width_in_luma_samples` and `pic_height_in_luma_samples` should not exceed the values ​​of `pic_width_max_in_luma_samples` and `pic_height_max_in_luma_samples`. Table 4 illustrates the RPR-related signaling in the SPS and PPS.

[0079] Table 4. RPR Signaling in SPS and PPS

[0080]

[0081] Reference image resampling process

[0082] When resolution changes within a bitstream, the current image may have one or more reference images of different sizes. According to the current RPR design, when the image resolution changes, all MVs (Modular Values) of the current image are normalized to the sample grid of the current image, rather than the sample grid of the reference images. This makes the image resolution change transparent to the MV prediction process.

[0083] When the image resolution changes, in addition to the MV, samples in a reference block must also be upsampled / downsampled during motion compensation of the current block. In VVC, the scaling ratios, i.e., refPicWidthInLumaSample / picWidthInLuma and refPicHeightInLumaSample / picHeightInLumaSample, are limited to the range [1 / 8,2].

[0084] In the current RPR design, different interpolation filters are applied to interpolate the reference samples when the current image and its reference image are at different resolutions. Specifically, when the resolution of the reference image is equal to or less than the resolution of the current image, the default 8-tap and 4-tap interpolation filters are used to generate inter-frame prediction samples for luminance and chrominance samples, respectively. However, the default motion interpolation filter does not exhibit strong low-pass characteristics. When the resolution of the reference image is higher than that of the current image, using the default motion interpolation filter will result in non-negligible aliasing, which becomes more severe as the downsampling ratio increases. Therefore, to improve the inter-frame prediction efficiency of RPR, two different sets of downsampling filters are applied when the reference image has a higher resolution than the current image. Specifically, when the downsampling ratio is equal to or greater than 1.5:1, the 8-tap and 4-tap Lanczos filters shown in Tables 5 and 6 are used.

[0085] Table 5 shows the brightness interpolation filters when the downsampling ratio is equal to or greater than 1.5:1.

[0086]

[0087] Table 6. Chromaticity interpolation filters with a downsampling ratio equal to or greater than 1.5:1

[0088]

[0089] When the downsampling ratio is equal to or greater than 2:1, use the following 8-tap and 4-tap downsampling filters derived by applying a cosine window function to a 12-tap SHM downsampling filter (as shown in Tables 7 and 8).

[0090] Table 7. Luminance interpolation filters with a downsampling ratio equal to or greater than 2:1.

[0091]

[0092] Table 8. Chromaticity interpolation filters with a downsampling ratio equal to or greater than 2:1

[0093]

[0094] Finally, the downsampling filter above is only applied to generate luma and chroma prediction samples for non-affine inter-frame blocks. For affine mode, the default 8-tap and 4-tap motion interpolation filters are still applied to the downsampling.

[0095] Problems in existing RPR designs

[0096] The objective of this disclosure is to improve the encoding and decoding efficiency of affine modes when applying RPR. Specifically, it identifies the following problems in existing RPR designs in VVC:

[0097] First, as discussed earlier, when the reference image resolution is higher than the current image resolution, only additional downsampling filters are applied to motion compensation in non-affine mode. For affine mode, 6-tap and 4-tap motion interpolation filters are applied. Assuming these filters are derived from the default motion interpolation filters, they do not exhibit strong low-pass characteristics. Therefore, compared to non-affine mode, the predicted samples in affine mode will exhibit more severe aliasing artifacts due to the well-known Nyquist-Shannon sampling theorem. Thus, for better encoding and decoding performance, it is desirable to also apply appropriate low-pass filters to motion compensation in affine mode when downsampling is required.

[0098] Second, according to the existing RPR design, the fractional pixel position of the reference sample is determined based on the position of the current sample, the MV (memory value), and the resolution scaling ratio between the reference image and the current image. Therefore, when downsampling the reference block, this results in higher memory bandwidth consumption and computational complexity for interpolating the reference sample of the current block. Assume the size of the current block is M (width) × N (height). When the reference image size is the same as the current image, an integer sample of size (M+7) × (N+7) needs to be accessed in the reference image, and motion compensation for the current block requires 8 × (M × (N+7)) + 8 × M × N multiplications. If the downsampling scaling ratio is s, then the corresponding memory bandwidth and multiplications increase to (s × M + 7) × (s × N + 7) and 8 × (M × (s × N + 7)) + 8 × M × N. Tables 9 and 10 compare the number of integer samples and multiplications per sample for motion compensation of different block sizes when the RPR downsampling scaling ratios are 1.5X and 2X, respectively. In Tables 9 and 10, the columns under the name "RPR 1X" correspond to the case where the reference image and the current image have the same resolution, i.e., without RPR. The column "Ratio to RPR 1X" depicts the memory bandwidth / multiplications when the RPR downsampling ratio is greater than 1, compared to the corresponding worst-case number of multiplications in the regular inter-frame mode without RPR (i.e., 16×4 bidirectional prediction). As can be seen, there is a significant increase in memory bandwidth and computational complexity when the reference image has a higher resolution than the current image, compared to the worst-case complexity of regular inter-frame prediction. The peak increase comes from 16×4 bidirectional prediction, where memory bandwidth and multiplications are 231% and 127% of the worst-case memory bandwidth and multiplications for bidirectional prediction, respectively.

[0099] Table 9 shows the memory bandwidth consumption per sample when the RPR ratio is 1.5X and 2X.

[0100]

[0101] Table 10 shows the number of multiplications per sample when the RPR ratio is 1.5X and 2X.

[0102]

[0103] Third, in the existing RPR design, VVC only supports adaptive switching of image resolution within the same bitstream, while the bit depth used for encoding and decoding the video sequence remains the same. However, according to the "Requirements for Future Video Codec Standards" in the CfP used to release the VVC standard, it is clearly stated that "this standard should support fast representation switching in adaptive streaming services that provide multiple representations of the same content—each representation having different attributes (e.g., spatial resolution or sample bit depth)." In practical video applications, due to Single Instruction Multiple Data (SIMD) operations, allowing changes in the codec bit depth within the encoded video sequence provides video encoders / decoders, especially software codec implementations, with a more flexible performance / complexity tradeoff.

[0104] RPR encoding / decoding improvements

[0105] This disclosure proposes solutions to improve the efficiency of RPR encoding / decoding in VVC and reduce its memory bandwidth and computational complexity. More specifically, the techniques proposed in this disclosure can be summarized as follows:

[0106] First, in order to improve the RPR encoding and decoding efficiency of affine mode, a new low-pass interpolation filter is proposed to replace the existing 8-tap luminance and 4-tap chrominance interpolation filters used for affine when the resolution of the reference image is higher than that of the current image, i.e. when downsampling is required.

[0107] Second, in order to simplify RPR, it is proposed to disable RPR-based inter-frame prediction for certain CU sizes. Compared with the regular inter-frame mode without RPR, these CU sizes lead to a significant increase in memory bandwidth and computational complexity.

[0108] Third, a method is proposed that allows for dynamic changes in the internal bit depth used to encode and decode a video sequence.

[0109] Downsampling filter for affine mode

[0110] As mentioned above, regardless of whether the current image and its reference image have the same resolution, the default 6-tap and 4-tap motion interpolation filters are always applied to affine patterns. Similar to the interpolation filters used in HEVC, the default motion interpolation filter in VVC does not exhibit strong low-pass characteristics. When the spatial scaling ratio is close to 1, this default motion interpolation filter can provide acceptable predicted sample quality. However, when the downsampling ratio from the reference image to the current image resolution becomes larger, based on the Nyquist-Shannon sampling theorem, aliasing artifacts become much more severe if the same default motion interpolation filter is used. Especially when the applied MV points to a reference sample at an integer sample location, the default motion interpolation does not apply any filtering at all. This can lead to a significant degrade in the quality of predicted samples for affine blocks.

[0111] To mitigate aliasing artifacts caused by downsampling, this disclosure proposes replacing the existing default 6-tap / 4-tap interpolation filters with different interpolation filters that have stronger low-pass characteristics for motion compensation in affine modes. Furthermore, to maintain the same memory bandwidth and computational complexity as conventional motion compensation procedures, the proposed downsampling filter length is the same as that of existing interpolation filters used for affine modes, i.e., 6 taps for the luminance component and 4 taps for the chrominance component.

[0112] Figure 7 A method for decoding video signals is shown. This method can be applied, for example, to a decoder.

[0113] In step 710, the decoder can obtain a reference image I associated with a video block within the video signal.

[0114] In step 712, the decoder can obtain the reference sample I(i,j) of the video block from the reference block in the reference image I. i and j can, for example, represent the coordinates of a sample within the video block.

[0115] In step 714, when the video block is encoded and decoded in non-affine inter-frame mode and the resolution of the reference image I is greater than the resolution of the current image, the decoder can obtain a first downsampling filter and a second downsampling filter to generate luminance and chrominance inter-frame prediction samples of the video block, respectively.

[0116] In step 716, when the video block is encoded and decoded in affine mode and the resolution of the reference image is greater than the resolution of the current image, the decoder can obtain a third downsampling filter and a fourth downsampling filter to generate inter-frame prediction samples of the luminance and chrominance of the video block, respectively.

[0117] In step 718, the decoder can obtain inter-frame prediction samples of the video block based on the application of the third and fourth downsampling filters to the reference sample I(i,j).

[0118] Affine brightness downsampling filter

[0119] There are several ways to derive the luminance downsampling filter for affine mode.

[0120] Method 1: In one or more embodiments of this disclosure, a luminance downsampling filter for affine mode is proposed to be directly derived from an existing luminance downsampling filter for conventional inter-frame mode (i.e., non-affine mode). Specifically, this method derives a new 6-tap luminance downsampling filter from the 8-tap luminance downsampling filters in Table 5 (for scaling 1.5X) and Table 7 (for scaling 2X) by summing the two leftmost / rightmost filter coefficients of the 8-tap filter to obtain a single filter coefficient for the 6-tap filter. Tables 11 and 12 illustrate the proposed 6-tap luminance downsampling filters at spatial scaling ratios of 1.5:1 and 2:1, respectively.

[0121] Table 11. 6-tap brightness downsampling filter, scaling ratio equal to or greater than 1.5:1

[0122]

[0123] Table 12 shows the 6-tap luminance downsampling filter when the scaling ratio is equal to or greater than 2:1.

[0124]

[0125] Method 2: In one or more embodiments of this disclosure, a 6-tap affine downsampling filter is directly derived from an SHM filter derived based on a cosine windowed sinc function. Specifically, in this method, the affine downsampling filter is derived based on the following equation:

[0126]

[0127] Where L is the filter length, and h(n) is the frequency response of the ideal low-pass filter, which is calculated as follows:

[0128] h(n) = s·f c ·sinc(s·f c ·n), n=-∞,…,+∞ (4)

[0129] f c The cutoff frequency is s, and the scaling factor is s. w(n) is the cosine window function, defined as:

[0130]

[0131] In one example, assuming f is 0.9 and L = 6, Tables 13 and 14 illustrate the 6-tap luminance downsampling derived at spatial scaling ratios of 1.5X (i.e., s = 1.5) and 2X (i.e., s = 2).

[0132] Table 136 shows the brightness downsampling filter with a scaling ratio equal to or greater than 1.5:1.

[0133]

[0134] Table 14 shows the 6-tap brightness downsampling filter when the scaling ratio is equal to or greater than 2:1.

[0135]

[0136] It should be noted that in Tables 13 and 14, the filter coefficients are derived with a precision of 7-bit sign variables, which is the same precision as the downsampling filters used in the RPR design.

[0137] Figure 8 A method for decoding video signals is shown. This method can be applied, for example, to a decoder.

[0138] In step 810, the decoder can obtain the frequency response of the ideal low-pass filter based on the cutoff frequency and scaling factor.

[0139] In step 812, the decoder can obtain the cosine window function based on the filter length.

[0140] In step 814, the decoder can obtain a third downsampling filter based on the frequency response and the cosine window function.

[0141] Affine chromaticity downsampling filter

[0142] The following section presents three methods for downsampling chroma reference blocks when the resolution of the reference image is higher than that of the current image.

[0143] Method 1: In the first method, it is proposed to reuse the existing 1.5X (Table 6) and 2X (Table 8) 4-tap chromaticity downsampling filters designed for non-affine modes under RPR for downsampling of reference samples in affine modes.

[0144] Method 2: In the second method, it is proposed to reuse the default 4-tap chromaticity interpolation filter (Table 3) to downsample the reference samples of the affine mode.

[0145] Method 3: In the third method, a chroma downsampling filter is proposed based on the cosine windowed sinc function described in (3)-(5). Tables 15 and 16 respectively depict the derived 4-tap chroma downsampling filters with scaling ratios of 1.5X and 2X when the cutoff frequency of the cosine windowed sinc function is assumed to be 0.9.

[0146] Table 15. 4-tap chroma downsampling filter with a scaling ratio equal to or greater than 1.5:1

[0147]

[0148] Table 16. 4-tap chroma downsampling filters with a scaling ratio equal to or greater than 2:1

[0149]

[0150] Block size for constraints used in RPR mode

[0151] As analyzed in the "Problem Statement" section, existing RPRs introduce a significant increase in complexity when downsampling occurs (e.g., the number of integer samples accessed for motion compensation and the number of multiplications required). Specifically, the memory bandwidth and number of multiplications required when downsampling the reference block are 231% and 127% of the worst-case bidirectional prediction, respectively.

[0152] In one or more embodiments, it is proposed to disable bidirectional prediction during inter-frame prediction for certain block shapes—e.g., 4×N, N×4, and / or 8×8—when the resolution of the reference image is higher than that of the current image (but unidirectional prediction is still allowed). Tables 17 and 18 show the corresponding per-sample memory bandwidth and number of multiplications when bidirectional prediction is disabled for 4×N, N×4, and 8×8 block sizes during inter-frame prediction using RPR. As can be seen, with the proposed constraints, the memory bandwidth and number of multiplications are reduced to 130% and 107% of the worst-case bidirectional prediction for 1.5X downsampling and 116% and 113% of the worst-case bidirectional prediction for 2X downsampling, respectively.

[0153] Table 17 shows the per-sample memory bandwidth consumption for 1.5X and 2X downsampling ratios after applying block size constraints to RPR.

[0154]

[0155] Table 18 shows the number of per-sample multiplications when applying block size constraints to RPR for downsampling ratios of 1.5X and 2X.

[0156]

[0157] Although the examples above only disable bidirectional prediction in RPR mode for 4×N, N×4, and 8×8 block sizes, the constraints proposed still apply to other block sizes and inter-frame coding / decoding modes (e.g., unidirectional / bidirectional prediction, merged / non-merged modes, etc.) for technicians with state-of-the-art modern video technology.

[0158] Adaptive bit depth switching

[0159] In existing RPR designs, VVC only supports adaptive switching of image resolution within the same bitstream, while the bit depth used for encoding and decoding video sequences remains the same. However, as previously analyzed, allowing switching of encoding and decoding bit depth within the same bitstream can provide more flexibility for practical codec / decoder devices and offer different trade-offs between encoding / decoding performance and computational complexity.

[0160] In this section, an Adaptive Bit Depth Switching (ABS) method is proposed to allow changing the internal codec bit depth without requiring the introduction of an IRAP picture such as an Instant Decoder Refresh (IDR) picture.

[0161] Figure 6 A hypothetical example is depicted where the current image 620 and its reference images 610 and 630 are encoded and decoded at different internal bit depths. Figure 6 Reference image 610Ref0 using 8-bit encoding / decoding, current image 620 using 10-bit encoding / decoding, and reference image 630Ref1 using 12-bit encoding / decoding are shown. Specifically, in the following sections, high-order syntax signaling and modifications to the motion compensation process of the current VVC framework are proposed to support the proposed ABS capability.

[0162] High-order ABS signaling

[0163] For the proposed ABS signaling, in SPS, a new syntax element `sps_max_bit_depth_minus8` is proposed to replace the existing bit depth syntax element `bit_depth_minus8`, which specifies the maximum internal codec bit depth of the image used to reference the SPS codec. Then, when the codec bit depth changes, a new PPS syntax `pps_bit_depth_minus8` is sent to specify the different codec bit depth of the image referencing the PPS.

[0164] There is a bitstream consistency requirement, meaning the value of pps_bit_depth_minus8 should not exceed the value of sps_max_bit_depth_minus8. Table 19 illustrates the proposed ABS signaling in SPS and PPS.

[0165] Table 19 lists the proposed ABS signaling in SPS and PPS.

[0166]

[0167] Predicted sample bit depth adjustment

[0168] When there is a change in codec bit depth within a video sequence, a current image can be predicted from another reference image whose reconstructed samples are represented with a different bit depth precision. When this occurs, the predicted samples generated from motion compensation of the reference image should be adjusted to the codec bit depth of the current image.

[0169] Interaction with other codec tools

[0170] Assuming ABS is applied to a reference image and the current image can be represented with different precisions, some existing codec tools in VVC that derive certain codec parameters using reference samples may not function correctly. For example, in current VVC, Bidirectional Optical Flow (BDOF) and Decoder-Side Motion Vector Correction (DMVR) are two decoder-side techniques that use temporal prediction samples to improve inter-frame encoding and decoding efficiency. Specifically, the BDOF tool uses L0 and L1 prediction samples to calculate per-sample corrections to improve prediction sample quality, while DMVR relies on L0 and L1 prediction samples to correct motion vector precision at the sub-block level. Based on the above considerations, it is proposed that when either of the two prediction signals is encoded and decoded at a different bit depth than the current image, the BDOF and DMVR processes are always bypassed for an inter-frame block.

[0171] Figure 9 A computing environment 910 coupled to a user interface 960 is shown. The computing environment 910 may be part of a data processing server. The computing environment 910 includes a processor 920, a memory 940, and an I / O interface 950.

[0172] Processor 920 typically controls the overall operation of computing environment 910, such as operations associated with display, data acquisition, data communication, and image processing. Processor 920 may include one or more processors that execute instructions to perform all or some of the steps of the methods described above. Furthermore, processor 920 may include one or more modules that facilitate interaction between processor 920 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, GPU, etc.

[0173] Memory 940 is configured to store various types of data to support the operation of computing environment 910. Memory 940 may include predetermined software 942. Examples of such data include instructions for running any application or method on computing environment 910, video datasets, image data, and so on. Memory 940 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0174] I / O interface 950 provides an interface between processor 920 and peripheral interface modules such as keyboard, click wheel, buttons, etc. Buttons may include, but are not limited to, home button, start scan button, and stop scan button. I / O interface 950 can be coupled to encoder and decoder.

[0175] In one embodiment, a non-transitory computer-readable storage medium is also provided, including a plurality of programs, such as those contained in memory 940, executable by processor 920 in computing environment 910, for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0176] The non-transitory computer-readable storage medium stores a plurality of programs executable by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the above-described method for motion prediction.

[0177] In one embodiment, the computing environment 910 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0178] The description of this disclosure has been given for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings given in the foregoing description and associated drawings.

[0179] The examples have been chosen and described to explain the principles of this disclosure and to enable those skilled in the art to understand the disclosure in various different implementations and to best utilize the underlying principles and different implementations with different modifications suitable for the intended specific purpose. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed implementations, and modifications and other implementations are contemplated to be included within the scope of this disclosure.

Claims

1. A method for decoding a video signal, comprising: obtaining a reference picture associated with a video block within a video signal; obtaining reference samples for the video block from the reference picture; determining a luma interpolation filter for a video block coded in an affine motion mode based on a scaling ratio derived from resolutions of a current picture and the reference picture, wherein a first filter coefficient of the luma interpolation filter is equal to a sum of first two filter coefficients of a first luma interpolation filter, a last filter coefficient of the luma interpolation filter is equal to a sum of last two filter coefficients of the first luma interpolation filter, and other filter coefficients of the luma interpolation filter are equal in order to other filter coefficients of the first luma interpolation filter, wherein the first luma interpolation filter is used for a video block coded in a non-affine motion mode when the scaling ratio is equal to or greater than a first value; and obtaining luma inter prediction samples for the video block by applying the luma interpolation filter to the reference samples.

2. The method of claim 1, wherein the determining the luma interpolation filter for the video block coded in the affine motion mode based on the scaling ratio comprises: determining a second luma interpolation filter as the luma interpolation filter for the video block coded in the affine motion mode in response to the scaling ratio being equal to or greater than the first value, wherein the second luma interpolation filter is different from a third luma interpolation filter for the video block coded in the affine motion mode when the scaling ratio is not equal to or greater than the first value.

3. The method of claim 1, further comprising: determining a chroma interpolation filter for the video block coded in the affine motion mode based on a comparison between resolutions of the reference picture and the current picture; and obtaining chroma inter prediction samples for the video block by applying the chroma interpolation filter to the reference samples.

4. A method for encoding a video signal, comprising: determining a reference picture associated with a video block within a video signal; obtaining reference samples for the video block from the reference picture; determining a luma interpolation filter for a video block coded in an affine motion mode based on a scaling ratio derived from resolutions of a current picture and the reference picture, wherein a first filter coefficient of the luma interpolation filter is equal to a sum of first two filter coefficients of a first luma interpolation filter, a last filter coefficient of the luma interpolation filter is equal to a sum of last two filter coefficients of the first luma interpolation filter, and other filter coefficients of the luma interpolation filter are equal in order to other filter coefficients of the first luma interpolation filter, wherein the first luma interpolation filter is used for a video block coded in a non-affine motion mode when the scaling ratio is equal to or greater than a first value; and obtaining luma inter prediction samples for the video block by applying the luma interpolation filter to the reference samples.

5. The method of claim 4, wherein the determining the luma interpolation filter for the video block coded in the affine motion mode based on the scaling ratio comprises: determining a second luma interpolation filter as a luma interpolation filter for a video block coded in affine motion mode in response to the scaling ratio being equal to or greater than a first value, wherein the second luma interpolation filter is different from a third luma interpolation filter for a video block coded in affine motion mode when the scaling ratio is not equal to or greater than the first value.

6. The method of claim 4, further comprising: determining a chroma interpolation filter for a video block coded in affine motion mode based on a comparison between resolutions of the reference picture and the current picture; and obtaining chroma inter prediction samples of the video block by applying the chroma interpolation filter to the reference samples.

7. A computing device comprising: one or more processors; and a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors are configured such that, when executing the instructions, they perform the method of any of claims 1-6.

8. A non-transitory computer-readable storage medium storing a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the method of any of claims 1-6 and store a bitstream to be generated according to the method of any of claims 1-6.

9. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the method of any of claims 1-6.

10. A non-transitory computer-readable storage medium storing a bitstream generated by a computing device according to the method of any of claims 1-6.

11. A method of storing a bitstream, comprising: performing the method of any of claims 4-6 to generate a bitstream, and storing the bitstream.

12. A method of storing a bitstream, comprising: performing the following steps to generate a bitstream: determining a reference picture associated with a video block within a video signal, obtaining reference samples of the video block from the reference picture, determining a luma interpolation filter for a video block coded in affine motion mode based on a scaling ratio derived from resolutions of a current picture and the reference picture, wherein a first filter coefficient of the luma interpolation filter is equal to a sum of a first two filter coefficients of a first luma interpolation filter, a last filter coefficient of the luma interpolation filter is equal to a sum of a last two filter coefficients of the first luma interpolation filter, and other filter coefficients of the luma interpolation filter are equal in order to other filter coefficients of the first luma interpolation filter, wherein the first luma interpolation filter is used for a video block coded in non-affine motion mode when the scaling ratio is equal to or greater than a first value, and obtaining luma inter prediction samples of the video block by applying the luma interpolation filter to the reference samples; storing the bitstream, wherein the bitstream is to be decoded according to the method of any of claims 1-3. ​ 13. A method of transmitting a bitstream, comprising: performing the method of any of claims 4-6 to generate a bitstream, and transmitting the bitstream.

14. A method of transmitting a bitstream, comprising: performing the following steps to generate a bitstream: determining a reference picture associated with a video block within a video signal, obtaining reference samples of the video block from the reference picture, determining a luma interpolation filter for a video block coded in an affine mode of motion based on a scaling ratio derived from resolutions of a current picture and the reference picture, wherein a first filter coefficient of the luma interpolation filter is equal to a sum of first two filter coefficients of a first luma interpolation filter, a last filter coefficient of the luma interpolation filter is equal to a sum of last two filter coefficients of the first luma interpolation filter, and other filter coefficients of the luma interpolation filter are equal to corresponding other filter coefficients of the first luma interpolation filter in order, wherein the first luma interpolation filter is used for a video block coded in a non-affine mode of motion when the scaling ratio is equal to or larger than a first value, and obtaining luma inter prediction samples of the video block by applying the luma interpolation filter to the reference samples; transmitting the bitstream, wherein the bitstream is to be decoded by the method of any of claims 1-3.

Citation Information

Patent Citations

  • Method and apparatus for encoding / decoding scalable video signal

    CN105379277A

  • Intra block copy mode for screen content coding

    CN107646195A