Adaptive resolution change in video processing

By combining image-level resampling with a cosine windowed sinc filter, the complexity and information loss issues in traditional video coding when adapting to resolution changes are resolved, achieving efficient adaptive resolution switching and supporting fast and seamless switching of video quality.

CN120915941APending Publication Date: 2025-11-07HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511079479.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-13
Filing Date
2020-05-29
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

When adapting to changes in resolution, existing video coding technologies often employ traditional block-level resampling methods, which complicate encoder design, reduce coding efficiency, and are incompatible with some coding tools. Furthermore, interpolation filters may introduce aliasing issues during downsampling.

Method used

An image-level resampling method is adopted to resample the reference image before encoding or decoding to make it consistent with the resolution of the current image. Motion compensation interpolation is performed using a cosine windowed sinc filter, combined with progressive downsampling techniques to reduce information loss.

Benefits of technology

It enables transparent operation of the encoder and decoder during adaptive resolution changes, improves encoding efficiency, avoids the complexity and information loss caused by block-level operations, and supports fast and seamless switching of video quality across multiple representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915941A_ABST
    Figure CN120915941A_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for performing adaptive resolution changes during video encoding and decoding. The methods include comparing resolutions of a target image and a first reference image; in response to the fact that the target image and the first reference image have different resolutions, resampling the first reference image to generate a second reference image; and encoding or decoding a target image by using the second reference image.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 865,927, filed June 24, 2019, and U.S. Provisional Patent Application No. 62 / 900,439, filed September 13, 2019, both of which are incorporated by reference herein in their entirety. TECHNICAL FIELD

[0002] The present disclosure relates generally to video processing, and more specifically, to methods and systems for performing adaptive resolution change in video coding BACKGROUND

[0003] A video is a set of still images (or “frames”) that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is commonly referred to as encoding, and the decompression process is commonly referred to as decoding. There are multiple video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, Versatile Video Coding (VVC / H.266) standard, and AVS standards, specify specific video coding formats, developed by standardization organizations. As more and more video standards employ advanced video coding techniques, the coding efficiency of new video coding standards is also increasing. SUMMARY

[0004] Embodiments of the present invention provide a method of performing adaptive resolution change in a video coding process. In one exemplary embodiment, the method comprises comparing resolutions of a target image and a first reference image; in response to the target image having a different resolution than the first reference image, resampling the first reference image to generate a second reference image; and encoding or decoding the target image using the second reference image.

[0005] Embodiments of the present invention also provide a device for adaptive resolution change in video coding. In one exemplary embodiment, the device comprises one or more memories storing computer instructions; and one or more processors configured to execute the computer instructions to cause the device to: compare resolutions of a target image and a first reference image; in response to the target image having a different resolution than the first reference image, resample the first reference image to generate a second reference image; and encode or decode the target image using the second reference image.

[0006] Embodiments of the present disclosure also provide a non-transitory computer- readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for adaptive resolution change. In one example embodiment, the method comprises comparing resolutions of a target image and a first reference image; in response to the target image and the first reference image having different resolutions, resampling the first reference image to produce a second reference image; encoding or decoding the target image using the second reference image. BRIEF DESCRIPTION OF DRAWINGS

[0007] Embodiments of the present disclosure and various aspects thereof are illustrated in the following detailed description and the accompanying drawings. Various features shown in the figures are not to scale.

[0008] Figure 1 is a schematic diagram illustrating an example video encoder consistent with embodiments of the present disclosure.

[0009] Figure 2 is a schematic diagram illustrating an example video decoder consistent with embodiments of the present disclosure.

[0010] Figure 3 shows an example where resolutions of reference images differ from a current image consistent with embodiments of the present disclosure.

[0011] Figure 4 is a table illustrating a sub-pixel motion compensation interpolation filter for luma components in Versatile Video Coding (VVC) consistent with embodiments of the present disclosure.

[0012] Figure 5 is a table illustrating a sub-pixel motion compensation interpolation filter for chroma components in VVC consistent with embodiments of the present disclosure.

[0013] Figure 6 shows an example reference picture buffer consistent with embodiments of the present disclosure.

[0014] Figure 7 is an example table illustrating a set of supported resolutions including three different resolutions consistent with embodiments of the present disclosure.

[0015] Figure 8 shows an example decoded picture buffer (DPB) when both a resampled reference picture and an original reference picture are stored consistent with embodiments of the present disclosure.

[0016] Figure 9 shows a progressive down-sampling consistent with embodiments of the present disclosure.

[0017] Figure 10An exemplary video encoding process with resolution change, consistent with embodiments of the present disclosure, is shown.

[0018] Figure 11 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0019] Figure 12 An exemplary video encoding process with peak signal-to-noise ratio (PSNR) calculation, consistent with embodiments of the present disclosure, is shown.

[0020] Figure 13 A frequency response of an exemplary low-pass filter, consistent with embodiments of the present disclosure, is shown.

[0021] Figure 14 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0022] Figure 15 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0023] Figure 16 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0024] Figure 17 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0025] Figure 18 Table 1 is a table showing an exemplary down-sampling filter, consistent with embodiments of the present disclosure.

[0026] Figure 19 An exemplary luma sample interpolation filtering process for reference down-sampling, consistent with embodiments of the present disclosure, is shown.

[0027] Figure 20 An exemplary chroma sample interpolation filtering process for reference down-sampling, consistent with embodiments of the present disclosure, is shown.

[0028] Figure 21 An exemplary luma sample interpolation filtering process for reference down-sampling, consistent with embodiments of the present disclosure, is shown.

[0029] Figure 22 An exemplary chroma sample interpolation filtering process for reference down-sampling, consistent with embodiments of the present disclosure, is shown.

[0030] Figure 23 An exemplary luma sample interpolation filtering process for reference down-sampling, consistent with embodiments of the present disclosure, is shown.

[0031] Figure 24An exemplary chroma sample interpolation filter process for reference downsampling consistent with embodiments of the disclosure is shown.

[0032] Figure 25 An exemplary chroma fractional sample position calculation for reference downsampling consistent with embodiments of the disclosure is shown.

[0033] Figure 26 Table 1 is a table showing an 8-tap filter for MC interpolation consistent with embodiments of the disclosure with a reference downsampling ratio of 2: 1.

[0034] Figure 27 Table 2 is a table showing an 8-tap filter for MC interpolation consistent with embodiments of the disclosure with a reference downsampling ratio of 1.5: 1.

[0035] Figure 28 Table 1 is a table showing an 8-tap filter for MC interpolation consistent with embodiments of the disclosure with a reference downsampling ratio of 2: 1.

[0036] Figure 29 Table 2 is a table showing an 8-tap filter for MC interpolation consistent with embodiments of the disclosure with a reference downsampling ratio of 1.5: 1.

[0037] Figure 30 Table 3 is a table showing 6-tap filter coefficients for luma 4x4 block MC interpolation with a reference downsampling ratio of 2: 1 consistent with embodiments of the disclosure.

[0038] Figure 31 Table 4 is a table showing 6-tap filter coefficients for luma 4x4 block MC interpolation with a reference downsampling ratio of 1.5: 1 consistent with embodiments of the disclosure. DETAILED DESCRIPTION

[0039] Reference will now be made in detail to the example embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers represent the same or similar elements between the several drawings. The implementations set forth in the following description of example embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present disclosure as recited by the attached claims. Certain aspects of the present disclosure are described in more detail below. To the extent not explicitly recited in the claims, terms and / or definitions provided herein prevail.

[0040] A video is a set of still images (or "frames") arranged in a time sequence to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in a time sequence, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with display functionality) can be used to display such images in a time sequence. Furthermore, in some applications, a video capture device can transmit the captured video in real time to a video playback device (e.g., a computer with a display) for monitoring, conferencing, live broadcasting, etc.

[0041] To reduce the storage space and transmission bandwidth required for such applications, a video can be compressed. For example, a video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or by special-purpose hardware. A compression module is generally referred to as an "encoder," and a decompression module is generally referred to as a "decoder." An encoder and a decoder can be collectively referred to as a "codec." An encoder and a decoder can be implemented in any of a variety of suitable hardware, software, or combinations thereof. For example, a hardware implementation of an encoder and a decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. A software implementation of an encoder and a decoder can include program code, computer-executable instructions, firmware, or any suitable computer- implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress a video from a first encoding standard and re-compress the decompressed video using a second encoding standard, in which case the codec can be referred to as a "transcoder."

[0042] A video encoding process can identify and retain useful information that can be used to reconstruct images. If information that is ignored in the video encoding process cannot be completely reconstructed, the encoding process can be referred to as "lossy." Otherwise, it can be referred to as "lossless." Most encoding processes are lossy, which is a trade-off for reducing the required storage space and transmission bandwidth.

[0043] In many cases, useful information of an image being encoded (referred to as a "current image") can include changes relative to a reference image (e.g., a previously encoded or reconstructed image). Such changes can include changes in position, brightness, or color of pixels, with changes in position being the most interesting. Changes in position of a group of pixels representing an object can reflect a motion of the object between the reference image and the current image.

[0044] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, the Joint Exploration Model (“JEM”) reference software has been developed by JVET to explore technologies beyond HEVC. As coding technologies are incorporated into JEM, JEM achieves higher coding performance than HEVC. VCEG and MPEG have also officially started developing the next generation of video compression standard, the Versatile Video Coding (VVC / H.266) standard.

[0045] The VVC standard continues to include more coding technologies that provide better compression performance. VVC can be implemented in the same video coding system as used in modern video compression standards, such as HEVC, H.264 / AVC, MPEG2, H.263, etc. Figure 1 is a schematic diagram illustrating an exemplary video encoder 100 consistent with the disclosed embodiments. For example, the video encoder 100 can perform intra- or inter- coding of blocks within a video frame, including video blocks, or partitions or sub-partitions of video blocks. Intra-coding can rely on spatial prediction to reduce or remove spatial redundancy in a video within a given video frame. Inter-coding can rely on temporal prediction to reduce or remove temporal redundancy in a video in neighboring frames of a video sequence. Intra modes can refer to a plurality of spatial-based compression modes, while inter modes (e.g., single or bi-prediction) can refer to a plurality of temporal-based compression modes.

[0046] Referring to Figure 1 , the input video signal 102 can be processed block-by-block. For example, a video block unit can be a 16x16 pixel block (e.g., a macroblock (MB)). In HEVC, extended block sizes (e.g., coding units (CUs)) can be used to compress video signals of resolutions, such as 1080p and higher. In HEVC, a CU can include up to 64x64 luma samples and corresponding chroma samples. In VVC, the size of a CU can be further increased to include 128x128 luma samples and corresponding chroma samples. The CU can be divided into prediction units (PUs), to which a separate prediction method can be applied. Each input video block (e.g., MB, CU, PU, etc.) can be processed by using a spatial prediction unit 160 or a temporal prediction unit 162.

[0047] The spatial prediction unit 160 performs spatial prediction (e.g., intra-prediction) on a current CU using information about the same image / slice that contains the current CU. Spatial prediction can use pixels from neighboring blocks that have already been coded in the same video image / slice to predict the current video block. Spatial prediction can reduce spatial redundancy inherent in a video signal. Temporal prediction (e.g., inter-prediction or motion-compensated prediction) can use samples from already coded video images to predict the current video block. Temporal prediction can reduce temporal redundancy inherent in a video signal.

[0048] Temporal prediction unit 162 performs temporal prediction (e.g., inter prediction) for the current CU using information from a different picture / slice than the picture / slice containing the current CU. Temporal prediction of a video block can be signaled by one or more motion vectors. The motion vectors can indicate the amount and direction of motion between the current block and its prediction block(s) in the reference frame(s). If multiple reference pictures are supported, one or more reference picture indices can be sent for the video block. The reference index(es) can be used to identify which reference picture(s) in decoded picture buffer (DPB) 164 (also referred to as reference picture store 164) the temporal prediction signal can come from. After spatial or temporal prediction, mode decision and encoder control unit 180 in the encoder can select the prediction mode, e.g., based on rate-distortion optimization methods. The prediction block can be subtracted from the current video block at adder 116. The prediction residual can be transformed by transform unit 104 and quantized by quantization unit 106. The quantized residual coefficients can be inverse quantized at inverse quantization unit 110 and inverse transformed at inverse transform unit 112 to form a reconstructed residual. The reconstructed block can be added to the prediction block at adder 126 to form a reconstructed video block. Loop filtering such as a deblocking filter and an adaptive loop filter 166 can be applied on the reconstructed video block before it is put into the reference picture store 164 and used for encoding future video blocks. To form the output video bitstream 120, the encoding mode (e.g., inter or intra), prediction mode information, motion information, and quantized residual coefficients can be sent to entropy encoding unit 108 to be compressed and packed to form the bitstream 120.

[0049] In accordance with the disclosed embodiments, the above-described units of video encoder 100 can be implemented as software modules (e.g., computer programs implementing different functions), hardware components (e.g., different circuit blocks for performing respective functions), or a mix of software and hardware.

[0050] Figure 2 is a schematic diagram illustrating an exemplary video decoder 200, in accordance with the disclosed embodiments. Referring to Figure 2 , video bitstream 202 can be unpacked or entropy decoded at entropy decoding unit 208. The encoding mode or prediction information can be sent to spatial prediction unit 260 (e.g., if intra coded) or temporal prediction unit 262 (e.g., as inter coded) to form a prediction block. If inter coded, the prediction information can include a prediction block size, one or more motion vectors (e.g., which can indicate the direction and amount of motion), or one or more reference indices (e.g., which can indicate which reference picture the prediction signal is to be obtained from).

[0051] Motion compensation prediction can be implemented by time prediction unit 262 to form a time prediction block. Residual transform coefficients can be sent to inverse quantization unit 210 and inverse transform unit 212 to reconstruct the residual block. The prediction block and residual block can be added together at 226. The reconstructed block can be loop-filtered (via loop filter 266) before being stored in decoded image buffer (DPB) 264 (also called reference image storage 264). The reconstructed video in DPB 264 can be used to drive a display device or to predict future video blocks. Decoded video 220 can be displayed on a monitor.

[0052] Consistent with the disclosed embodiments, the aforementioned units of the video decoder 200 can be implemented as software modules (e.g., computer programs that implement different functions), hardware components (e.g., different circuit blocks for performing corresponding functions), or a combination of software and hardware.

[0053] One of the goals of the VVC standard is to provide video conferencing applications with the ability to tolerate network and device diversity. In particular, the VVC standard needs to provide the ability to quickly adapt to changes in the network environment, including rapidly reducing the coding bit rate when network conditions deteriorate and rapidly increasing video quality when network conditions improve. Furthermore, for adaptive streaming services that provide multiple representations of the same content, each of the multiple representations may have different attributes (e.g., spatial resolution or sample bit depth), and video quality may vary from low to high. Therefore, the VVC standard needs to support fast representation switching for adaptive streaming services. During the switch from one representation to another (e.g., from one resolution to another), the VVC standard needs to enable an efficient prediction structure without compromising the ability to switch quickly and seamlessly.

[0054] In some embodiments consistent with this disclosure, in order to change the resolution, the encoder (e.g., Figure 1 The encoder 100 sends an Instant Decoder Refresh (IDR) encoded image to clear the contents of the reference image buffer (e.g., ...). Figure 1 DPB164 and Figure 2 DPB 264 in (DDR). When receiving an IDR-encoded image, the decoder (e.g., Figure 2 The decoder 200 in the code marks all images in the reference buffer as "not used for reference". All subsequently transmitted images can be decoded without referencing any frame that was decoded before the IDR image. The first image in the encoded video sequence is always the IDR image.

[0055] In some embodiments consistent with the present disclosure, adaptive resolution change (ARC) techniques can be used to allow a video stream to change spatial resolution between coded pictures within the same video sequence without the need for new IDR pictures and without the need for multiple layers as in scalable video codecs. According to ARC techniques, at a resolution switch point, a currently coded picture is either predicted from a reference picture with the same resolution, if available, or from a reference picture of a different resolution by resampling the reference picture. In Figure 3 An adaptive resolution change is shown in Fig. 1, where the resolution of the reference picture "Ref 0" is the same as the resolution of the currently coded picture. However, the resolutions of the reference pictures "Ref 1" and "Ref 2" are different from the resolution of the current picture. In order to generate the motion compensated prediction signal for the current picture, both "Ref 1" and "Ref 2" are resampled to the resolution of the current picture.

[0056] Consistent with the present disclosure, when the resolution of a reference frame is different from the resolution of a current frame, one way to generate a motion compensated prediction signal is based on image resampling, where the reference picture is first resampled to the same resolution as the current picture, the existing motion compensation process with motion vectors can be applied. The motion vectors can be scaled (if they are sent in units before the resampling is applied) or not scaled (if they are sent in units after the resampling is applied). For image-based resampling, and in particular for downsampling of the reference picture (i.e. the resolution of the reference picture is larger than the resolution of the current picture), information can be lost in the reference resampling step before the motion compensation interpolation, because downsampling is typically implemented by low-pass filtering and decimation.

[0057] Another way is block-based resampling, where the resampling is performed at the block level. This is done by checking the one or more reference pictures used by the current block, and if one or both of them have a different resolution than the current picture, the resampling is performed in conjunction with the sub-pixel motion compensation interpolation process.

[0058] The present disclosure provides both image-based resampling methods and block-based resampling methods for use with ARC. The following description first addresses the disclosed image-based resampling methods, and then addresses the disclosed block-based resampling methods.

[0059] The disclosed image-based resampling method can address some of the problems caused by the conventional block-based resampling method. First, in the conventional block-level resampling method, when the resolution of the reference picture is different from the resolution of the current picture, on-the-fly block-level resampling can be performed. However, the block-level resampling process can complicate the encoder design because during the motion search, the encoder can have to resample the blocks at each search point. Motion search is typically a time-consuming process for the encoder, and thus the on-the-fly requirement during the motion search can complicate the motion search process. As an alternative to the on-the-fly block-level resampling, the encoder can pre-resample the reference picture so that the resampled reference picture has the same resolution as the current picture. However, this can degrade the coding efficiency because it can cause different prediction signals during the motion estimation and the motion compensation.

[0060] Second, in the conventional block-based resampling method, the design of the block-level resampling is incompatible with some other useful coding tools, such as subblock-based temporal motion vector prediction (SbTMVP), affine motion compensation prediction, decoder-side motion vector refinement (DMVR), etc. When ARC is enabled, these coding tools can be disabled. But such disabling can significantly degrade the coding performance.

[0061] Third, while any sub-pixel motion compensation interpolation filter known in the art can be used for block resampling, it can not be equally suitable for block upsampling and block downsampling. For example, in the case where the resolution of the reference picture is lower than the current picture, the interpolation filter used for motion compensation can be used for upsampling. However, in the case where the resolution of the reference picture is higher than the current picture, the interpolation filter is not suitable for downsampling because the interpolation filter cannot filter integer positions, which can cause aliasing. For example, Table 1 Figure 4 ) shows exemplary interpolation filter coefficients for integer positions of the luma component with various fractional sample position values. Table 2 Figure 5 ) shows exemplary interpolation filter coefficients for the chroma component with various fractional sample position values. As shown in Table 1 Figure 4 ) and Table 2 Figure 5 ), at the fractional sample position 0 (i.e., the integer position), the interpolation filter is not applied.

[0062] To avoid the above problems associated with the conventional block-based resampling method, the present disclosure provides an image-level resampling method that can be used for ARC. The image resampling process can involve upsampling or downsampling. Upsampling is increasing the spatial resolution while maintaining the two-dimensional (2D) representation of the image. In the upsampling process, the resolution of the reference picture is increased by interpolating the unavailable samples from the neighboring available samples. In the downsampling process, the resolution of the reference picture is decreased.

[0063] According to the disclosed embodiments, resampling can be performed at the picture level. In picture-level resampling, if the resolution of a reference picture is different from the resolution of the current picture, the reference picture is resampled to the resolution of the current picture. Motion estimation and / or compensation of the current picture can be performed based on the resampled reference picture. In this way, ARC can be implemented in a “transparent” manner at the encoder and the decoder, as the block-level operations are independent of the resolution change.

[0064] Picture-level resampling can be performed on-the-fly, i.e., while predicting the current picture. In some example embodiments, only the original (unresampled) reference pictures are stored in the decoded picture buffer (DPB), e.g., DPB 164 in the encoder 100( Figure 1 ) and DBP 264 in the decoder 200( Figure 2 ). The DPB is managed in the same way as in the current version of VVC. In the disclosed picture-level resampling approach, before the encoding or decoding of a picture is performed, the encoder or the decoder resamples a reference picture if the resolution of the reference picture in the DPB is different from the resolution of the current picture. In some embodiments, a resampled picture buffer can be used to store all resampled reference pictures of the current picture. The resampled pictures of the reference pictures are stored in the resampled picture buffer, and the motion search / compensation of the current picture is performed using pictures from the resampled picture buffer. After the encoding or decoding is completed, the resampled picture buffer is removed. Figure 6 An example of on-the-fly picture-level resampling is shown, in which a low resolution reference picture is resampled and stored in a resampled picture buffer. As shown in Figure 6 , the DPB of the encoder or the decoder contains 3 reference pictures. The resolutions of reference pictures “Ref 0” and “Ref 2” are the same as the resolution of the current picture. Therefore, “Ref 0” and “Ref 2” do not need to be resampled. However, the resolution of reference picture “Ref 1” is different from the resolution of the current picture, and therefore needs to be resampled. As a result, only the resampled “Ref 1” is stored in the reference picture buffer, while “Ref 0” and “Ref 2” are not stored in the reference picture buffer.

[0065] In accordance with some example embodiments, in the on-the-fly image level resampling, the number of resampled reference pictures that a picture can use is limited. For example, a maximum number of resampled reference pictures for a given current picture can be pre-set. For example, the maximum number can be set to one. In this case, a bitstream constraint can be imposed such that the encoder or decoder can allow at most one of the reference pictures to have a resolution different from that of the current picture, and all other reference pictures must have the same resolution. Because this maximum number indicates the size of the resampled picture buffer and the maximum number of resampling that can be performed by any picture in the current video sequence, the maximum number is directly related to the decoder complexity in the worst case. Therefore, this maximum number can be signaled as part of the sequence parameter set (SPS) or the picture parameter set (PPS), and it can be specified as part of the profile / tier definition.

[0066] In some example embodiments, two versions of a reference picture are stored in the DPB. One version has the original resolution, and the other version has the maximum resolution. If the resolution of the current picture is different from the original resolution or the maximum resolution, the encoder or decoder can perform on-the-fly downsampling from the stored maximum resolution picture. For picture output, the DPB always outputs the original (un-resampled) reference picture.

[0067] In some example embodiments, the resampling ratio can be chosen arbitrarily, and the vertical and horizontal scaling ratios are allowed to be different. Because of the application of the image level resampling, the block level operations of the encoder / decoder are independent of the resolution of the reference picture, allowing the arbitrary resampling ratio to be enabled without further complicating the block level design logic of the encoder / decoder.

[0068] In some example embodiments, the signaling of the coded resolution and the maximum resolution can be performed as follows. The maximum resolution of any picture in the video sequence is signaled in the SPS. The coded resolution of a picture can be signaled in the PPS or in the slice header. In either case, a flag is signaled whether the coded resolution is the same as the maximum resolution or not. If the coded resolution is different from the maximum resolution, the coded width and height of the current picture are signaled additionally. If the PPS signaling is used, the signaled coded resolution is applied to all pictures that refer to this PPS. If the slice header signaling is used, the signaled coded resolution is applied only to the current picture itself. In some embodiments, the difference between the current coded resolution and the signaled maximum resolution can be signaled in the SPS.

[0069] In some example embodiments, instead of using arbitrary resampling rates, the resolution is limited to a predefined set of N supported resolutions, where N is the number of supported resolutions within a sequence. The value of N and the supported resolutions can be signaled in the SPS. The coded resolution of a picture is signaled through the PPS or slice header. In either case, instead of signaling the actual coded resolution, the corresponding resolution index is signaled. Table 3 Figure 7 shows an example of the set of supported resolutions, where the coded sequence allows for 3 different resolutions. If the coded resolution of the current picture is 1440x816, the corresponding index value (=1) is signaled through the PPS or slice header.

[0070] In some example embodiments, both resampled and non-resampled reference pictures are stored in the DPB. For each reference picture, one original (i.e., non-resampled) picture and N-1 resampled copies are stored. If the resolution of a reference picture is different from the resolution of the current picture, the current picture is coded using the resampled picture in the DPB. For picture output, the DPB outputs the original (non-resampled) reference picture. As an example, Figure 8 the occupancy of the DPB for the set of supported resolutions given in Table 3 Figure 7 is shown. As Figure 8 indicated, the DPB contains N (e.g., N=3) copies of each picture used as a reference picture.

[0071] The choice of the value of N depends on the application. A larger value of N makes the resolution selection more flexible, but increases the complexity and memory requirements. A smaller value of N is suitable for devices with limited complexity, but limits the selectable resolutions. Thus, in some embodiments, the encoder can be given the flexibility to decide the value of N based on the application and device capabilities and signal it through the SPS.

[0072] Consistent with the present disclosure, the downsampling of a reference picture can be performed progressively (i.e., gradually). In a conventional (i.e., direct) downsampling, the downsampling of an input picture is performed in one step. If the downsampling rate is high, a single-step downsampling requires a longer tap downsampling filter to avoid severe aliasing problems. However, a longer tap filter is computationally expensive. In some example embodiments, to preserve the quality of the downsampled picture, if the downsampling rate is higher than a threshold (e.g., higher than 2:1 downsampling), a progressive downsampling is used where the downsampling is performed gradually. For example, a downsampling filter sufficient for 2:1 downsampling can be repeatedly applied to achieve a downsampling rate larger than 2:1. Figure 9An example of progressive down-sampling is shown, where 4x down-sampling is achieved in both horizontal and vertical dimensions on 2 images. The first image is 2x down-sampling 2 (bi-directional), and the second image is 2x down-sampling 2 (bi-directional).

[0073] Consistent with the present disclosure, the disclosed picture-level resampling method can be used with other coding tools. For example, in some example embodiments, scaled motion vectors based on resampled reference pictures can be used during temporal motion vector predictor (TMVP) or advanced temporal motion vector predictor (ATMVP).

[0074] Next, a test method to evaluate the coding tools for ARC is described. Since adaptive resolution change is mainly used to adapt to network bandwidth, the following test conditions can be considered to compare the coding efficiency of different ARC schemes when network bandwidth changes.

[0075] In some embodiments, the resolution is changed to half (in both vertical and horizontal directions) at a certain time instance and then restored to the original resolution after a certain period of time. Figure 10 An example of resolution change is shown. At time tl, the resolution is changed to half, and then restored to full resolution at time t2.

[0076] In some embodiments, when the resolution is reduced, if the down-sampling rate is too large (e.g., the down-sampling rate in a given dimension is larger than 2: 1), progressive down-sampling is used for a certain time. However, when up-sampling is used, progressive up-sampling is not used.

[0077] In some embodiments, the scalable HEVC test model (SMH) down-sampling filter can be used for source resampling. Table 4 Figure 11 ) shows the detailed filter coefficients with different fractional sampling positions.

[0078] In some example embodiments, two peak signal-to-noise ratios (PSNRs) are calculated to measure the video quality. In Figure 12 an illustration of PSNR calculation is shown. The first PSNR is calculated between the resampled source image and the decoded image. The second PSNR is calculated between the original source and the up-sampled decoded source.

[0079] In some embodiments, the number of images to be coded is the same as the current test condition for VVC.

[0080] Next, the disclosed block-based resampling method is described. In block-based resampling, combining resampling and motion compensated interpolation into one filtering operation can reduce the information loss described above. Take the case where the motion vector of the current block has half-pel precision in one dimension (e.g., the horizontal dimension), and the width of the reference picture is twice the width of the current picture. In this case, compared to image-level resampling, halving the width of the reference picture to match the width of the current picture, and then doing half-pel motion interpolation, the block-based resampling method will directly extract the odd positions in the reference picture as the reference block with half-pel precision. At the 15th JVET meeting, VVC adopted a block-based ARC resampling method, in which motion compensation (MC) interpolation and reference resampling are combined and performed in a one-step filter. In VVC Draft 6, the existing filter for MC interpolation without reference resampling is reused for MC interpolation with reference resampling. The same filter is used for reference upsampling and downsampling. The details of the filter selection are described as follows.

[0081] For luma components: if the half-pel AMVR mode is selected and the interpolation position is half-pel, use the 6-tap filter [3, 9, 20, 20, 9, 3]; if the motion compensation block size is 4x4, use the following 6-tap filter shown in Table 5 ( Figure 13 ); otherwise, use the 8-tap filter shown in Table 6 ( Figure 14 ). For chroma components, use the 4-tap filter shown in Table 7 ( Figure 15 ).

[0082] In VVC, the same filter is used for MC interpolation without reference resampling and MC interpolation with reference resampling. Although the VVC MCIF is designed based on DCT upsampling, it can not be appropriate to use it as a one-step filter that combines reference downsampling and MC interpolation. For example, for 0-phase filtering, i.e., the scaled mv is integer, the VVC 8-tap MCIF coefficients are [0, 0, 0, 64, 0, 0, 0, 0], which means that the predicted samples are directly copied from the reference samples. Although this can not be a problem for MC interpolation without reference downsampling or with reference upsampling, it can cause aliasing in the case of reference downsampling due to the lack of a low-pass filter before decimation.

[0083] The present disclosure provides a method for MC interpolation with reference downsampling using a cosine-windowed sinc filter.

[0084] A windowed sinc filter is a bandpass filter that separates one frequency band from other frequency bands. A windowed sinc filter is a low-pass filter whose frequency response allows all frequencies below the cutoff frequency to pass with amplitude 1, and blocks all frequencies above the cutoff frequency with zero amplitude, asFigure 16 are shown.

[0085] The filter kernel, also called the impulse response of the filter, is obtained by inverse Fourier transforming the frequency response of an ideal low-pass filter. The impulse response of a low-pass filter has the general form of a sinc function: where fc is the cut-off frequency, taking values in [0, 1], and r is the down-sampling rate, e.g. 1.5 for 1.5:1 down-sampling, and 2 for 2:1 down-sampling. The sinc function is defined as:

[0086] The sinc function is infinite. To make the filter kernel finite in length, a window function is used to truncate the filter kernel to L points. To obtain a smooth tapered curve, a cosine window function is used, which is given by:

[0087] The kernel of a cosine-windowed sinc filter is the product of the ideal response function h(n) and the cosine window function w(n):

[0088] Two parameters are chosen for the windowed sinc kernel, the cut-off frequency fc and the kernel length L. For the down-sampling filter used in the Scalable HEVC Test Model (SHM), fc = 0.9 and L = 13.

[0089] The filter coefficients obtained in (4) are real numbers. Applying the filter is equivalent to computing a weighted average of the reference samples, with weights being the filter coefficients. To enable efficient computation in digital computers or hardware, the coefficients are normalized, multiplied by a scaler and rounded to integers, so that the sum of the coefficients equals 2^N, where N is an integer. The filtered samples are divided by 2^N (equivalent to a right shift of N bits). For example, in VVC Draft 6, the sum of the interpolation filter coefficients is 64.

[0090] In some disclosed embodiments, the down-sampling filter is used in the SHM for VVC motion compensated interpolation with reference down-sampling for both luma and chroma components, and the existing MCIF is used for motion compensated interpolation with reference up-sampling. With kernel length L = 13, the first coefficient is small and rounded to zero, the filter length can be reduced to 12 without affecting the filter performance.

[0091] As an example, the filter coefficients for 2:1 down-sampling and 1.5:1 down-sampling are shown in Table 8 Figure 17 and Table 9 Figure 18 respectively.

[0092] In addition to the values of the coefficients, there are other differences between the design of the SHM filter and the existing MCIF.

[0093] As a first difference, the SHM filter needs to filter at both integer sample positions and fractional sample positions, while the MCIF only needs to filter at fractional sample positions. An example of the modification to the VVC Draft 6 luma sample interpolation filter process for the reference downsampling case is described in Table 10 Figure 19 ) below.

[0094] An example of the modification to the VVC Draft 6 chroma sample interpolation filter process for the reference downsampling case is described in Table 11 Figure 20 ) below.

[0095] As a second difference, the filter coefficients of the SHM filter sum to 128, while the filter coefficients of the existing MCIF sum to 64. In VVC Draft 6, in order to reduce the loss due to rounding errors, the intermediate prediction signal is kept at a higher precision (expressed with a higher bit-depth) than the output signal. The precision of the intermediate signal is referred to as the internal precision. In one embodiment, in order to keep the internal precision the same as in VVC Draft 6, the output of the SHM filter needs to be right-shifted by 1 additional bit compared to using the existing MCIF. An example of the modification to the VVC Draft 6 luma sample interpolation filter process for the reference downsampling case is shown in Table 12 Figure 21 ) below.

[0096] An example of the modification to the VVC Draft 6 chroma sample interpolation filter process for the reference downsampling case is shown in Table 13 Figure 22 ) below.

[0097] According to some embodiments, the internal precision can be increased by 1 bit, and an additional right-shift by 1 bit can be applied to convert the internal precision to the output precision. An example of the modification to the VVC Draft 6 luma sample interpolation filter process for the reference downsampling case is shown in Table 14 Figure 23 ) below.

[0098] An example of the modification to the VVC Draft 6 chroma sample interpolation filter process for the reference downsampling case is shown in Table 15 Figure 24 ) below.

[0099] As a third difference, the SHM filter has 12 taps. Thus, to generate an interpolated sample, 11 neighboring samples are needed (5 on the left, 6 on the right, or 5 on the top or 6 on the bottom). In comparison with MCIF, additional neighboring samples are fetched. In VVC Draft 6, the chroma mv precision is 1 / 32. However, the SHM filter only has 16 phases. Thus, the chroma mv can be rounded to 1 / 16 for reference downsampling. This can be done by right shifting the last 5 bits of the chroma mv by 1 bit. An example of the modification to the VVC Draft 6 chroma fractional sample position calculation for the reference downsampling case is shown in Table 16 Figure 25 .

[0100] According to some embodiments, to align with the existing MCIF design in VVC Draft, we propose to use an 8-tap cosine-windowed-sine filter. The filter coefficients can be derived by setting L = 9 in the cosine-windowed-sine filter function in Equation (4). The sum of the filter coefficients can be set to 64 to further align with the existing MCIF filter. Example filter coefficients for 2:1 and 1.5:1 ratios are shown in Table 17 Figure 26 and Table 18 Figure 27 , respectively.

[0101] According to some embodiments, to accommodate the 1 / 32 sampling precision of the chroma component, a 32-phase cosine-windowed-sine filter set can be used in the chroma motion compensated interpolation with reference downsampling. Examples of filter coefficients for 2:1 and 1.5:1 ratios are shown in Table 19 Figure 28 and Table 20 Figure 29 , respectively.

[0102] According to some embodiments, for 4x4 luma blocks, a 6-tap cosine-windowed-sine filter can be used for MC interpolation with reference downsampling. Examples of filter coefficients for 2:1 and 1.5:1 ratios are shown in Table 21 Figure 30 and Table 22 Figure 31 , respectively.

[0103] In some embodiments, a non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by an apparatus (e.g., the disclosed encoders and decoders) for performing the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, flexible disk, hard disk, solid-state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, PROM, and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, cache, register, any other memory chip or cartridge, and a carrier wave transporting data or instructions. This apparatus can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0104] It should be noted that the relational terms herein, such as first, second, and the like, are used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the words "comprises," "has," "contains," and "includes," and variations thereof, are all open-ended, and specifically do not exclude the presence of unrecited items or additional items.

[0105] As used herein, the term "or" encompasses all possible combinations, unless otherwise specifically stated, or unless the context requires otherwise. For example, if a database is stated to include A or B, then, unless otherwise specifically stated or unless the context requires otherwise, the database can include A; or B; or both A and B. As a second example, if a database is stated to include A, B, or C, then, unless otherwise specifically stated or unless the context requires otherwise, the database can include A; or B; or C; or A and B; or A and C; or B and C; or A and B and C.

[0106] It can be understood that the above-described embodiments can be realized by hardware, or software (program code), or a combination of hardware and software. If realized by software, it can be stored in the above-mentioned computer-readable medium. The software can perform the disclosed methods when executed by a processor. The computing units and other functional units described in the present disclosure can be realized by hardware, or software, or a combination of hardware and software. It can also be understood by those of ordinary skill in the art that a plurality of the above-mentioned modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into a plurality of sub-modules / sub-units.

[0107] In the foregoing specification, specific details of embodiments have been set forth. Modifications and alterations can occur to others upon reading the specification. It is intended that the application be construed as including all such modifications and alterations. It is intended that the application be construed as including all such modifications and alterations. The specification and examples given herein are to be considered illustrative and not restrictive, and the scope of the application is to be affected by that of the appended claims, along with the full scope of equivalents to which such claims are entitled. The sequence of steps shown in the figures is also intended to be illustrative only and is not intended to be limiting to any particular sequence of steps. Thus, those skilled in the art will appreciate that the steps can be performed in different orders while still implementing the same method.

[0108] Embodiments can be further described using the following clauses: 1. A computer-implemented method of video processing, comprising: comparing resolutions of a target picture and a first reference picture; in response to the target picture having a different resolution than the first reference picture, resampling the first reference picture to produce a second reference picture; and encoding or decoding the target picture using the second reference picture. 2. The method of clause 1, further comprising: storing the second reference picture in a first buffer, the first buffer being different from a second buffer storing decoded pictures used for predicting future pictures; and removing the second reference picture from the first buffer upon completion of encoding or decoding of the target picture. 3. The method of any of clauses 1 and 2, wherein the encoding or decoding the target picture comprises: encoding or decoding the target picture using no more than a predetermined number of resampled reference pictures. 4. The method of clause 3, wherein the encoding or decoding the target picture comprises: signaling the predetermined number as part of a sequence parameter set or a picture parameter set. 5. The method of clause 1, wherein the resampling the first reference picture to generate a second reference picture comprises: storing a first version and a second version of the first reference picture, the first version having an original resolution and the second version having a maximum resolution that allows resampling reference pictures; and resampling the maximum resolution version to generate the second reference picture. 6. The method of clause 5, further comprising: storing the first version and the second version of the first reference picture in a decoded picture buffer; and outputting the first version for encoding of a future picture. 7. The method of clause 1, wherein the resampling the first reference picture to generate a second reference picture comprises: generating the second reference picture at a supported resolution. 8. The method of clause 7, further comprising: the signaled information as part of a sequence parameter set or a picture parameter set indicates: the supported resolution, and a pixel size corresponding to the supported resolution. 9. The method of clause 8, wherein the information comprises: an index corresponding to at least one supported resolution. 10. The method of any of clauses 8 and 9, further comprising: the number of supported resolutions is set according to a configuration of a video application or a video device. 11. The method of any of clauses 1-10, wherein the resampling the first reference picture to generate a second reference picture: progressively down-samples the first reference picture to generate the second reference picture. 12. An apparatus comprising: one or more memories storing computer instructions; and one or more processors configured to execute the computer instructions to cause the device to: compare resolutions of a target picture and a first reference picture; in response to the target picture and the first reference picture having different resolutions, resample the first reference picture to generate a second reference picture; and encode or decode the target picture using the second reference picture. 13. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: comparing resolutions of a target picture and a first reference picture; in response to the target picture and the first reference picture having different resolutions, resampling the first reference picture to generate a second reference picture; and encoding or decoding the target picture using the second reference picture. 14. A computer-implemented method of video processing, comprising: applying a bandpass filter to the reference image in response to the target image and the reference image having different resolutions to perform motion compensated interpolation and generate a reference block; and encoding or decoding a block of the target image using the reference block. 15. The method of clause 14, wherein the bandpass filter is a cosine-tapered sine filter. 16. The method of clause 15, wherein the cosine-tapered sine filter has a kernel function wherein fc is a cutoff frequency of the cosine-tapered sine filter, L is a kernel length, and r is a down-sampling rate. 17. The method of clause 16, wherein fc is equal to 0.9 and L is equal to 13. 18. The method of any of clauses 15-17, wherein the cosine-tapered sine filter is an 8-tap filter. 19. The method of any of clauses 15-17, wherein the cosine-tapered sine filter is a 4-tap filter. 20. The method of any of clauses 15-17, wherein the cosine-tapered sine filter is a 32-phase filter. 21. The method of clause 20, wherein the 32-phase filter is used in chroma motion compensated interpolation. 22. The method of any of clauses 14-21, wherein the applying a bandpass filter to the reference image comprises: obtaining a luma sample or a chroma sample at a fractional sample location. 23. An apparatus comprising: one or more memories storing computer instructions; and one or more processors configured to execute the computer instructions to cause the device to: apply a bandpass filter to the reference image in response to the target image and the reference image having different resolutions to perform motion compensated interpolation and generate a reference block. and encode or decode a block of the target image using the reference block. 24. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: 25. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: comparing resolutions of a target image and a first reference image; in response to the target picture having a different resolution than the first reference picture, resampling the first reference picture to generate a second reference picture; and encoding or decoding a target picture using the second reference picture.

[0109] In the drawings and specification, there have been disclosed exemplary embodiments. However, many variations and modifications can be made to these embodiments. Consequently, it is intended that the scope of the application be limited only by the broadest interpretation of the appended claims to be accorded.

Claims

1. A computer-implemented method of video processing, applied to an encoder, comprising: determining a coding resolution of pictures associated with a video sequence; in response to the coding resolution being different from a maximum resolution of any picture in the video sequence, additionally signaling a coding width and height of a current picture; wherein the coding resolution is signaled in a PPS; the maximum resolution is signaled in an SPS.

2. The method of claim 1, wherein, a flag is signaled to indicate whether the coding resolution is the same as the maximum resolution.

3. The method of claim 1 or 2, wherein, the coding resolution is signaled in a PPS, the signaled coding resolution is applied to all pictures referring to the PPS.

4. The method of claim 1, wherein, a difference between the current coding resolution and the signaled maximum resolution is signaled in an SPS.

5. A computer-implemented method of video processing, applied to a decoder, comprising: decoding a coding resolution of pictures associated with a video sequence; in response to the coding resolution being different from a maximum resolution of any picture in the video sequence, additionally receiving a width and height of a current picture; wherein the coding resolution is received in a PPS; the maximum resolution is received in an SPS.

6. The method of claim 5, wherein, a flag is received to indicate whether the coding resolution is the same as the maximum resolution.

7. The method of claim 5 or 6, wherein, the coding resolution is received in a PPS, the received coding resolution is applied to all pictures referring to the PPS.

8. The method of claim 5, wherein, a difference between the current coding resolution and the received maximum resolution is received in an SPS.

9. A non-transitory computer-readable medium having stored thereon a set of instructions and a bitstream, the set of instructions executable by one or more processors to perform a method to generate the bitstream, the method comprising: determining a coding resolution of pictures associated with a video sequence; in response to the coding resolution being different from a maximum resolution of any picture in the video sequence, additionally signaling a coding width and height of a current picture; wherein the coding resolution is signaled in a PPS; the maximum resolution is signaled in an SPS.