Filter performing motion compensation interpolation by resampling
By using bandpass filters for resampling and motion compensation interpolation in video encoding, the problems of high bandwidth and large storage of high-definition videos are solved, and the encoding efficiency and storage requirements are improved. It is suitable for applications such as video surveillance.
Patent Information
- Application Number
- CN202510859389.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-27
- Filing Date
- 2020-08-20
- Publication Date
- 2025-08-01
AI Technical Summary
When processing high-definition videos, existing video encoding technologies face the problems of high bandwidth and large storage requirements. Especially in video surveillance applications, high-definition video bitstreams require high bandwidth for real-time transmission and large-scale storage, and existing methods are difficult to effectively reduce the bit rate of encoded videos.
The motion compensation interpolation method is performed using resampling based on a bandpass filter, and the reference block is generated by downsampling and interpolation of the reference picture, and the block of the target picture is processed to reduce information loss and improve encoding efficiency.
It effectively reduces the encoding bit rate of high-definition video, reduces storage and transmission requirements, and improves the efficiency of video encoding. Especially in video surveillance applications, it adapts to reference pictures of different resolutions for motion compensation interpolation, reducing information loss.
Smart Images

Figure CN120416508A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Application No. 62 / 904,608, filed on September 23, 2019, and U.S. Provisional Application No. 62 / 906,930, filed on September 27, 2019, both of which are incorporated herein by reference. Background Art
[0002] This application relates to video processing, and more particularly, to filters for performing motion compensated interpolation by resampling.
[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is generally referred to as encoding, and the decompression process is generally referred to as decoding. There are currently various video coding formats that employ standardized video coding techniques, and the most common ones are video coding formats based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, the AVS standard, etc., are developed by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Application
[0004] Embodiments of this application provide a computer-implemented method for performing motion compensated interpolation by resampling. The method may include: applying a band-pass filter to a reference picture for a target picture and a reference picture with different resolutions, and performing motion compensated interpolation by reference downsampling to generate a reference block; and processing a block of the target picture using the reference block.
[0005] Embodiments of this application also provide a system for performing motion compensated interpolation. The system may include: a memory storing a set of instructions; and at least one processor configured to execute the set of instructions to cause the device to perform: applying a band-pass filter to a reference picture for a target picture and a reference picture with different resolutions, and performing motion compensated interpolation by reference downsampling to generate a reference block; and processing a block of the target picture using the reference block.
[0006] An embodiment of the present application also provides a non - transitory computer - readable medium that stores a set of instructions executable by at least one processor of a computer system to cause the computer system to execute a method for processing video content. The method may include: for a target picture and a reference picture with different resolutions, applying a band - pass filter to the reference picture and performing motion - compensated interpolation through reference downsampling to generate a reference block; and processing a block of the target picture using the reference block. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments and aspects of the present application are shown in the following detailed description and the drawings. The various features shown in the drawings are not drawn to scale.
[0008] Figure 1 The structure of an exemplary video sequence consistent with an embodiment of the present application is shown.
[0009] Figure 2A A schematic diagram of an exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.
[0010] Figure 2B A schematic diagram of another exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.
[0011] Figure 3A A schematic diagram of an exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.
[0012] Figure 3B A schematic diagram of another exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present application is shown.
[0013] Figure 4 It is a block diagram of an exemplary apparatus for encoding or decoding video consistent with an embodiment of the present application.
[0014] Figure 5 A schematic diagram of a reference picture and a current picture consistent with an embodiment of the present application is shown.
[0015] Figure 6 An exemplary interpolation filter coefficient table based on a 6 - tap DCT for a 4×4 luminance component consistent with an embodiment of the present application is shown.
[0016] Figure 7 An exemplary 8 - tap interpolation filter coefficient table for a luminance component consistent with an embodiment of the present application is shown.
[0017] Figure 8 An exemplary 4 - tap 32 - phase interpolation filter coefficient table for a chrominance component consistent with an embodiment of the present application is shown.
[0018] Figure 9 Shows the frequency response of an exemplary ideal low - pass filter consistent with the embodiments of the present application.
[0019] Figure 10 Shows an exemplary 12 - tap cosine - window function difference interpolation filter coefficient table for 2:1 downsampling consistent with the embodiments of the present application.
[0020] Figure 11 Shows an exemplary 12 - tap cosine - window function difference interpolation filter coefficient table for 1.5:1 downsampling consistent with the embodiments of the present application.
[0021] Figure 12 Shows an exemplary luminance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0022] Figure 13 Shows an exemplary chrominance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0023] Figure 14 Shows an exemplary luminance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0024] Figure 15 Shows an exemplary chrominance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0025] Figure 16 Shows an exemplary luminance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0026] Figure 17 Shows an exemplary chrominance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0027] Figure 18 Shows an exemplary chrominance sample interpolation filtering process for reference downsampling consistent with the embodiments of the present application.
[0028] Figure 19 Shows an exemplary 8 - tap filter for performing MC interpolation consistent with the embodiments of the present application, with a reference downsampling rate of 2:1.
[0029] Figure 20 Shows an exemplary 8 - tap filter for performing MC interpolation consistent with the embodiments of the present application, with a reference downsampling rate of 1.5:1.
[0030] Figure 21 Shows an exemplary 8 - tap filter for performing MC interpolation consistent with the embodiments of the present application, with a reference downsampling rate of 2:1 and a phase of 32.
[0031] Figure 22 Shows an exemplary 8 - tap filter for performing MC difference consistent with an embodiment of the present application, with a reference down - sampling rate of 1.5:1 and a phase of 32.
[0032] Figure 23 Shows an exemplary 6 - tap filter coefficient table for 4×4 luma block MC interpolation consistent with an embodiment of the present application, with a reference down - sampling ratio of 2:1 and 16 phases.
[0033] Figure 24 Shows an exemplary 6 - tap filter coefficient table for luma 4×4 block MC interpolation with a reference down - sampling rate of 1.5:1 and 16 phases, consistent with an embodiment of the present application.
[0034] Figure 25 Shows a table of an exemplary 8 - tap filter consistent with an embodiment of the present application, which is used for MC interpolation with a reference down - sampling ratio of 2:1.
[0035] Figure 26 Shows an exemplary 8 - tap filter for MC interpolation consistent with an embodiment of the present application, with a reference down - sampling rate of 1.5:1.
[0036] Figure 27 Shows an exemplary 6 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 16 phases and a reference down - sampling ratio of 2:1.
[0037] Figure 28 Shows an exemplary 6 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 16 phases and a reference down - sampling ratio of 1.5:1.
[0038] Figure 29 Shows an exemplary 4 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 32 phases and a reference down - sampling ratio of 2:1.
[0039] Figure 30 Shows an exemplary 4 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 32 phases and a reference down - sampling ratio of 1.5:1.
[0040] Figure 31 Shows an exemplary 8 - tap filter for MC interpolation consistent with an embodiment of the present application, with a reference down - sampling rate of 2:1.
[0041] Figure 32Shows an exemplary 8 - tap filter for MC interpolation consistent with an embodiment of the present application, with a reference down - sampling rate of 1.5:1.
[0042] Figure 33 Shows an exemplary 6 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 16 phases and a reference down - sampling ratio of 2:1.
[0043] Figure 34 Shows an exemplary 6 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 16 phases and a reference down - sampling ratio of 1.5:1.
[0044] Figure 35 Shows an exemplary 4 - tap filter for MC interpolation consistent with an embodiment of the present application, which has 32 phases and a reference down - sampling ratio of 2:1.
[0045] Figure 36 Shows an exemplary 4 - tap filter for MC interpolation consistent with an embodiment of the present application, with a reference down - sampling rate of 1.5:1 and 32 phases.
[0046] Figure 37 Shows an example of filter coefficient signaling consistent with an embodiment of the present application.
[0047] Figure 38 Shows an exemplary syntax structure for marking resampling rates and corresponding filter banks consistent with an embodiment of the present application.
[0048] Figure 39 Is a flowchart of an exemplary method for processing video content consistent with an embodiment of the present application.
[0049] Figure 40 Is a flowchart of an exemplary method for rounding filter coefficients of a cosine window function difference filter consistent with an embodiment of the present application. Detailed Description of the Invention
[0050] Reference may now be made in detail to the example embodiments, which are illustrated as in the accompanying drawings. The following description refers to the accompanying drawings, where the same numbers in different drawings represent the same or similar elements, unless otherwise specified. The embodiments set forth in the following description of the example embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with aspects related to the present application as set forth in the appended claims. Specific aspects of the present application are described in more detail below. If there is a conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0051] Video coding systems are commonly used to compress digital video signals, e.g., to reduce storage space consumption or reduce transmission bandwidth consumption associated with such signals. As high-definition (HD) video (e.g., with a resolution of 1920×1080 pixels) becomes increasingly popular in various applications of video compression, such as online video streaming, video conferencing, or video surveillance, there is a continuous need to develop video coding tools that can improve the compression efficiency of video data.
[0052] For example, video surveillance applications are being used more and more widely in many application scenarios (e.g., security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are growing rapidly. Many video surveillance application scenarios prefer to provide users with high-definition video to capture more information, with more pixels per frame to capture such information. However, high-definition video bitstreams may have a high bit rate, which requires high-bandwidth transmission and large-space storage. For example, a surveillance video stream with an average resolution of 1920×1080 may require a bandwidth of up to 4Mbps for real-time transmission. In addition, video surveillance usually monitors continuously 7×24, which poses a great challenge to the storage system if video data is to be stored. Therefore, the need for high bandwidth and large storage for high-definition video has become a major limitation to its large-scale deployment in video surveillance.
[0053] Video is a set of static pictures (or "frames") arranged in a time series to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these pictures in chronological order, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display function) can be used to display such pictures in chronological order. In addition, in some applications, video capture devices can transmit the captured video in real time to a video playback device (e.g., a computer with a display), such as for surveillance, conferencing, or live streaming.
[0054] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for decompression is generally called a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be called a "transcoder".
[0055] The video encoding process can identify and retain useful information that can be used to reconstruct the picture and ignore unimportant information during the reconstruction process. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off made to reduce the required storage space and transmission bandwidth.
[0056] The useful information of the picture being encoded (referred to as the "current picture") includes the changes relative to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes can include changes in the position of pixels, brightness changes, or color changes, with the position change being the most concerned. The position change of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0057] A picture encoded without referring to another picture (i.e., it is its own reference picture) is called an "I-picture". A picture encoded using a previous picture as a reference picture is called a "P-picture". A picture encoded using a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B-picture".
[0058] As described above, video surveillance using high-definition video faces challenges of high bandwidth and large storage requirements. To address these challenges, the bitrate of the encoded video can be reduced. Among I-, P-, and B-pictures, the I-picture has the highest bitrate. Since the background of most surveillance videos is almost static, one way to reduce the overall bitrate of the encoded video is to use fewer I-pictures for video encoding.
[0059] However, the improvement using fewer I-pictures may be negligible because I-pictures usually do not dominate in the encoded video. For example, in a typical video bitstream, the ratio of I, B, and P pictures can be 1:20:9, where the I-picture can account for less than 10% of the total bitrate. In other words, in such an example, even if all I-pictures are removed, the reduced bitrate cannot exceed 10%.
[0060] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present application is illustrated. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real-life video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video input interface (e.g., a video broadcast transceiver) to receive videos from a video content provider.
[0061] As Figure 1 shown, the video sequence 100 can include a series of pictures arranged in time along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, and there are more pictures between pictures 106 and 108. In Figure 1 this, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, as indicated by the arrow, and its reference picture is picture 102. Picture 106 is a B-picture, as indicated by the arrow, and its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be directly before or after the picture. For example, the reference picture of picture 104 can be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are only examples, and the present application does not limit the embodiments of the reference pictures to the Figure 1 example shown in
[0062] Due to the computational complexity of such tasks, video codecs typically do not encode or decode an entire picture at once. Instead, they can divide the picture into basic segments and encode or decode the picture segment by segment. In this application, these basic segments are referred to as basic processing units ("BPUs"). For example, Figure 1 The structure 110 in shows an example structure of a picture (e.g., any one of pictures 102 - 108) of the video sequence 100. In structure 110, the picture is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in a picture can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing units for a picture can be selected based on a balance between coding efficiency and the level of detail to be retained in the basic processing units.
[0063] A basic processing unit can be a logical unit that can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color picture can include a luminance component (Y) representing achromatic luminance information, one or more chrominance components representing color information (e.g., Cb and Cr), and associated syntax elements, where the luminance and chrominance components can have basic processing units of the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components can be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luminance and chrominance components.
[0064] Video coding has multiple operation stages, examples of which are in Figures 2A - 2B and Figures 3A - 3BShown in. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments called "basic processing subunits" in this application. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit can have the same or smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which can be stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be carried out to a further level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0065] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, and the basic processing unit may be too large to make such a decision. The encoder can split the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.
[0066] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further split the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at the level of this segment.
[0067] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown (in [figure reference]), the encoder can perform a transformation operation on a residual basic processing unit (e.g., a CU). However, in some cases, these basic processing units may still be too large to process. The encoder can further divide the basic processing unit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transformation operation can be performed at the level of this segment. It should be noted that the partitioning scheme of the same basic processing unit can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.
[0068] In Figure 1 the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and its boundaries are shown as dashed lines. In different schemes, different basic processing units of the same picture can be divided into different basic processing subunits.
[0069] In some embodiments, to provide parallel processing and fault tolerance capabilities for video encoding and decoding, a picture can be divided into multiple processing regions. In this way, for a certain region of the picture, the encoding or decoding process can be independent of information from any other region of the picture. In other words, each region of the picture can be processed independently. In this way, the codec can process different regions of the picture in parallel, thereby improving the encoding efficiency. In addition, when the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, thereby providing fault tolerance capabilities. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions; "slices" and "tilings". It should also be noted that different pictures of the video sequence 100 can have different partitioning schemes for dividing the picture into regions.
[0070] For example, in Figure 1 the structure 110 is divided into three regions 114, 116, and 118, and its boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and regions of the structure 110 in [figure reference] are only examples, and the present application does not limit their implementation.
[0071] Figure 2A shows a schematic diagram of an example encoding process 200A consistent with an embodiment of the present application. For example, the encoding process 200A can be performed by an encoder. AsFigure 2A As shown, the encoder may encode video sequence 202 into video bitstream 228 according to encoding process 200A. Similar to Figure 1 the video sequence 100 in Figure 1 , the video sequence 202 may include a set of pictures arranged in chronological order (referred to as "original pictures"). Similar to
[0072] In Figure 2A , the encoder may input the basic processing units of the original pictures of video sequence 202 (referred to as "original BPUs") into prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder may subtract prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder may input residual BPU 210 into transform stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The encoder may input prediction data 206 and quantized transform coefficients 216 into binary encoding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder may input quantized transform coefficients 216 into inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which is used in the next iteration of process 200A in prediction stage 204. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0073] The encoder may iteratively execute process 200A to encode each original BPU of the original picture (in the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder may continue to encode the next picture in video sequence 202.
[0074] Referring to process 200A, the encoder can receive video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein can refer to receiving, inputting, obtaining, retrieving, fetching, reading, accessing, or any action used for inputting data in any way.
[0075] In prediction stage 204, in the current iteration, the encoder can receive the original BPU and prediction reference 224, and perform a prediction operation to generate prediction data 206 and prediction BPU 208. The prediction reference 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from prediction data 206 and prediction reference 224 that can be used to reconstruct the original BPU into prediction BPU 208.
[0076] Ideally, the predicted BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is usually slightly different from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the value of the pixel corresponding to the original BPU (e.g., grayscale value or RGB value) from the value of the pixel corresponding to the predicted BPU208. Each pixel of the residual BPU 210 can have a residual value that is the result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared with the original BPU, the prediction data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality, thereby compressing the original BPU.
[0077] To further compress the residual BPU 210, in transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", each basis pattern associated with a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component of the variation of the residual BPU210 (e.g., the frequency of brightness variation). No basis pattern can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0078] Different transformation algorithms can use different basic patterns. In the transformation stage 212, various transformation algorithms can be used, such as discrete cosine transform, discrete sine transform, etc. The transformation in the transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basic pattern by the corresponding correlation coefficients and sum the products to generate a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (so the basic pattern is the same). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly reducing the quality. Therefore, the residual BPU 210 is further compressed.
[0079] The encoder can further compress the transformation coefficients in the quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequency). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without significantly degrading the decoding quality. For example, in the quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After this operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, while the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the zero-valued quantized transformation coefficients 216 to further compress the transformation coefficients through this transformation coefficient. The quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0080] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. Generally, the quantization stage 214 can cause the most information loss in the process 200A. The greater the information loss, the fewer bits the quantized transformation coefficients 216 may require. To obtain different degrees of information loss, the encoder can use different quantization parameter values or any other parameters in the quantization process.
[0081] In the binary coding stage 226, the encoder can use binary coding techniques (such as entropy coding, variable - length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm) to code the predicted data 206 and the quantized transform coefficients 216. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the encoder can code other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. The encoder can use the output data of the binary coding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 can be further packed for network transmission.
[0082] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0083] It should be noted that other variants of process 200A can also be used to code the video sequence 202. In some embodiments, the encoder can execute the various stages of process 200A in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit Figure 2A one or more of the stages.
[0084] Figure 2B A schematic diagram of another example coding process 200B consistent with the embodiments of the present application is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0085] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-frame prediction") can use pixels from one or more already-encoded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the spatial redundancy inherent in the picture. Temporal prediction (e.g., inter-picture prediction or "inter-frame prediction") can use regions from one or more already-encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the temporal redundancy inherent in the picture.
[0086] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-frame prediction. For the original BPU of the picture being encoded, the prediction reference 224 can include one or more neighboring BPUs that have already been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder can generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform extrapolation at the pixel level, for example, by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard being used. For intra-frame prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, the extrapolation parameters, the direction of the neighboring BPUs used relative to the original BPU, etc.
[0087] For another example, in the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed by the BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the reconstructed BPU. When all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference pictures (referred to as the "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at the position in the reference picture that has the same coordinates as the original BPU in the current picture and may extend outward a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller, equal to, larger, or different in shape) from the original BPU. Since the reference picture and the current picture are temporally separated in the time axis (e.g., as Figure 1 shown), it can be considered that over time, the matching region "moves" to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When using multiple reference pictures (e.g., as in Figure 1 Picture 106), the encoder may search for the matching region and determine its associated motion vector for each reference picture. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference pictures.
[0088] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, etc.
[0089] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference picture according to the motion vector, where the encoder may predict the original BPU of the current picture. When using multiple reference pictures (e.g., as in Figure 1In the picture 106), the encoder can move the matching region of the reference picture according to the corresponding motion vector and average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of each matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching regions.
[0090] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same time direction as the current picture. For example, Figure 1 The picture 104 in is an unidirectional inter-frame prediction picture, where the reference picture (i.e., picture 102) is before picture 104. Bidirectional inter-frame prediction can use one or more reference pictures in two time directions relative to the current picture. For example, Figure 1 The picture 106 in is a bidirectional inter-frame prediction picture, where the reference pictures (i.e., pictures 104 and 108) are in two time directions relative to picture 104.
[0091] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode to minimize the value of the cost function. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0092] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can directly input the prediction reference 224 into the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs have been encoded and reconstructed), the encoder can input the prediction reference 224 into the loop filter stage 232. At this time, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking effect) introduced by inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The reference picture for loop filtering can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter prediction reference picture for future pictures of the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength), as well as the quantized transform coefficients 216, prediction data 206, and other information, in the binary coding stage 226.
[0093] Figure 3A FIG. shows a schematic diagram of an example decoding process 300A consistent with an embodiment of the present application. Process 300A can be a decompression process corresponding to the compression process 200A in FIG. 2. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figures 2A - 2B in the quantization stage 214), generally, the video stream 304 is different from the video sequence 202. Similar to Figures 2A - 2B processes 200A and 200B in FIG., the decoder can perform process 300A on each picture encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode a basic processing unit in one iteration of the decoding process 300A. In some embodiments, the decoder can perform process 300A in parallel on regions (e.g., regions 114-118) of each picture encoded in the video bitstream 228.
[0094] In Figure 3AIn [the process], the decoder may input a part of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as "encoded BPU") into the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode this part into prediction data 206 and quantized transform coefficients 216. The decoder may input the quantized transform coefficients 216 into the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The decoder may input the prediction data 206 into the prediction stage 204 to generate the predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may input the predicted reference 224 into the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0095] The decoder may iteratively execute process 300A to decode each encoded BPU of the encoded picture and generate the predicted reference 224 for decoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder may output the picture to the video stream 304 for display and continue to decode the next encoded picture in the video bitstream 228.
[0096] In the binary decoding stage 302, the decoder may perform the inverse operations of the binary encoding techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context - adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may also decode other information in the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the quantization parameter process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder may unpack the video bitstream 228 before inputting it into the binary decoding stage 302.
[0097] Figure 3B A schematic diagram of another example decoding process 300B consistent with the embodiments of the present application is shown. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0098] In process 300B, for the coded basic processing unit (referred to as the "current BPU") of the coded picture being decoded (referred to as the "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder uses to code the current BPU. For example, if the encoder uses intra prediction to code the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., an identification value) indicating intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, extrapolation parameters, the directions of the adjacent BPUs relative to the original BPU, and the like. For example, if the encoder uses inter prediction to code the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., an identification value) indicating inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, the number of reference pictures associated with the current BPU, the weights respectively associated with the reference pictures, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference pictures, one or more motion vectors respectively associated with the matching regions, and the like.
[0099] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B and will not be elaborated further below. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A .
[0100] In process 300B, the decoder can input the prediction reference 224 into the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly input the prediction reference 224 into the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs have been decoded), the encoder can input the prediction reference 224 into the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder canFigure 2B Apply the loop filter to the prediction reference 224 in the manner described. The reference picture for loop filtering can be stored in buffer 234 (e.g., the decoded picture buffer in a computer memory) for later use (e.g., as an inter prediction reference picture for future coded pictures of video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter prediction is used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).
[0101] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video that is consistent with embodiments of the present application. As Figure 4 shown, apparatus 400 may include a processor 402. When processor 402 executes the instructions described herein, apparatus 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on a chip (SoCs), application specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may also be a group of processors grouped as a single logical controller. For example, as Figure 4 shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0102] Device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in processes 200A, 200B, 300A, or 300B) and data to be processed (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or manipulate the data for processing. The memory 404 may include high-speed random access storage devices or non-volatile storage devices. In some embodiments, the memory 404 may include any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped as a single logical component ( Figure 4 not shown).
[0103] The bus 410 may be a communication device for transferring data between components inside the device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., universal serial bus port, peripheral component interconnect express port), etc.
[0104] For ease of explanation without ambiguity, in this application, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuit may be implemented entirely in hardware, or a combination of software, hardware, or firmware. In addition, the data processing circuit may be a single independent module, or may be fully or partially combined into any other component of the device 400.
[0105] The device 400 may further include a network interface 406 to provide wired or wireless communication related to a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0106] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4 shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays. Displays or light-emitting diode displays), video input devices (e.g., cameras or input interfaces communicatively coupled to video archives), etc.
[0107] It should be noted that a video codec (e.g., the codec performing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0108] One of the key requirements of the VVC standard is to provide the ability to tolerate network and device diversity for video conferencing applications and to be able to quickly adapt to changing network environments, including quickly reducing the encoding bitrate when network conditions deteriorate and quickly improving video quality when network conditions improve. The expected video quality can range from very low to very high. The standard should also support fast representation switching for adaptive streaming services that provide multiple representations of the same content, each with different attributes (e.g., spatial resolution or sampling bit depth). When switching from one representation to another (e.g., from one resolution to another), the standard should be able to use an effective prediction structure without compromising the ability to switch quickly and seamlessly.
[0109] The purpose of adaptive resolution change (ARC) is to allow a stream to change spatial resolution between encoded pictures in the same video sequence, such as whether a new IDR frame is required in a scalable video codec and whether multiple layers are required. Instead, at a switching point, the resolution of the picture changes, and prediction can be made from reference pictures with the same resolution (if available) and reference pictures with different resolutions. If the reference pictures have different resolutions, they are resampled, as Figure 5 shown. For example, as Figure 5 shown, the resolution of reference picture 506 is the same as the resolution of current picture 502, while the resolutions of reference pictures 504 and 508 are different from the resolution of current picture 502. After reference pictures 504 and 508 are resampled to the resolution of current picture 502, motion compensation prediction can be performed on these reference pictures. Therefore, adaptive resolution change (ARC) is sometimes also referred to as reference picture resampling (RPR), and these two terms are used interchangeably in this application.
[0110] When the resolution of the reference frame is different from that of the current frame, one way to generate a motion-compensated prediction signal is based on picture resampling, where the reference picture is first resampled to the same resolution as the current picture, and an existing motion compensation process with motion vectors can be applied. The motion vectors can be scaled (if sent in units before applying resampling) or not scaled (if sent in units after applying resampling). For picture-based resampling, especially for downsampling of the reference picture (i.e., the resolution of the reference picture is greater than that of the current picture), information may be lost in the reference resampling step before motion compensation interpolation because downsampling is usually achieved through low-pass filtering and decimation).
[0111] Another method is block-based resampling, i.e., resampling is performed at the block level. This is done by examining the reference pictures used by the current block, and if one or both of the reference pictures have a different resolution from the current picture, resampling is performed in combination with a sub-pixel motion compensation interpolation process.
[0112] In block-based resampling, combining resampling and motion compensation interpolation into a single filtering operation can reduce the above-mentioned information loss. Take the following situation as an example: the motion vector of the current block has half-pixel accuracy in one dimension (e.g., the horizontal dimension), and the width of the reference picture is twice that of the current picture. In this case, compared with picture-level resampling that reduces the width of the reference picture by half to match the width of the current picture and then performs half-pixel motion interpolation, the block-based resampling method can directly take the odd positions in the reference picture as reference blocks with half-pixel accuracy. At the 15th JVET meeting, the VVC adopted a block-based ARC resampling method, where motion compensation (MC) interpolation and reference resampling are combined and performed in a single-step filter. In VVC draft 6, the existing filter for MC interpolation for non-reference resampling is reused for MC interpolation with reference resampling. The same filter is used for reference upsampling and reference downsampling. The details of the filter selection are described below.
[0113] For the luminance component, if the half-pixel AMVR mode is selected and the interpolation position is half-pixel, a 6-tap filter [3, 9, 20, 20, 9, 3] is used. If the motion compensation block size is 4×4, a 6-tap filter as shown in Figure 6 Table 6 is used. Otherwise, an 8-tap filter as shown in Figure 7 Table 7 is used.
[0114] For the chrominance component, a 4-tap filter as shown in Figure 8 Table 8 is used.
[0115] In VVC, the same filter is used for MC interpolation without reference resampling and MC interpolation with reference resampling. Although the VVC motion compensation interpolation filter (MCIF) is designed based on DCT upsampling, it may not be suitable to use it as a one-step filter that combines reference downsampling and MC interpolation. For example, for phase 0 filtering (e.g., the scaled motion vector is an integer), the VVC 8-tap MCIF coefficients are [0, 0, 0, 64, 0, 0, 0, 0], which means that the predicted sample is directly copied from the reference sample. Although there may be no problem with MC interpolation without reference downsampling or reference upsampling, in the case of reference downsampling, due to the lack of a low-pass filter before decimation, this may cause aliasing artifacts.
[0116] This application provides a method for MC interpolation with reference downsampling using a cosine window function difference filter.
[0117] The window function difference filter is a band-pass filter that separates one frequency band from other frequency bands. The window function difference filter is a low-pass filter whose frequency response allows all frequencies below the cut-off frequency to pass through with an amplitude of 1 and stops all frequencies above the cut-off frequency with an amplitude of 0, as Figure 9 shown.
[0118] The filter kernel, also known as the impulse response of the filter, is obtained by taking the inverse Fourier transform of the frequency response of an ideal low-pass filter. The impulse response of the low-pass filter is in the general form of an interpolation function (sinc function) based on the following formula (1). where fc is the cut-off frequency, whose value is in [0, 1], and r is the downsampling rate, i.e., 1.5 for 1.5:1 downsampling and 2 for 2:1 downsampling. The interpolation function is defined based on the following formula (2).
[0119] The interpolation function is infinite. To make the filter kernel have a finite length, a window function is used to truncate the filter kernel to L points. To obtain a smooth tapered curve, a cosine window function based on the following formula (3) is used.
[0120] The kernel of the cosine window function difference filter is the product of the ideal response function h(n) and the cosine window function w(n), which is obtained by the following formula (4).
[0121] The window function difference kernel has two parameters available for selection, namely the cut-off frequency fc and the kernel length L. By adjusting the values of L and fc, the desired filter response can be obtained. For example, for the downsampling filter used in the scalable HEVC test model (SHM), fc = 0.9 and L = 13.
[0122] The filter coefficients obtained in formula (4) are real numbers. Applying the filter is equivalent to calculating the weighted average of the reference samples, with the weights being the filter coefficients. For efficient calculation in a digital computer or hardware, the coefficients are normalized, multiplied by a scalar, and rounded to integers such that the sum of the coefficients equals 2^N, where N is an integer. The filtered sample is divided by 2^N (equivalent to a right shift of N bits). For example, in VVC draft 6, the sum of the interpolation filter coefficients is 64.
[0123] In some embodiments, for the luminance and chrominance components, a downsampling filter can be used in the SHM for reference downsampling in VVC motion compensation interpolation, and an existing MCIF can be used for reference upsampling in motion compensation interpolation. When the kernel length L = 13, the first coefficient is very small and rounds to zero, and the filter length can be reduced to 12 without affecting the filter performance.
[0124] For example, Figure 10 Table 10 of Figure 11 and Table 11 of
[0125] respectively show the filter coefficients for 2:1 downsampling and 1.5:1 downsampling.
[0126] As a first difference, the SHM filter needs to perform filtering at both integer sampling positions and fractional sampling positions, while the MCIF only needs to perform filtering at fractional sampling positions. The following Figure 12 Table 12 of
[0127] The following Figure 13 Table 13 of
[0128] As a second difference, for the SHM filter, the sum of the filter coefficients is 128, while for the existing MCIF, the sum of the filter coefficients is 64. In VVC Draft 6, to reduce the loss caused by rounding errors, the intermediate prediction signal is maintained at a higher precision (represented by a higher bit depth) than the output signal. The precision of the intermediate signal is referred to as the internal precision. In some embodiments, to maintain the same internal precision as in VVC Draft 6, compared with using the existing MCIF, the output of the SHM filter needs to be right-shifted by an additional 1 bit. Below Figure 14 Table 14 below shows an example of the modification of the interpolation filtering process for the luma samples of VVC Draft 6 for the reference downsampling case.
[0129] Below Figure 15 Table 15 below shows an example of the modification of the interpolation filtering process for the chroma samples of VVC Draft 6 for the reference downsampling case.
[0130] According to some embodiments, the internal precision is increased by 1 bit, and a 1-bit right shift can be used to convert the internal precision to the output precision. Figure 16 Table 16 below shows an example of the modification of the interpolation filtering process for the luma samples of VVC Draft 6 for the reference downsampling case.
[0131] Below Figure 17 Table 17 below shows an example of the modification of the interpolation filtering process for the chroma samples of VVC Draft 6 for the reference downsampling case.
[0132] As a third difference, the SHM filter has 12 taps. Therefore, to generate an interpolated sample, 11 adjacent samples are required (5 on the left, 6 on the right, or 5 above or 6 below). Compared with MCIF, more adjacent samples are extracted. In VVC Draft 6, the chroma mv precision is 1 / 32. However, the SHM filter has only 16 phases. Therefore, the chroma mv can be rounded to 1 / 16 and used as a reference for downsampling. This can be achieved by right-shifting the last 5 bits of the chroma mv by 1 bit. Figure 18 Table 18 below shows an example of the modification of the calculation of the chroma fractional sampling positions of VVC Draft 6 for the reference downsampling case.
[0133] According to some embodiments, to be consistent with the existing MCIF design in the VVC draft, it is recommended to use an 8-tap cosine window function interpolation filter. The filter coefficients can be derived by setting L = 9 in the cosine window function interpolation function in Equation (4). The sum of the filter coefficients can be set to 64 to further align with the existing MCIF filter. In the disclosed embodiments, the filters for complementary phases can be symmetric (e.g., the filter coefficients for complementary phases are reversed) or asymmetric. Figure 19 Table 19 andFigure 20 Table 20 shows example filter coefficients for ratios of 2:1 and 1.5:1, respectively.
[0134] According to some embodiments, to accommodate the 1 / 32 sampling precision of the chrominance component, a 32-phase cosine window function difference filter bank can be used in reference downsampled chrominance motion compensation interpolation. Figure 21 The following Table 21 and Figure 22 Table 22 show example filter coefficients for ratios of 2:1 and 1.5:1, respectively.
[0135] According to some embodiments, for a 4×4 luma block, a 6-tap cosine window function difference filter can be used for MC interpolation with reference downsampling. Figure 23 The following Table 23 and Figure 24 Table 24 show examples of filter coefficients for ratios of 2:1 and 1.5:1.
[0136] In VVC, different interpolation filters are used in different motion compensation cases. For example, in conventional motion compensation, an 8-tap DCT-based interpolation filter is used. In the motion compensation of 4×4 sub-blocks, a 6-tap DCT-based interpolation filter is used. In the motion compensation with MVD precision of 1 / 2 and MVD phase of 1 / 2, different 6-tap filters are used. In chrominance motion compensation, a 4-tap DCT-based interpolation filter is used.
[0137] In some embodiments, when the reference picture has a higher resolution than the current coded picture, the DCT-based interpolation filter can be replaced with a cosine window function difference filter with the same filter length. In these embodiments, the 8-tap DCT-based interpolation filter can be replaced with an 8-tap cosine window function difference filter for conventional motion compensation; the 6-tap DCT-based interpolation filter can be replaced with a 6-tap cosine window function difference filter for 4×4 sub-block motion compensation. The 4-tap DCT-based interpolation filter can be replaced with a 4-tap cosine window function difference filter for chrominance motion compensation. In addition, when the MVD precision is 1 / 2 and the phase is 1 / 2, a 6-tap DCT-based interpolation filter can be used for motion compensation. The filter selection may depend on the ratio between the reference picture resolution and the current picture resolution. Exemplary 8-tap, 6-tap, and 4-tap cosine window function difference filters for 2:1 and 1.5:1 downsampling rates are as shown in Figures 19 - 24 Table 19-24. As shown in Figures 25 - 30 Table 25-30, the filters are symmetric in complementary phases.
[0138] The filter coefficients obtained in Equation (4) are real values. After normalization and scaling, the filter coefficients can be rounded to integers. The sum of the filter coefficients, also known as the filter gain, represents the filter precision. In digital processing, the filter gain is typically set to 2^N. Due to the rounding operation, the sum of the filter coefficients may not be equal to the filter gain (2^N). Additional adjustment is made to the filter coefficients such that the sum of the filter coefficients equals the gain (2^N). In some embodiments, the above adjustment includes determining an appropriate rounding direction, which may include rounding up or rounding down. Thus, the adjusted filter coefficients can be summed to the filter gain to minimize / maximize the cost function. In some embodiments, the cost function can be the sum of absolute differences (SAD) or the sum of squared errors (SSE) between the rounded filter coefficients and the true filter coefficients before rounding. In some embodiments, the cost function can be the SAD or SSE between the frequency response of the rounded filter coefficients and the frequency response of a reference value. The reference value can be the true filter coefficients before rounding, or the frequency response of an ideal filter, or the frequency response of a filter longer than the designed filter. For example, when designing an 8-tap filter, the reference may be a 12-tap filter. Based on the above rounding method, another set of cosine window function difference filters with lengths of 8 taps, 6 taps, and 4 taps are as shown in Figures 31 - 36 Tables 31 - 36.
[0139] Since in adaptive resolution change applications, the downsampling of the source image and the upsampling of the reconstructed image processes may not be defined in the standard, which means that the user can use any resampling filter. Since the filter selected by the user may not match the reference resampling filter, the performance may be degraded due to filter mismatch. In some disclosed embodiments, the coefficients of the reference resampling filter are marked in the bitstream. And the encoder / decoder can use the user-defined filter for reference resampling. The signaling of the filter coefficients is as shown in Figure 37 and 38 Tables 37 - 38.
[0140] The syntax structures exemplified in Tables 37 and 38 can be presented in high-level syntax, such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice headers, etc.
[0141] The use of the user-defined RPR filter and the presence of the above exemplary syntax structures can be controlled by flags in the high-level syntax, including SPS, PPS, APS, slice headers, etc.
[0142] Consistent with the disclosed embodiments, when one or more marked filters use a fixed filter length and filter norm, a hardware implementation may be more efficient. Thus, in some embodiments, the filter norm and / or filter taps may not be marked but implicit (e.g., implicit as default values defined in a standard). Filter specifications, filter taps, filter precision, or maximum bit depth may be part of the encoder compliance requirements.
[0143] In some embodiments, the encoder may send a Supplemental Enhancement Information (SEI) message to suggest what non-standard upsampling filter the receiving device should use after decoding a picture. Such SEI messages are optional. However, with an SEI message, a decoder using the suggested upsampling filter can obtain better video quality. The filter coefficient() syntax structure in Table 31( Figure 31 ) can be used to specify the upsampling filter suggested in such an SEI message.
[0144] Figure 39 is a flowchart of an exemplary method 3900 for processing video content consistent with embodiments of the present application. In some embodiments, method 3900 may be performed by a codec (e.g., using the encoder of the encoding process 200A or 200B in Figures 2A - 2B or the decoder of the decoding process 300A or 300B in Figures 3A - 3B ). For example, a codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or transcoding a video sequence. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a decoded compressed video sequence (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence, which may be captured by a surveillance device (e.g., the video input device in Figure 4 ) associated with a processor (e.g., processor 402) of the device. The video sequence may include multiple pictures. The device may perform method 3900 at the picture level. For example, in method 3900, the device may process one picture at a time. For example, in method 3900, the device may process multiple pictures at a time. Method 39 includes the following steps.
[0145] In step 3902, for a target picture and a reference picture with different resolutions, a band-pass filter may be applied to the reference picture to perform motion compensation interpolation using reference downsampling to generate a reference block.
[0146] The band - pass filter can be a cosine window function difference filter generated based on an ideal low - pass filter and a window function. For example, the kernel function f(n) of the cosine window function difference filter can be the product of the ideal low - pass filter h(n) and the window function w(n), which can be obtained based on formula (4).
[0147] Therefore, where fc is the cut - off frequency of the cosine window function difference filter f(n), L is the kernel length, and r is the down - sampling rate of the reference down - sampling.
[0148] The filter coefficients of the cosine window function difference filter can be rounded off before application. Figure 40 is a flowchart of an exemplary method 4000 for rounding the filter coefficients of a cosine window function difference filter that is consistent with the embodiments of the present application. It can be understood that method 4000 can be implemented independently or as part of method 3900. Method 4000 may include the following steps.
[0149] In step 4002, the true filter coefficients of the band - pass filter can be obtained.
[0150] In step 4004, multiple rounding directions of the true filter coefficients can be determined respectively.
[0151] In step 4006, by rounding the true filter coefficients according to the multiple rounding directions respectively, multiple combinations of the rounded filter coefficients can be generated.
[0152] In step 4008, among the multiple combinations, the combination of the rounded filter coefficients that minimizes or maximizes the cost function can be selected. In some embodiments, the cost function can be associated with the rounded filter coefficients and a reference value. The reference value can be the true filter coefficients of the band - pass filter before rounding, the frequency response of an ideal filter, or the frequency response of a filter longer than the band - pass filter. For example, when the reference value can be the true filter coefficients of the band - pass filter before rounding, the cost function can be the sum of absolute differences (SAD) or the sum of squared errors (SSE) between the rounded filter coefficients before rounding and the true filter coefficients. As another example, when the reference value is the frequency response of an ideal filter, the cost function can be the SAD or SSE between the frequency response of the rounded filter coefficients and the frequency response of the ideal filter. As described above, the reference value can also be the frequency response of a filter longer than the band - pass filter. For example, if the band - pass filter is an 8 - tap filter, the reference can be a 12 - tap filter, whose length is longer than the 8 - tap band - pass filter.
[0153] The selected combinations of the rounded filter coefficients of the bandpass filter can be marked, together with the syntax structures associated with the filter coefficients, in at least one of the sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), and slice header.
[0154] It should be understood that, as shown in Table 19-36 ( Figures 19 - 36 ), the cosine window function difference filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter. In some embodiments, based on the ratio between the resolution of the reference picture and the resolution of the target picture, the bandpass filter can be determined as one of an 8-tap filter, a 6-tap filter, and a 4-tap filter.
[0155] When applying the bandpass filter to the reference picture, luminance samples or chrominance samples can be obtained at fractional sampling positions. A lookup table using fractional sampling positions (e.g., Table 19-36) can be referred to determine the filter coefficients of the obtained luminance samples or chrominance samples.
[0156] Returning to Figure 39 , at step 3904, the reference block can be used to process the block of the target picture. For example, the block of the target picture can be encoded or decoded using the reference block.
[0157] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and these instructions can be executed by a device (e.g., the encoder and decoder disclosed in this application) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, floppy disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical medium with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, or any other flash memory, NVRAM, caches, registers, any other storage chip or cartridge memory, and the same network version. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0158] The above embodiments can be further described using the following clauses: 1. A computer-implemented method for performing motion compensation interpolation, comprising: Applying a bandpass filter to the reference picture for a target picture and a reference picture with different resolutions, and performing motion compensation interpolation by reference downsampling to generate a reference block; and Processing the block of the target picture using the reference block. 2. The method according to clause 1, wherein the bandpass filter is a cosine window function difference filter generated based on an ideal low-pass filter and a window function. 3. The method according to claim 2, wherein the cosine window function difference filter is based on a kernel function wherein where fc is the cut-off frequency of the cosine window function difference filter, L is the kernel length, and r is the downsampling rate of the reference downsampling. 4. The method according to claim 1, further comprising: obtaining the true filter coefficients of the band-pass filter; respectively determining a plurality of rounding directions of the true filter coefficients; generating a plurality of combinations of rounded filter coefficients by rounding the true filter coefficients according to the plurality of rounding directions; and selecting, from the plurality of combinations, the combination of the rounded filter coefficients that minimizes or maximizes a cost function. 5. The method according to claim 4, wherein the cost function is associated with the rounded filter coefficients and a reference value. 6. The method according to claim 5, wherein the reference value is the true filter coefficient of the band-pass filter before rounding, the frequency response of an ideal filter, or the frequency response of a filter longer than the length of the band-pass filter. 7. The method according to claim 2, wherein the cosine window function difference filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter. 8. The method according to claim 1, wherein applying the band-pass filter to the reference picture comprises: obtaining luminance samples or chrominance samples at fractional sampling positions. 9. The method according to claim 1, further comprising: determining the band-pass filter as one of an 8-tap filter, a 6-tap filter, and a 4-tap filter based on the ratio between the resolution of the reference picture and the resolution of the target picture. 10. The method according to claim 4, further comprising: marking the selected combination of the rounded filter coefficients of the band-pass filter in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header. 11. A system for performing motion compensation interpolation, comprising: a memory storing a set of instructions; and at least one processor configured to execute the set of instructions to cause the system to perform: applying a band-pass filter to the reference picture for target pictures and reference pictures having different resolutions, performing motion compensation interpolation by reference downsampling to generate a reference block; and Process the blocks of the target picture using the reference blocks. 12. The system according to claim 11, wherein the band - pass filter is a cosine window function difference filter generated based on an ideal low - pass filter and a window function. 13. The system according to claim 12, wherein the cosine window function difference filter is based on a kernel function wherein where fc is the cut - off frequency of the cosine window function difference filter, L is the kernel length, and r is the down - sampling rate of the reference down - sampling. 14. The system according to claim 11, wherein the at least one processor is configured to execute a set of instructions to cause the system to further perform: Obtain the true filter coefficients of the band - pass filter; Determine the multiple rounding directions of the true filter coefficients respectively; Generate multiple combinations of rounded filter coefficients by rounding the real filter coefficients according to the multiple rounding directions; and Among the multiple combinations, select the combination of rounded filter coefficients that minimizes or maximizes a cost function. 15. The system according to claim 14, wherein the cost function is associated with the rounded filter coefficients and a reference value. 16. The system according to claim 15, wherein The reference value is the true filter coefficient of the band - pass filter before rounding, the frequency response of an ideal filter, or the frequency response of a filter longer than the length of the band - pass filter. 17. The system according to claim 12, wherein the cosine window function difference filter is an 8 - tap filter, a 6 - tap filter, or a 4 - tap filter. 18. The system according to claim 11, wherein applying the band - pass filter to the reference picture includes: Obtain luminance samples or chrominance samples at fractional sampling positions. 19. The system according to claim 11, wherein the at least one processor is configured to execute a set of instructions to cause the system to further perform: Based on the ratio between the resolution of the reference picture and the resolution of the target picture, determine the band - pass filter as one of an 8 - tap filter, a 6 - tap filter, and a 4 - tap filter. 20. A non - transitory computer - readable medium storing a set of instructions that can be executed by at least one processor of a computer system to cause the computer system to execute a method for processing video content, the method comprising: In response to a target picture and a reference picture having different resolutions, a band-pass filter is applied to the reference picture, and motion compensation interpolation is performed through reference downsampling to generate a reference block; and Process the blocks of the target picture using the reference block.
[0159] It should be noted that relational terms such as "first" and "second" in this document are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, words such as "comprising", "having", "containing", and "including" and other similar forms have the same meaning and are open-ended, because one or more items after any of these words are not intended to exhaustively list such items or be limited to the listed items.
[0160] As used herein, unless otherwise expressly stated, the term "or" encompasses all possible combinations, unless infeasible. For example, if it is stated that a component may include A or B, then unless otherwise expressly stated or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless otherwise expressly stated or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0161] It can be understood that the above embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When executed by a processor, the software can execute the disclosed method. The computing units and other functional units described in this application can be implemented by hardware, or by software, or by a combination of hardware and software. Those of ordinary skill in the art can also understand that the above multiple modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0162] In the foregoing specification, embodiments have been described with reference to many specific details, which may vary with different embodiments. Certain adjustments and modifications can be made to the described embodiments. For those skilled in the art, other embodiments considering the specifications and practices of this application disclosed herein are obvious. The above specification and embodiments are only regarded as examples, and the true scope and spirit of this application are indicated by the claims. The order of steps shown in the figures is also intended for illustrative purposes only and is not intended to be limited to any specific order of steps. Therefore, those skilled in the art can understand that these steps can be executed in a different order while implementing the same method.
[0163] Exemplary embodiments have been disclosed in the accompanying drawings and the specification. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are for general and descriptive purposes only and not for the purpose of limitation.
Claims
1. A computer-implemented method for processing video content, comprising: encoding a block of an image by applying one or more filters to a reference image; wherein applying the one or more filters generates luminance samples at fractional sample positions, and wherein the one or more filters include an 8-tap filter that has a plurality of coefficients for each 1 / 16 fractional sample position, and wherein the coefficients associated with a fractional sample position are {-4, 2, 20, 28, 20, 2, -4, 0}.
2. The method according to claim 1, wherein the one or more filters include a 4-tap filter that has a plurality of coefficients [p0, p1, p2, p3] for each edge fractional sample position p, as follows:
3. The method according to claim 1, characterized in that the plurality of coefficients are signaled in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header.
4. The method according to claim 1, characterized in that applying the one or more filters includes: selecting the one or more filters based on the sampling rate of the reference image.
5. A computer-implemented method for processing video content, comprising: decoding a block of an image by applying one or more filters to a reference image; wherein applying the one or more filters generates luminance samples at fractional sample positions, and wherein the one or more filters include an 8-tap filter that has a plurality of coefficients for each 1 / 16 fractional sample position, and wherein the coefficients associated with a fractional sample position are {-4, 2, 20, 28, 20, 2, -4, 0}.
6. The method according to claim 5, wherein the one or more filters include a 4-tap filter that has a plurality of coefficients [p0, p1, p2, p3] for each edge fractional sample position p, as follows:
7. The method according to claim 5, characterized in that the plurality of coefficients are signaled in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header.
8. The method according to claim 5, characterized in that applying the one or more filters includes: selecting the one or more filters based on the sampling rate of the reference image.
9. A non-transitory computer-readable storage medium storing computer instructions and a bitstream, the computer instructions being executed by a processor to implement the following method to generate the bitstream, the method comprising: encoding a block of an image by applying one or more filters to a reference image; Wherein, applying the one or more filters generates luminance samples at fractional sample positions, and wherein the one or more filters include an 8-tap filter, the 8-tap filter having a plurality of coefficients for each 1 / 16 fractional sample position, wherein the coefficients associated with a fractional sample position are {-4, 2, 20, 28, 20, 2, -4, 0}.
10. The non-transitory computer-readable storage medium according to claim 9, wherein The method further includes: Applying the one or more filters includes: selecting the one or more filters based on the sampling rate of the reference image.