Filters for motion compensation interpolation with reference down-sampling

Motion compensation interpolation with reference resampling using a band-pass filter addresses high bandwidth and storage issues in high-definition video by enhancing compression efficiency, reducing bit rate and storage needs.

JP2025111417APending Publication Date: 2025-07-30ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025041986
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-27
Filing Date
2025-03-17
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

High bandwidth and storage requirements for high-definition video applications, particularly in video surveillance, due to the high bit rate of I pictures and the inefficiency of existing video coding techniques.

Method used

Implementing motion compensation interpolation with reference resampling using a band-pass filter to process video content, allowing for downsampling of reference blocks to generate a reference block based on target and reference pictures with different resolutions.

Benefits of technology

Reduces the bit rate and storage requirements for high-definition video by improving compression efficiency, enabling effective processing of high-resolution video data without significant quality degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111417000001_ABST
    Figure 2025111417000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for processing video content using motion compensation interpolation.SOLUTION: The methods include: in response to a target picture and a reference picture having different resolutions, applying a band-pass filter to the reference picture in order to perform motion compensation interpolation with reference down-sampling to generate a reference block; and encoding or decoding a block of the target picture using the reference block.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 904,608, filed on September 23, 2019, and U.S. Provisional Patent Application No. 62 / 906,930, filed on September 27, 2019, both of which are hereby incorporated by reference in their entirety.

[0002] Technical Field

[0002] This disclosure generally relates to video processing, and more particularly, to filters for motion - compensated interpolation with reference resampling.

Background Art

[0003] Background

[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. Most commonly, there are various video coding formats that use standardized video coding techniques based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify a particular video coding format, have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes higher.

Summary of the Invention

Means for Solving the Problems

[0004] Summary of the Disclosure

[0004] Embodiments of the present disclosure provide a computer-implemented method for performing motion compensation interpolation with reference resampling. The method may include applying a band-pass filter to a reference picture and processing a block of a target picture using the reference block to perform motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions.

[0005]

[0005] Embodiments of the present disclosure further provide a system for performing motion compensation interpolation. The system may include a memory storing a set of instructions and at least one processor configured to execute the set of instructions to cause the apparatus to apply a band-pass filter to a reference picture and process a block of a target picture using the reference block to perform motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions.

[0006]

[0006] Embodiments of the present disclosure further provide a non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer system to cause the computer system to execute a method for processing video content. The method may include applying a band-pass filter to a reference picture and processing a block of a target picture using the reference block to perform motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions.

[0007] Brief Description of the Drawings

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the figures are not drawn to scale.

Brief Description of the Drawings

[0008]

Figure 1

[0008] Shows the structure of an exemplary video sequence that conforms to an embodiment of the present disclosure.

Figure 2A

[0009] Shows a schematic diagram of an exemplary encoding process of a hybrid video coding system that conforms to an embodiment of the present disclosure.

Figure 2B

[0010] Shows a schematic diagram of another exemplary encoding process of a hybrid video coding system that conforms to an embodiment of the present disclosure.

Figure 3A

[0011] Shows a schematic diagram of an exemplary decoding process of a hybrid video coding system that conforms to an embodiment of the present disclosure.

Figure 3B

[0012] Shows a schematic diagram of another exemplary decoding process of a hybrid video coding system that conforms to an embodiment of the present disclosure.

Figure 4

[0013] Is a block diagram of an exemplary device for encoding or decoding video that conforms to an embodiment of the present disclosure.

Figure 5

[0014] Shows a schematic diagram of a reference picture and a current picture that conform to an embodiment of the present disclosure.

Figure 6

[0015] Shows a table of exemplary 6-tap DCT-based interpolation filter coefficients for a 4×4 luma block that conforms to an embodiment of the present disclosure.

Figure 7

[0016] Shows a table of exemplary 8-tap interpolation filter coefficients for the luma component that conforms to an embodiment of the present disclosure.

Figure 8

[0017] Shows a table of exemplary 4-tap 32-phase interpolation filter coefficients for the chroma component that conforms to an embodiment of the present disclosure.

Figure 9

[0018] Shows the frequency response of an exemplary ideal low-pass filter that conforms to an embodiment of the present disclosure.

Figure 10

[0019] Shows a table of exemplary 12-tap cosine window sinc interpolation filter coefficients for 2:1 downsampling that conforms to an embodiment of the present disclosure.

Figure 11

[0020] Table showing exemplary 12 - tap cosine window sinc interpolation filter coefficients for 1.5:1 downsampling that conforms to an embodiment of the present disclosure.

Figure 12

[0021] Shows an exemplary luma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 13

[0022] Shows an exemplary chroma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 14

[0001] Shows an exemplary luma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 15

[0002] Shows an exemplary chroma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 16

[0003] Shows an exemplary luma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 17

[0004] Shows an exemplary chroma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 18

[0005] Shows an exemplary chroma sample interpolation filtering process for reference downsampling that conforms to an embodiment of the present disclosure.

Figure 19

[0006] Shows an exemplary 8 - tap filter for MC interpolation with reference downsampling at a ratio of 2:1.

Figure 20

[0007] Shows an exemplary 8 - tap filter for MC interpolation with reference downsampling at a ratio of 1.5:1.

Figure 21

[0008] Illustrative 8-tap filters for MC interpolation with reference downsampling at a ratio of 2:1 with 32 phases, consistent with embodiments of the present disclosure.

Figure 22

[0009] Illustrative 8-tap filters for MC interpolation with reference downsampling at a ratio of 1.5:1 with 32 phases, consistent with embodiments of the present disclosure.

Figure 23

[0010] Table showing exemplary 6-tap filter coefficients for 4×4 luma block MC interpolation with a reference downsampling ratio of 2:1 with 16 phases, consistent with embodiments of the present disclosure.

Figure 24

[0011] Table showing exemplary 6-tap filter coefficients for luma 4×4 block MC interpolation with a reference downsampling ratio of 1.5:1 with 16 phases, consistent with embodiments of the present disclosure.

Figure 25

[0012] Table showing exemplary 8-tap filters for MC interpolation with reference downsampling at a ratio of 2:1, consistent with embodiments of the present disclosure.

Figure 26

[0013] Illustrative 8-tap filters for MC interpolation with reference downsampling at a ratio of 1.5:1, consistent with embodiments of the present disclosure.

Figure 27

[0014] Illustrative 6-tap filters for MC interpolation with reference downsampling at a ratio of 2:1 with 16 phases, consistent with embodiments of the present disclosure.

Figure 28

[0015] Illustrative 6-tap filters for MC interpolation with reference downsampling at a ratio of 1.5:1 with 16 phases, consistent with embodiments of the present disclosure.

Figure 29

[0016] Illustrative 4-tap filters for MC interpolation with reference downsampling at a ratio of 2:1 with 32 phases, consistent with embodiments of the present disclosure.

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

[0023] Illustrative 4-tap filter for MC interpolation with reference downsampling at a ratio of 1.5:1 having 32 phases, in accordance with an embodiment of the present disclosure.

Figure 37

[0024] An example of signaling of filter coefficients, in accordance with an embodiment of the present disclosure.

Figure 38

[0025] Illustrative syntax structure for signaling resampling ratio and corresponding filter set, in accordance with an embodiment of the present disclosure.

Figure 39

[0026] Flowchart of an illustrative method for processing video content, in accordance with an embodiment of the present disclosure.

Figure 40

[0027] A flowchart of an exemplary method for rounding the filter coefficients of a cosine window sinc filter that conforms to an embodiment of the present disclosure.

DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description

[0028] Here, reference is made in detail to the exemplary embodiments shown in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the exemplary embodiments do not represent all implementations that conform to the present invention. Rather, they are merely examples of devices and methods that conform to aspects related to the present invention recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0010]

[0029] Video coding systems are often used to compress digital video signals, for example, to reduce the storage space consumed or to reduce the consumption of the transmission bandwidth associated with such signals. In various applications of video compression such as online video streaming, video conferencing, or video surveillance, as the popularity of high-definition (HD) video (e.g., having a resolution of 1920×1080 pixels) increases, there is a continuous demand to develop video coding tools that can improve the compression efficiency of video data.

[0011]

[0030] For example, the application of video surveillance is being used more extensively and widely in many application scenarios (such as security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are increasing rapidly. Many video surveillance application scenarios choose to provide users with HD video in order to capture more information, and HD video has more pixels per frame in order to capture such information. However, an HD video bitstream can have a high bitrate that requires high bandwidth for transmission and large space for storage. For example, a surveillance video stream with an average resolution of 1920×1080 may require a bandwidth of 4 Mbps for real-time transmission. Furthermore, video surveillance generally performs round-the-clock monitoring, which can be a major challenge for the storage system when storing video data. Therefore, the demand for high bandwidth and large storage space of HD video has become the main limitation for the large-scale deployment of HD video in video surveillance.

[0012]

[0031] Video is a set of still pictures (or "frames") arranged in chronological order for storing visual information. A video capture device (such as a camera) can be used to capture and store those pictures in chronological order, and a video playback device (such as a TV, computer, smartphone, tablet computer, video player, or any end-user terminal with a display function) can be used to display such pictures in chronological order. Furthermore, in some applications, for surveillance, meetings, live broadcasts, etc., the video capture device can transmit the captured video to the video playback device (such as a computer with a monitor) in real time.

[0013]

[0032] To reduce the memory space and transmission bandwidth required by such applications, the video can be compressed before being stored and transmitted and decompressed before being displayed. This compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. A module for compression is generally called an "encoder", and a module for decompression is generally called a "decoder". The encoder and decoder can be collectively called a "codec". The encoder and decoder can be implemented as various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or algorithms or processes implemented by any suitable computer fixed in a computer-readable medium. The compression and decompression of video can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder".

[0014]

[0033] The video encoding process can identify and retain the useful information that can be used to reconstruct the picture and ignore the information that is not important for reconstruction. If the unimportant information that is ignored cannot be fully reconstructed, such an encoding process can be called "irreversible". Otherwise, such an encoding process can be called "reversible". Most encoding processes are irreversible, which is a trade-off to reduce the required memory space and transmission bandwidth.

[0015]

[0034] The useful information of the symbolized picture (referred to as the "current picture") includes the changes with respect to the reference picture (for example, a picture that was encoded and reconstructed in the past). Such changes can include changes in pixel position, luminance, or color, among which the position change is the most important. The position change of the group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.

[0016]

[0035] A picture that is encoded without referring to another picture (i.e., such a picture is its own reference picture) is called an "I picture". A picture that is encoded using a past picture as a reference picture is called a "P picture". A picture that is encoded using both a past picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".

[0017]

[0036] As described above, video surveillance using HD video faces the problems of high bandwidth and large storage requirements. To address this problem, the bit rate of the encoded video can be reduced. Among I pictures, P pictures, and B pictures, I pictures have the highest bit rate. Since the background of most surveillance videos is almost static, one way to reduce the overall bit rate of the encoded video may be to use fewer I pictures for video encoding.

[0018]

[0037] However, in the encoded video, since I pictures are generally not the main ones, the improvement measure of using fewer I pictures may be minor. For example, in a typical video bitstream, the ratio of I pictures, B pictures, and P pictures may be 1:20:9, and the I pictures may account for less than 10% of the total bit rate. In other words, in such an example, even if all I pictures are removed, the reduced bit rate may be only 10%.

[0019]

[0038] FIG. 1 shows the structure of an example of a video sequence 100 that conforms to an embodiment of the present disclosure. The video sequence 100 can be a live relay video or a captured and archived video. The video 100 can be a real video, a video generated by a computer (e.g., a computer game video), or a combination thereof (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing videos captured in the past (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.

[0020]

[0039] As shown in FIG. 1, the video sequence 100 can include a series of pictures that are temporally arranged along a timeline including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, and there are more pictures between picture 106 and picture 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and as indicated by the arrow, its reference picture is picture 102. Picture 106 is a B picture, and as indicated by the arrow, its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be the picture immediately before or after that picture. For example, the reference picture of picture 104 can be a picture preceding picture 102. The reference pictures of pictures 102-106 are merely examples, and it should be noted that the present disclosure does not limit the embodiments of the reference pictures to the examples shown in FIG. 1.

[0021]

[0040] Typically, a video codec does not encode or decode an entire picture at once because such a task is computationally complex. Instead, a video codec can divide a picture into basic segments and encode or decode the picture segment by segment. In the present disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, the structure 110 in FIG. 1 shows an example of the structure of a picture (e.g., any one of pictures 102-108) of the video sequence 100. In the structure 110, the picture is divided into 4×4 basic processing units, and the boundaries thereof are indicated by dashed lines. In some embodiments, the basic processing unit can be referred to as a "macroblock" within some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), and can be referred to as a "coding tree unit" ("CTU") within some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size within the picture, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail to be maintained within the basic processing unit.

[0022]

[0041] The basic processing unit can be a logical unit that can include various types of video data groups stored in a computer memory (e.g., within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements, and the luma component and chroma components can have the same-sized basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma component and chroma components can be referred to as "coding tree blocks" ("CTBs"). Any operation performed on the basic processing unit can be repeatedly performed on each of its luma component and chroma components.

[0023]

[0042] The encoding of the video has multiple operation stages, examples of which are detailed in FIGS. 2A-2B and FIGS. 3A-3B. For each stage, the size of the basic processing unit may still be too large to process, and thus can be further divided into segments called "basic processing sub-units" in the present disclosure. In some embodiments, the basic processing sub-unit can be called a "block" within some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or can be called a "coding unit" ("CU") within other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit can have a size equal to or smaller than that of the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit that can include various types of video data groups (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., within a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed at further levels according to the need for processing. It should also be noted that the various stages can divide the basic processing unit using various methods.

[0024]

[0043] For example, (an example of which is detailed in FIG. 2B) in the mode determination stage, the coder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, and the basic processing unit may be too large to make such a decision. The coder can divide the basic processing unit into a plurality of basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.

[0025]

[0044] In another example, (one example of which is detailed in FIG. 2A) in the prediction stage, the coder can perform a prediction operation at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform the prediction operation at that level.

[0026]

[0045] In another example, (one example of which is detailed in FIG. 2A) in the transformation stage, the coder can perform a transformation operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transformation operation at that level. It should be noted that the splitting method for the same basic processing subunit can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.

[0027]

[0046] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and its boundaries are indicated by dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.

[0028]

[0047] In some embodiments, in order to provide parallel processing and error resilience functions for video encoding and decoding, pictures can be divided into regions for processing, so that for the regions of a picture, the encoding or decoding process can be made independent of the information of any other region of the picture. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of a picture in parallel, thus improving the encoding efficiency. Further, if the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, thus providing an error resilience function. In some video coding standards, pictures can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions, namely "slices" and "tiles". It should also be noted that the various pictures of video sequence 100 may have various splitting methods for dividing the pictures into regions.

[0029]

[0048] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, and its boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are only examples, and the present disclosure does not limit its embodiments.

[0030]

[0049] FIG. 2A shows a schematic diagram of an example of an encoding process 200A that conforms to an embodiment of the present disclosure. For example, the encoding process 200A can be executed by an encoder. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in temporal order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can execute process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder can execute process 200A in an iterative manner, and the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can execute process 200A in parallel for the regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0031]

[0050] In FIG. 2A, the coder can feed the basic processing unit of the original picture of video sequence 202 (referred to as the "original BPU") to prediction stage 204 to generate prediction data 206 and predicted BPU 208. The coder can subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The coder can feed the residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The coder can feed the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path". During process 200A, after quantization stage 214, the coder can feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224 for use in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used to ensure that both the coder and the decoder use the same reference data for prediction.

[0032]

[0051] The coder can iteratively execute process 200A to encode each original BPU of the original picture (within the forward path) and generate predicted reference 224 for encoding the next original BPU of the original picture (within the reconstruction path). After encoding all the original BPUs of the original picture, the coder can proceed to encode the next picture in video sequence 202.

[0033]

[0052] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive (s)" can refer to receiving for the purpose of inputting data, inputting, obtaining, retrieving, getting, reading out, accessing, or any action of any method.

[0034]

[0053] In the prediction stage 204, in the current iteration, the coder can receive the original BPU and prediction criterion 224 and perform a prediction operation to generate prediction data 206 and predicted BPU 208. The prediction criterion 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of the prediction stage 204 is to reduce the redundancy of information by extracting the prediction data 206, and the prediction data 206 can be used to reconstruct the original BPU as the predicted BPU 208 from the prediction data 206 and the prediction criterion 224.

[0035]

[0054] Ideally, the predicted BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such a difference, after generating the predicted BPU 208, the coder can subtract it from the original BPU to generate a residual BPU 210. For example, the coder can subtract the pixel value (e.g., grayscale value or RGB value) of the predicted BPU 208 from the corresponding pixel value of the original BPU. As a result of such subtraction between the corresponding pixel of the original BPU and the predicted BPU 208, each pixel of the residual BPU 210 can have a residual value. Compared with the original BPU, the prediction data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality.

[0036]

[0055] To further compress the residual BPU 210, in the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component of the variation of the residual BPU 210 (e.g., the luminance variation frequency). None of the basis patterns can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.

[0037]

[0056] Various transformation algorithms can use various basis patterns. For example, in the transformation stage 212, various transformation algorithms such as the discrete cosine transform, the discrete sine transform, etc. can be used. The transformation in the transformation stage 212 is reversible. That is, the coder can restore the residual BPU 210 by the inverse operation of the transformation (referred to as "inverse transformation"). For example, to restore the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basis patterns by their respective associated coefficients and add the products to yield a weighted sum. In the video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients that can reconstruct the residual BPU 210 from which the decoder can reconstruct the residual BPU 210 without receiving the basis patterns from the coder. Compared with the residual BPU 210, the transformation coefficients may have fewer bits, but those transformation coefficients can be used to reconstruct the residual BPU 210 without significantly degrading the quality. Therefore, the residual BPU 210 is further compressed.

[0038]

[0057] The coder can further compress the transform coefficients in the quantization stage 214. In the transform process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the coder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the coder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as the "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some of the transform coefficients of the high-frequency basis pattern can be converted to zero, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients within the inverse operation of quantization (referred to as "inverse quantization").

[0039]

[0058] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be non-reversible. Typically, the quantization stage 214 can contribute to the largest information loss within the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain various levels of information loss, the coder can use various values of the quantization parameter or any other parameter of the quantization process.

[0040]

[0059] In the binary coding stage 226, the coder can encode the predicted data 206 and the quantized transform coefficients 216 using binary coding techniques such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the coder can encode other information such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transformation in the transformation stage 212, the parameters of the quantization process (e.g., quantization parameters), coder control parameters (e.g., bitrate control parameters), etc. in the binary coding stage 226. The coder can generate a video bitstream 228 using the output data of the binary coding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0041]

[0060] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the coder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transformation stage 220, the coder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction criterion 224 used in the next iteration of process 200A.

[0042]

[0061] It should be noted that other variations of process 200A can be used to encode the video sequence 202. In some embodiments, the encoder can execute the steps of process 200A in a different order. In some embodiments, one or more steps of process 200A can be combined into a single step. In some embodiments, a single step of process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. In some embodiments, process 200A can include additional steps. In some embodiments, process 200A can omit one or more steps in FIG. 2A.

[0043]

[0062] FIG. 2B shows a schematic diagram of another example 200B of an encoding process that conforms to an embodiment of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision step 230 and divides the prediction step 204 into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of process 200B additionally includes a loop filter step 232 and a buffer 234.

[0044]

[0063] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use the pixels of one or more already encoded adjacent BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions of one or more already encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0045]

[0064] Referring to process 200B, in the forward path, the coder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the coder can perform intra prediction. For the original BPU of the coded picture, the prediction reference 224 may include one or more adjacent BPUs that are coded (within the forward path) and reconstructed (within the reconstruction path) within the same picture. The coder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the coder can perform extrapolation at the pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation may be located relative to the original BPU in various directions, such as the vertical direction (e.g., above the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined within the video coding standard being used. In intra prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used relative to the original BPU, etc.

[0046]

[0065] In another example, in the temporal prediction stage 2044, the coder can perform inter prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that are encoded (within the forward path) and reconstructed (within the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the coder can generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the predicted BPU 208. When all the reconstructed BPUs of the same picture are generated, the coder can generate a picture reconstructed as a reference picture. The coder can perform an operation of "motion estimation" to find a matching region within the range of the reference picture (referred to as the "search window"). The position of the search window within the reference picture can be determined based on the position of the original BPU within the current picture. For example, the search window can be centered at a position having the same coordinates as the original BPU within the current picture within the reference picture and can be expanded over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (for example, by using a pel recursive algorithm, a block matching algorithm, etc.), the coder can determine that region as the matching region. The matching region may have dimensions different from (for example, smaller than, equal to, larger than, or of a different shape than) the original BPU. Since the reference picture and the current picture are temporally separated within the timeline (as shown in FIG. 1, for example), it can be considered that the matching region "moves" to the position of the original BPU as time passes. The coder can record such a direction and distance of motion as a "motion vector". When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can search for a matching region for each reference picture and obtain its associated motion vector. In some embodiments, the coder can assign weights to the pixel values of the matching regions of the individual matching reference pictures.

[0047]

[0066] Motion estimation can be used, for example, to identify various types of motion such as translation, rotation, scaling, etc. In inter prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the matching region, the motion vector related to the matching region, the number of reference pictures, the weights related to the reference pictures, and the like.

[0048]

[0067] To generate the predicted BPU 208, the coder can perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the coder can move the matching region of the reference picture according to the motion vector, in which the coder can predict the original BPU of the current picture. When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can move the matching regions of the reference pictures according to the individual motion vectors and average the pixel values of the matching regions. In some embodiments, when the coder assigns weights to the pixel values of the matching regions of the individual matching reference pictures, the coder can add the weighted sum of the pixel values of the moved matching regions.

[0049]

[0068] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction with respect to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter prediction picture where the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions with respect to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.

[0050]

[0069] Continuing to refer to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform rate distortion optimization techniques, in which the coder can select a prediction mode to minimize the value of a cost function according to the bit rate of candidate prediction modes and the distortion of the reconstructed reference pictures under the candidate prediction modes. Depending on the selected prediction mode, the coder can generate the corresponding predicted BPU 208 and predicted data 206.

[0051]

[0070] In the reconstruction path of process 200B, when the intra prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current BPU that is encoded and reconstructed within the current picture), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). When the inter prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking artifacts) caused by inter prediction. For example, the encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as an inter prediction reference picture for future pictures of video sequence 202). The encoder can store one or more reference pictures in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode loop filter parameters (e.g., loop filter strength) at the binary coding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.

[0052]

[0071] FIG. 3A shows a schematic diagram of an example of a decoding process 300A that conforms to an embodiment of the present disclosure. Process 300A can be a decompression process corresponding to the compression process 200A of FIG. 2A. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 can be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), generally, video stream 304 is not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, the decoder can execute process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder can execute process 300A in an iterative manner, and the decoder can decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder can execute process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0053]

[0072] In FIG. 3A, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded picture (referred to as an “encoded BPU”) to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode a portion thereof into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can feed the predicted reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.

[0054]

[0073] The decoder can repeatedly execute process 300A to decode each encoded BPU of the encoded picture and generate a predicted reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0055]

[0074] In the binary decoding stage 302, the decoder can perform the reverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible compression algorithm). In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of transform, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, when the video bitstream 228 is transmitted in packet units over the network, the decoder can depacketize the video bitstream 228 and then feed it to the binary decoding stage 302.

[0056]

[0075] FIG. 3B shows a schematic diagram of another example 300B of the decoding process that conforms to an embodiment of the present disclosure. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0057]

[0076] In Process 300B, for an encoded basic processing unit of a decoded encoded picture (referred to as the "current picture") (referred to as the "current BPU"), the prediction data 206 decoded by the decoder in the binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as a reference, the size of the adjacent BPUs, extrapolation parameters, the direction of the adjacent BPUs with respect to the original BPU, etc. In another example, if inter prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation may include, for example, the number of reference pictures related to the current BPU, the weights respectively related to the reference pictures, the positions (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively related to the matching regions, etc.

[0058]

[0077] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of the execution of such spatial or temporal prediction are shown in Figure 2B and will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate the predicted BPU 208. As described in Figure 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction reference 224.

[0059]

[0078] In process 300B, the decoder can feed the predicted reference 224 for performing a prediction operation within the next iteration of process 300B to the spatial prediction stage 2042 or the temporal prediction stage 2044. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the coder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture can be stored in the buffer 234 (e.g., the decoded picture buffer in computer memory) for later use (e.g., for use as an inter prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU, the prediction data can further include loop filter parameters (e.g., the strength of the loop filter).

[0060]

[0079] FIG. 4 is a block diagram of an example of a device 400 for encoding or decoding video that conforms to an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 can become a dedicated machine for encoding or decoding video. The processor 402 can be any type of circuit capable of manipulating or processing information. For example, the processor 402 can include any combination of any number of central processing units ("CPUs"), graphics processing units ("GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 can be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 can include a plurality of processors including processor 402a, processor 402b, and processor 402n.

[0061]

[0080] Machine 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps within processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and the processing data (e.g., via bus 410), and execute the program instructions to perform operations or processing on the processing data. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a memory bank (not shown in FIG. 4) grouped as a single logical component.

[0062]

[0081] A bus 410, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc., can be a communication device that transfers data between components within the device 400.

[0063]

[0082] For simplicity of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuit can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuit can be a single independent module, or can be fully or partially combined within any other component of the device 400.

[0064]

[0083] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (NFC) adapters, cellular network chips, and the like.

[0065]

[0084] In some embodiments, device 400 may optionally further include a peripheral device interface 408 for providing connection to one or more peripheral devices. As shown in FIG. 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touch pad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), and the like.

[0066]

[0085] It should be noted that the video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. In another example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0067]

[0086] One of the important requirements of the VVC standard is to enable video conferencing applications to tolerate network and device diversity, quickly reduce the encoded bitrate when network conditions deteriorate, and quickly increase the video quality when network conditions improve, so as to be able to quickly adapt to changing network environments. The expected video quality can vary from very low to very high. In the case of an adaptive streaming service that provides multiple representations of the same content, each with different characteristics (e.g., spatial resolution or sample bit depth), the standard should also support fast representation switching. During the switch from one representation to another (such as switching from one resolution to another), the standard should be able to use an efficient prediction structure without sacrificing the fast and seamless switching function.

[0068]

[0087] The purpose of adaptive resolution change (ARC) is to enable the stream to change the spatial resolution between encoded pictures within the same video sequence without requiring a new IDR frame or multiple layers as in a scalable video codec. Instead, at the switching point, the picture changes the resolution and can be predicted from reference pictures with the same resolution (if any) and reference pictures with different resolutions. If the reference pictures are of different resolutions, the reference pictures are resampled as shown in Figure 5. For example, as shown in Figure 5, the resolution of reference picture 506 is the same as that of the current picture 502, but the resolutions of reference pictures 504 and 508 are different from the resolution of the current picture 502. After resampling reference pictures 504 and 508 to match the resolution of the current picture 502, motion compensation prediction can be performed from those references. Therefore, adaptive resolution change (ARC) is sometimes also called reference picture resampling (RPR), and in this disclosure, these two terms are used interchangeably without distinction.

[0069]

[0088] When the resolution of the reference frame is different from that of the current frame, one way to generate a motion-compensated prediction signal is picture-based resampling. In picture-based resampling, first the reference picture is resampled to the same resolution as the current picture, and then the existing motion compensation process with motion vectors can be applied. The motion vectors may be scaled (if transmitted in units before the application of resampling) or may not be scaled (if transmitted in units after the application of resampling). In picture-based resampling, especially for downsampling of the reference picture (i.e., when the resolution of the reference picture is larger than that of the current picture), information may be lost within the reference resampling step before motion compensation interpolation because downsampling is usually achieved by low-pass filtering followed by decimation).

[0070]

[0089] Another way is block-based resampling. In block-based resampling, resampling is performed at the block level. This is done by examining the reference pictures used by the current block, and if one or both of them have a different resolution from the current picture, resampling is performed in combination with a motion compensation interpolation process that is more precise than the 1-pixel unit.

[0071]

[0090] In block-based resampling, combining resampling and motion-compensated interpolation into one filtering operation can reduce the information loss described above. Take the following example. The motion vector of the current block has a precision of half a pixel unit in one dimension (e.g., the horizontal dimension), and the width of the reference picture is twice the width of the current picture. Compared to picture-level resampling that reduces the width of the reference picture by half to match the width of the current picture and performs half-pixel unit motion interpolation, in this case, the block-based resampling method can directly obtain odd positions in the reference picture as reference blocks with a precision of half a pixel unit. At the 15th JVET meeting, a block-based resampling method for ARC was adopted in VVC. In this method, motion compensation (MC) interpolation and reference resampling are combined and executed within one-step filtering. In VVC draft 6, the existing filter for MC interpolation without reference resampling is reused for MC interpolation with reference resampling. The same filter is used for both reference upsampling and reference downsampling. Details regarding filter selection will be described below.

[0072]

[0091] For the luma component, when the AMVR mode of half a pixel unit is selected and the interpolation position is in half-pixel units, a 6-tap filter [3, 9, 20, 20, 9, 3] is used. When the size of the motion-compensated block is 4×4, the following 6-tap filter as shown in Table 6 of Figure 6 is used. Otherwise, the 8-tap filter shown in Table 7 of Figure 7 is used.

[0073]

[0092] For the chroma component, the following 4-tap filter as shown in Table 8 of Figure 8 is used.

[0074]

[0093] In VVC, the same filter is used for both MC interpolation without reference resampling and MC interpolation with reference resampling. The motion compensation interpolation filter (MCIF) of VVC is designed based on DCT upsampling, but it may be inappropriate to use MCIF as a one-step filter that combines reference downsampling and MC interpolation. For example, in phase 0 filtering (e.g., when the scaled motion vector is an integer), the VVC 8-tap MCIF coefficients are [0, 0, 0, 64, 0, 0, 0, 0], which means that the predicted sample is directly copied from the reference sample. This may not be a problem in MC interpolation without reference downsampling or with reference upsampling, but in the case of reference downsampling, aliasing artifacts may occur due to the absence of a low-pass filter before decimation.

[0075]

[0094] The present disclosure provides a method of using a cosine windowed sinc filter for MC interpolation with reference downsampling.

[0076]

[0095] The windowed sinc filter is a band-pass filter that separates one frequency band from another. The windowed sinc filter is a low-pass filter having a frequency response that allows all frequencies below the cut-off frequency to pass through with an amplitude of 1 and stops all frequencies above the cut-off frequency with zero amplitude, as shown in FIG. 9.

[0077]

[0096] By taking the inverse Fourier transform of the frequency response of an ideal low-pass filter, a filter kernel, also known as the impulse response of the filter, is obtained. The impulse response of the low-pass filter is in the general form of a sinc function based on the following equation (1).

Equation

Equation

[0078]

[0097] The sinc function is infinite. To create the filter kernel within a finite length, a window function is applied to truncate the filter kernel at the L-th point. To obtain a smooth taper curve, a cosine window function based on the following formula (3) is used.

Equation

[0079]

[0098] The kernel of the cosine window sinc filter is the product of the ideal response function h(n) based on the following formula (4) and the cosine window function w(n).

Equation

[0080]

[0099] There are two parameters, the cut-off frequency fc and the kernel length L, selected for the window sinc kernel. By adjusting the values of L and fc, a desired filter response can be achieved. For example, in the downsampling filter used in the scalable HEVC test model (SHM), fc = 0.9 and L = 13.

[0081]

[0100] The filter coefficients obtained by Equation (4) are real numbers. Applying the filter is equivalent to calculating the weighted average of the reference samples with the weights being the filter coefficients. For efficient calculation in a digital computer or hardware, the coefficients are normalized, scaled by a scalar, and rounded to an integer such that the sum of the coefficients is equal to 2^N (where N is an integer). The filtered samples are divided by 2^N (corresponding to an N-bit right shift). For example, in VVC Draft 6, the sum of the interpolation filter coefficients is 64.

[0082]

[0101] In some embodiments, for both the luma and chroma components, a downsampling filter can be used within SHM for VVC motion compensation interpolation with reference downsampling, and the existing MCIF for motion compensation interpolation with reference upsampling can be used. While the kernel length L = 13, the first coefficient is small and rounded to zero, and the filter length can be reduced to 12 without affecting the filter performance.

[0083]

[0102] As an example, the filter coefficients for 2:1 downsampling and 1.5:1 downsampling are shown in Table 10 of FIG. 10 and Table 11 of FIG. 11, respectively.

[0084]

[0103] In addition to the values of the coefficients, there are several other differences between the design of the SHM filter and the existing MCIF.

[0085]

[0104] As the first difference, the SHM filter requires filtering at both integer and fractional sample positions, while the MCIF requires filtering only at fractional sample positions. An example of the modification to the luma sample interpolation filtering process of VVC Draft 6 for the case of reference downsampling is described in Table 12 of the attached FIG. 12.

[0086]

[0105] An example of the modification to the chroma sample interpolation filtering process of VVC Draft 6 regarding the case of reference downsampling is described in Table 13 of the attached Figure 13.

[0087]

[0106] As a second difference, the sum of the filter coefficients is 128 in the SHM filter, while the sum of the filter coefficients is 64 in the existing MCIF. In VVC Draft 6, in order to reduce the loss caused by rounding error, the intermediate prediction signal is kept at a higher precision than the output signal (represented by a higher bit depth). The precision of the intermediate signal is called the internal precision. In some embodiments, in order to keep the internal precision the same as that in VVC Draft 6, compared with using the existing MCIF, the output of the SHM filter needs to be right-shifted by one additional bit. An example of the modification to the luma sample interpolation filtering process of VVC Draft 6 regarding the case of reference downsampling is shown in Table 14 of the attached Figure 14.

[0088]

[0107] An example of the modification to the chroma sample interpolation filtering process of VVC Draft 6 regarding the case of reference downsampling is shown in Table 15 of the attached Figure 15.

[0089]

[0108] According to some embodiments, the internal precision can be increased by 1 bit, and the additional 1-bit right shift can be used to convert the internal precision to the output precision. An example of the modification to the luma sample interpolation filtering process of VVC Draft 6 regarding the case of reference downsampling is shown in Table 16 of Figure 16.

[0090]

[0109] An example of the modification to the chroma sample interpolation filtering process of VVC Draft 6 regarding the case of reference downsampling is shown in Table 17 of the attached Figure 17.

[0091]

[0110] As a third difference, the SHM filter has 12 taps. Therefore, to generate an interpolated sample, 11 adjacent samples (5 on the left and 6 on the right, or 5 above, or 6 below) are required. Compared with MCIF, additional adjacent samples are obtained. In VVC draft 6, the chroma mv accuracy is 1 / 32. However, the SHM filter has only 16 phases. Therefore, the chroma mv can be rounded to 1 / 16 for reference downsampling. This rounding can be done by right-shifting the last 5 bits of the chroma mv by 1 bit. An example of the correction to the chroma fractional sample position calculation in VVC draft 6 for the case of reference downsampling is shown in Table 18 of FIG. 18.

[0092]

[0111] According to some embodiments, it is proposed to use an 8-tap cosine window sinc filter to be consistent with the existing MCIF design in the VVC draft. The filter coefficients can be derived by setting L = 9 in the cosine window sinc function in Equation (4). To be further consistent with the existing MCIF filter, the sum of the filter coefficients can be set to 64. In the disclosed embodiments, the filters for complementary phases can be symmetric (e.g., the filter coefficients for complementary phases are inverted) or asymmetric. Examples of filter coefficients with ratios of 2:1 and 1.5:1 are shown in Table 19 of FIG. 19 and Table 20 of FIG. 20, respectively.

[0093]

[0112] According to some embodiments, to conform to the 1 / 32 sample accuracy of the chroma component, a 32-phase cosine window sinc filter set can be used for chroma motion compensation interpolation with reference downsampling. Examples of filter coefficients with ratios of 2:1 and 1.5:1 are shown in Table 21 of the attached FIG. 21 and Table 22 of FIG. 22, respectively.

[0094]

[0113] According to some embodiments, in a 4×4 luma block, a 6-tap cosine window sinc filter can be used for MC interpolation with reference downsampling. Examples of filter coefficients for ratios of 2:1 and 1.5:1 are shown in Table 23 of the attached FIG. 23 and Table 24 of FIG. 24, respectively.

[0095]

[0114] In VVC, various interpolation filters are used in various motion compensation cases. For example, in normal motion compensation, an 8-tap DCT-based interpolation filter is used. In motion compensation for a 4×4 sub-block, a 6-tap DCT-based interpolation filter is used. In motion compensation where the MVD accuracy is 1 / 2 and the MVD phase is 1 / 2, a different 6-tap filter is used. In chroma motion compensation, a 4-tap DCT-based interpolation filter is used.

[0096]

[0115] In some embodiments, when the reference picture has a higher resolution than the currently encoded picture, the DCT-based interpolation filter can be replaced with a cosine window sinc filter of the same filter length. In these embodiments, for normal motion compensation, the 8-tap DCT-based interpolation filter can be replaced with an 8-tap cosine window sinc filter, and for motion compensation of a 4×4 sub-block, the 6-tap DCT-based interpolation filter can be replaced with a 6-tap cosine window sinc filter. For chroma motion compensation, the 4-tap DCT-based interpolation filter can be replaced with a 4-tap cosine window sinc filter. Further, when the MVD accuracy is 1 / 2 and the phase is 1 / 2, a 6-tap DCT-based interpolation filter can be used for motion compensation. The selection of the filter may depend on the ratio between the resolution of the reference picture and the resolution of the current picture. Exemplary 8-tap, 6-tap, and 4-tap cosine window sinc filters for downsampling ratios of 2:1 and 1.5:1 are shown in Tables 19 - 24 of FIGS. 19 - 24. The filters shown in Tables 25 - 30 of FIGS. 25 - 30 are contrastive on complementary phases.

[0097]

[0116] The filter coefficients obtained by Equation (4) are real numbers. After normalization and scaling, the filter coefficients can be rounded to integers. The sum of the filter coefficients, also known as the filter gain, represents the accuracy of the filter. In digital processing, the filter gain is usually set to 2^N. Due to the rounding operation, the sum of the filter coefficients may not be equal to the filter gain (2^N). Further adjustment is performed on the filter coefficients so that the sum of the filter coefficients is equal to the gain (2^N). In some embodiments, the adjustment includes determining the appropriate rounding direction. The rounding direction may include rounding up or rounding down. Therefore, the adjusted filter coefficients can be summed to the filter gain to minimize / maximize the cost function. In some embodiments, the cost function can be the sum of absolute differences (SAD) or sum of squared errors (SSE) between the rounded filter coefficients and the actual filter coefficients before rounding. In some embodiments, the cost function can be the SAD or SSE between the frequency response of the rounded filter coefficients and the reference frequency response. The reference can be the actual filter coefficients before rounding, or the frequency response of an ideal filter, or the frequency response of a filter having a length longer than the designed filter. For example, when designing an 8-tap filter, the reference can be a 12-tap filter. Based on the above rounding method, another set of cosine window sinc filters having lengths of 8 taps, 6 taps, and 4 taps are illustrated in Tables 31 to 36 of FIGS. 31 to 36.

[0098]

[0117] In the application of adaptive resolution change, since the downsampling process of the source picture and the upsampling process of the reconstructed picture may not be standardized in the standard, it means that the user can use any resampling filter. The filter selected by the user may not match the reference resampling filter, but the performance may degrade due to the filter mismatch. In some disclosed embodiments, the coefficients of the reference resampling filter are signaled within the bitstream. The encoder / decoder can use the filter defined by the user for reference resampling. The signaling of the filter coefficients is illustrated in Tables 37 and 38 of FIGS. 37 to 38.

[0099]

[0118] The syntax structures exemplified in Table 37 and Table 38 can be indicated by high-level syntax such as a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header.

[0100]

[0119] The use of the RPR filter determined by the user and the existence of the above-exemplified syntax structure can be controlled by flags in the high-level syntax including SPS, PPS, APS, slice headers, etc.

[0101]

[0120] Consistent with the disclosed embodiments, when one or more of the filtered signals signaled use a fixed filter length and filter norm, a hardware implementation can be more efficient. Thus, in some embodiments, the filter norm and / or filter taps may not be signaled and can instead be implied (e.g., implied to be default values defined within the standard). Filter norm, filter taps, filter accuracy, or maximum bit depth can be part of the coder compliance requirements.

[0102]

[0121] In some embodiments, after the picture is decoded, in order to propose which non-norm upsampling filter should be used at the receiving device, the coder can transmit a supplementary enhancement information (SEI) message. This SEI message is optional. However, with the SEI message, a decoder using the proposed upsampling filter can obtain better video quality. To specify the upsampling filter proposed within such an SEI message, the filter coefficients () syntax structure in Table 31 (Figure 31) can be used.

[0103]

[0122] FIG. 39 is a flowchart of an exemplary method 3900 for processing video content that is consistent with an embodiment of the present disclosure. In some embodiments, method 3900 may be performed by a codec (e.g., an encoder using the encoding process 200A or 200B of FIGS. 2A-2B, or a decoder using the decoding process 300A or 300B of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 3900 at the picture level. For example, the device can process pictures one by one within method 3900. In another example, the device can process a plurality of pictures at once within method 3900. Method 3900 can include the following steps.

[0104]

[0123] In step 3902, in response to the target picture and the reference picture having different resolutions, a band-pass filter can be applied to the reference picture to perform motion-compensated interpolation with reference downsampling to generate a reference block.

[0105]

[0124] The band-pass filter can be a cosine window sinc filter generated based on an ideal low-pass filter and a window function. For example, the kernel function f(n) of the cosine window sinc filter can be the product of an ideal low-pass filter h(n) and a window function w(n) based on Equation (4).

[0106]

[0125] Accordingly,

Equation

[0107]

[0126] The filter coefficients of the cosine window sinc filter can be rounded before application. FIG. 40 is a flowchart of an exemplary method 4000 for rounding the filter coefficients of a cosine window sinc filter that is consistent with an embodiment of the present disclosure. It will be understood that method 4000 can be implemented independently of method 3900 or as part of method 3900. Method 4000 can include the following steps.

[0108]

[0127] In step 4002, the actual filter coefficients of the band-pass filter can be obtained.

[0109]

[0128] In step 4004, a plurality of rounding directions for the actual filter coefficients can be determined respectively.

[0110]

[0129] In step 4006, by rounding the actual filter coefficients according to the plurality of rounding directions respectively, a plurality of combinations of the rounded filter coefficients can be generated.

[0111]

[0130] In step 4008, among multiple combinations, a combination of rounded filter coefficients that minimizes or maximizes the cost function can be selected. In some embodiments, the cost function may be related to the rounded filter coefficients and a reference. The reference can be the actual filter coefficients of the band-pass filter before rounding, the frequency response of an ideal filter, or the frequency response of a filter having a length longer than that of the band-pass filter. For example, if the reference can be the actual filter coefficients of the band-pass filter before rounding, the cost function can be the sum of absolute differences (SAD) or the sum of squared errors (SSE) between the rounded filter coefficients and the actual filter coefficients before rounding. As another example, if the reference is the frequency response of an ideal filter, the cost function can be the SAD or SSE between the frequency response of the rounded filter coefficients and the frequency response of the ideal filter. As discussed above, the reference can also be the frequency response of a filter having a length longer than that of the band-pass filter. For example, if the band-pass filter is an 8-tap filter, the reference can be a 12-tap filter having a length longer than that of the 8-tap band-pass filter.

[0112]

[0131] The selected combination of rounded filter coefficients of the band-pass filter can be signaled in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header, together with the syntax structure related to the filter coefficients.

[0113]

[0132] As shown in Tables 19 to 36 (Figures 19 to 36), it will be understood that the cosine window sinc filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter. In some embodiments, based on the ratio between the resolution of the reference picture and the resolution of the target picture, the band-pass filter can be determined to be one of an 8-tap filter, a 6-tap filter, and a 4-tap filter.

[0114]

[0133] When applying a band-pass filter to a reference picture, a luma sample or a chroma sample can be obtained at a fractional sample position. Using the fractional sample position, the filter coefficients of the obtained luma sample or chroma sample can be determined with reference to a reference table (for example, Tables 19 to 36).

[0115]

[0134] Referring again to FIG. 39, in step 3904, a block of the target picture can be processed using a reference block. For example, a block of the target picture can be encoded or decoded using the reference block.

[0116]

[0135] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by an apparatus (such as an encoder and a decoder disclosed) for performing the above method. General non-transitory media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, and EPROM, flash EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions of those.

[0117]

[0136] Embodiments can be further described using the following clauses: 1. A computer-implemented method for performing motion compensation interpolation, comprising: applying a band-pass filter to a reference picture in order to perform motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions; and processing a block of the target picture using the reference block A computer-implemented method including 2. The method according to clause 1, wherein the bandpass filter is a cosine window sinc filter generated based on an ideal lowpass filter and a window function. 3. The cosine window sinc filter has a kernel function [Number] (where [Number] where fc is the cut-off frequency of the cosine window sinc filter, L is the kernel length, and r is the downsampling ratio of the reference downsampling). The method according to clause 2, based on 4. Obtaining the actual filter coefficients of the bandpass filter, Determining respectively a plurality of rounding directions for the actual filter coefficients, Generating a plurality of combinations of rounded filter coefficients by rounding the actual filter coefficients according to the plurality of rounding directions, and Selecting, from the plurality of combinations, the combination of rounded filter coefficients that minimizes or maximizes the cost function The method according to clause 1, further including 5. The cost function is related to the rounded filter coefficients and the reference, the method according to clause 4. 6. The reference is the actual filter coefficients of the bandpass filter before rounding, the frequency response of the ideal filter, or the frequency response of a filter having a length longer than that of the bandpass filter, the method according to clause 5. 7. The cosine window sinc filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter, the method according to clause 2. 8. Applying the bandpass filter to the reference picture includes obtaining a luma sample or a chroma sample at a fractional sample position, the method according to clause 1. 9. The method according to clause 1, further comprising determining that the band - pass filter is one of an 8 - tap filter, a 6 - tap filter, and a 4 - tap filter based on a ratio between a resolution of a reference picture and a resolution of a target picture. 10. The method according to clause 4, further comprising signaling a selected combination of the rounded filter coefficients of the band - pass filter in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header. 11. A system for performing motion - compensated interpolation, comprising a memory for storing a set of instructions, and at least one processor, the at least one processor being configured to apply a band - pass filter to a reference picture and, in response to the target picture and the reference picture having different resolutions, perform motion - compensated interpolation with reference down - sampling to generate a reference block, and process a block of the target picture using the reference block by executing the set of instructions so as to cause the system to perform. System. 12. The system according to clause 11, wherein the band - pass filter is a cosine - windowed sinc filter generated based on an ideal low - pass filter and a window function. 13. The cosine - windowed sinc filter has a kernel function

Number

Number

[0118]

[0137] It should be noted that relative terms such as "first" and "second" in this specification are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between those entities or operations. Further, terms such as "comprising", "having", "containing", "including" and other similar forms are intended to be equivalent in meaning, and it is not intended that the items following any one of these terms are an exhaustive listing of such items, or are limited to only the items listed, and are intended to be non-limiting in the sense that they do not exclude other items.

[0119]

[0138] As used herein, unless otherwise specified, the term "or" includes all possible combinations except when it is not feasible. For example, when it is stated that a certain database may include A or B, unless otherwise specified or not feasible, the database can include A or B or both A and B. As a second example, when it is stated that a certain database may include A, B or C, unless otherwise specified or not feasible, the database can include A, or B, or C, or both A and B, or both A and C, or both B and C, or all of A, B and C.

[0120]

[0139] It will be understood that the embodiments described above can be implemented by hardware or software (program code) or a combination of hardware and software. When implemented by software, the software can be stored in the above computer-readable medium. The software can execute the disclosed method when executed by a processor. The computing unit and other functional units described in the present disclosure can be implemented by hardware or software or a combination of hardware and software. It will be understood by those skilled in the art that a plurality of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into a plurality of sub-modules / sub-units.

[0121]

[0140] In the above specification, the embodiments have been described with respect to a number of specific details that may vary for each implementation form. Certain adaptations and modifications may be made to the described embodiments. By studying this specification and practicing the invention disclosed herein, other embodiments may become apparent to those skilled in the art. This specification and the examples are to be considered solely as illustrative, and it is intended that the true scope and spirit of the present disclosure be indicated by the appended claims. The order of the steps shown in the figures is for illustrative purposes only and is not intended to be limited to the order of specific steps. Therefore, those skilled in the art can understand that those steps can be executed in different orders while implementing the same method.

[0122]

[0141] The exemplary embodiments have been disclosed in the drawings and in this specification. However, many modifications and variations can be made to those embodiments. Therefore, although specific terms have been used, they have been used for illustrative purposes only and not for purposes of limitation.

Claims

1. A computer-implemented method for performing motion compensation interpolation, comprising: applying a band-pass filter to the reference picture to generate a reference block for performing motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions; and processing a block of the target picture using the reference block A computer-implemented method comprising.

2. The method according to claim 1, wherein the band-pass filter is a cosine window sinc filter generated based on an ideal low-pass filter and a window function.

3. The cosine window sinc filter has a kernel function 【Number 1】 (where 【Number 2】 where fc is the cut-off frequency of the cosine window sinc filter, L is the kernel length, and r is the downsampling ratio of the reference downsampling) The method according to claim 2, based on.

4. obtaining the real filter coefficients of the band-pass filter; determining respectively a plurality of rounding directions for the real filter coefficients; generating a plurality of combinations of rounded filter coefficients by rounding the real filter coefficients according to the plurality of rounding directions; and selecting a combination of rounded filter coefficients that minimizes or maximizes a cost function among the plurality of combinations The method according to claim 1, further comprising.

5. The method according to claim 4, wherein the cost function is related to the rounded filter coefficients and the reference.

6. The method according to claim 5, wherein the reference is the real filter coefficient of the band-pass filter before the rounding, the frequency response of the ideal filter, or the frequency response of a filter having a length longer than the band-pass filter.

7. The method according to claim 2, wherein the cosine window sinc filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter.

8. Applying the band-pass filter to the reference picture includes obtaining luma samples or chroma samples at fractional sample positions. The method according to claim 1.

9. The method according to claim 1, further comprising determining the band-pass filter to be one of an 8-tap filter, a 6-tap filter, and a 4-tap filter based on a ratio between the resolution of the reference picture and the resolution of the target picture.

10. The method according to claim 4, further comprising signaling the selected combination of the rounded filter coefficients of the bandpass filter in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), and a slice header.

11. A system for performing motion compensation interpolation, comprising: a memory storing a set of instructions; and at least one processor, wherein the at least one processor is configured to execute the set of instructions to cause the system to: apply a bandpass filter to the reference picture to generate a reference block for performing motion compensation interpolation with reference downsampling to generate a reference block in response to the target picture and the reference picture having different resolutions; and process a block of the target picture using the reference block. System.

12. The system according to claim 11, wherein the bandpass filter is a cosine window sinc filter generated based on an ideal low-pass filter and a window function.

13. The cosine window sinc filter is a kernel function where 【Number 3】 where fc is the cut-off frequency of the cosine window sinc filter, L is the kernel length, and r is the downsampling ratio of the reference downsampling. 【Number 4】 The system according to claim 12, based on.

14. The at least one processor is further configured to execute the set of instructions to cause the system to: obtain the real filter coefficients of the bandpass filter; determine respectively a plurality of rounding directions for the real filter coefficients; generate a plurality of combinations of rounded filter coefficients by rounding the real filter coefficients according to the plurality of rounding directions; and select a combination of rounded filter coefficients that minimizes or maximizes a cost function from the plurality of combinations. The system according to claim 11, configured as such.

15. The system according to claim 14, wherein the cost function is related to the rounded filter coefficients and the reference.

16. ​ The system of claim 15, wherein the reference is the actual filter coefficient of the bandpass filter before the rounding, the frequency response of an ideal filter, or the frequency response of a filter having a length longer than that of the bandpass filter.

17. The system of claim 12, wherein the cosine window sinc filter is an 8-tap filter, a 6-tap filter, or a 4-tap filter.

18. The system of claim 11, wherein applying the bandpass filter to the reference picture includes obtaining a luma sample or a chroma sample at a fractional sample position.

19. The system of claim 11, wherein at least one processor is configured to execute a set of instructions to further cause the system to determine the bandpass filter to be one of an 8-tap filter, a 6-tap filter, and a 4-tap filter based on a ratio between the resolution of the reference picture and the resolution of the target picture.

20. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer system to cause the computer system to execute a method for processing video content, the method comprising: applying a bandpass filter to a reference picture to generate a reference block by performing motion-compensated interpolation with reference downsampling in response to the target picture and the reference picture having different resolutions; and processing a block of the target picture using the reference block The non-transitory computer-readable medium includes.

Citation Information

Patent Citations

  • Filters for motion-compensated interpolation with reference downsampling.

    JP7653415B2

  • Selective use of coding tools in video processing

    WO2020228660A1

  • Adaptive resolution change in video processing

    WO2020263499A1

  • Implicit signaling of adaptive resolution management based on frame type

    WO2021026363A1