Integerization for Interpolation Filter Design in Video Coding
Patent Information
- Application Number
- JP2024552444
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-13
- Filing Date
- 2023-03-02
- Publication Date
- 2026-02-20
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 268,919, entitled "Integerization Methods for Design of Interpolation Filters," filed on March 4, 2022, U.S. Provisional Application No. 63 / 362,824, entitled "Integerization Methods for Design of Interpolation Filter," filed on April 11, 2022, and U.S. Provisional Application No. 63 / 362,958, entitled "Integerization Methods for Design of Interpolation Filters," filed on April 13, 2022, the entire contents of which are incorporated herein by reference.
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to video processing and, more particularly, to integerization for interpolation filter design in video coding. [Background technology]
[0003] With the widespread use of camera-equipped devices such as smartphones, tablet computers, and computers, taking videos or pictures has become easier than ever. However, even a short video may have a significant amount of data. Video coding techniques (including video encoding and video decoding) can compress video data into smaller sizes, allowing various videos to be stored and transmitted. Video coding has been widely applied to various applications, such as digital television broadcasting, video transmission on the Internet and mobile networks, real-time applications (e.g., video chat, video conferencing, etc.), DVDs, Blu-ray discs, etc. In order to reduce the consumption of storage capacity for storing videos and / or network bandwidth for transmitting videos, it is expected to improve the efficiency of video coding schemes. Summary of the Invention
[0004] Some embodiments relate to integerization for interpolation filter design in video coding. In one example, a method for coding a video includes accessing a plurality of frames of a video, and performing inter prediction on the plurality of frames using an integerization interpolation filter set to generate prediction residuals for the plurality of frames. The integerization interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, and for each interpolation filter of the interpolation filter set, the steps perform: generating two integerization filter coefficient values for each filter coefficient of the interpolation filter; generating a filter candidate set based on the two integerization filter coefficient values of each filter coefficient; calculating an error metric for each filter candidate of the filter candidate set; and selecting an integerization interpolation filter for the interpolation filter from the filter candidate set. The selected integerization interpolation filter has the lowest error metric in the filter candidate set. The method further includes encoding the prediction residuals for the plurality of frames into a bitstream representing the video.
[0005] In another example, a non-transitory computer-readable medium has program code stored thereon. The program code causes one or more processing devices to perform operations including accessing a plurality of frames of a video, and performing inter prediction on the plurality of frames using an integerized interpolation filter set to generate prediction residuals for the plurality of frames. The integerized interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, generating, for each interpolation filter of the interpolation filter set, two integerized filter coefficient values for each filter coefficient of the interpolation filter, generating a filter candidate set based on the two integerized filter coefficient values of each filter coefficient, calculating an error metric for each filter candidate of the filter candidate set, and selecting an integerized interpolation filter for the interpolation filter from the filter candidate set. The selected integerized interpolation filter has the lowest error metric in the filter candidate set. The operations further include encoding the prediction residuals for the plurality of frames into a bitstream representing the video.
[0006] In another example, a system includes a processing device and a non-transitory computer-readable medium communicatively coupled to the processing device. The processing device is configured to perform operations by executing program code stored on the non-transitory computer-readable medium. The operations include accessing a plurality of frames of a video and performing inter prediction on the plurality of frames using an integerized interpolation filter set to generate prediction residuals for the plurality of frames. The integerized interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, generating, for each interpolation filter of the interpolation filter set, two integerized filter coefficient values for each filter coefficient of the interpolation filter, generating a filter candidate set based on the two integerized filter coefficient values of each filter coefficient, calculating an error metric for each filter candidate of the filter candidate set, and selecting an integerized interpolation filter for the interpolation filter from the filter candidate set. The selected integerized interpolation filter has the lowest error metric in the filter candidate set. The operations further include encoding the prediction residuals for the plurality of frames into a bitstream representing the video.
[0007] In another example, a method for decoding a video from a video bitstream includes decoding one or more frames of a video from the video bitstream, and performing inter prediction using an integerized interpolation filter set based on the one or more frames to decode another frame of the video. The integerized interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, generating two integerized filter coefficient values for each filter coefficient of the interpolation filter of the interpolation filter set, generating a filter candidate set based on the two integerized filter coefficient values of each filter coefficient, calculating an error metric for each filter candidate of the filter candidate set, and selecting an integerized interpolation filter for the interpolation filter from the filter candidate set. The selected integerized interpolation filter has the lowest error metric in the filter candidate set. The method further includes displaying the decoded one or more frames and the decoded another frame.
[0008] In another example, a non-transitory computer-readable medium has program code stored thereon, the program code causing one or more processing devices to perform operations including: decoding one or more frames of a video from a video bitstream; and performing inter prediction using an integerized interpolation filter set based on the one or more frames to decode another frame of the video. The integerized interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, generating, for each interpolation filter of the interpolation filter set, two integerized filter coefficient values for each filter coefficient of the interpolation filter; generating a filter candidate set based on the two integerized filter coefficient values of each filter coefficient; calculating an error metric for each filter candidate of the filter candidate set; and selecting an integerized interpolation filter for the interpolation filter from the filter candidate set. The selected integerized interpolation filter has the lowest error metric in the filter candidate set. The operations further include displaying the decoded one or more frames and the decoded another frame.
[0009] In another example, a system includes a processing device and a non-transitory computer-readable medium communicatively coupled to the processing device. The processing device is configured to perform operations by executing program code stored on the non-transitory computer-readable medium. The operations include decoding one or more frames of a video from a video bitstream, and performing inter prediction using an integerized interpolation filter set based on the one or more frames to decode another frame of the video. The integerized interpolation filter set is generated by: accessing an interpolation filter set, each of the interpolation filters of the interpolation filter set having floating-point filter coefficients, generating, for each interpolation filter of the interpolation filter set, two integerized filter coefficient values for each filter coefficient of the interpolation filter, generating a filter candidate set based on the two integerized filter coefficient values of each filter coefficient, calculating an error metric for each filter candidate of the filter candidate set, and selecting an integerized interpolation filter for the interpolation filter from the filter candidate set. The selected integerized interpolation filter has the lowest error metric in the filter candidate set. The operations further include displaying the decoded frame or frames and the decoded another frame.
[0010] These illustrative examples are not intended to limit or define the disclosure, but rather to provide examples to aid in understanding the disclosure. Additional examples are described in specific embodiments and are described in more detail therein. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present application. [Diagram 2] FIG. 2 is a block diagram illustrating an example of a video decoder configured to implement embodiments of the present application. [Diagram 3]1 illustrates an example of coding tree unit partitioning of a picture in a video, according to some embodiments of the present disclosure. [Figure 4] 1 illustrates an example of a coding unit split of a coding tree unit in accordance with some embodiments of the present disclosure. [Diagram 5] 1 illustrates an example process for generating an integerized interpolation filter according to some embodiments of the present disclosure. [Figure 6] 4 illustrates another example of a video encoding process according to some embodiments of the present disclosure. [Figure 7] 1 illustrates another example of a video decoding process according to some embodiments of the present disclosure. [Figure 8] 1 illustrates an example of a computing system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The features, embodiments, and advantages of the present disclosure can be better understood by reading the following specific embodiments with reference to the drawings.
[0013] Various embodiments provide an integerization of an interpolation filter design used in motion compensated inter prediction in video coding. As mentioned above, more and more video data is being generated, stored, and transmitted. It is beneficial to increase the efficiency of video coding techniques. One way to achieve this goal is to use inter prediction, where reconstructed pixels or samples of other frames are used to predict video pixels or samples of a current frame waiting to be decoded. To perform inter prediction, an interpolation filter is typically used, which uses sample values at integer pixel positions to determine predicted samples at sub-pel positions.
[0014] Interpolation filter design methods typically generate filters with floating-point filter coefficients. In practical video encoders, floating-point calculations are undesirable because floating-point calculations may produce different results on different computing architectures. Such instability of floating-point calculations limits the interoperability of video coding standards. Floating-point calculations are more computationally expensive than integer multiplication. To implement a filter set designed with these methods in a practical video coding standard, it is necessary to convert the floating-point filter coefficients into a fixed-point representation with a desired number of bits. However, existing integerization methods only use rounding operations, which may cause the resulting filter to lose the desired characteristics, leading to inaccurate predictions and reduced coding efficiency.
[0015] Various embodiments described in this application solve these problems by proposing an optimized integerization mechanism to minimize the interpolation error. For a given filter design method, the method may generate a set of filters with floating-point filter coefficients and generate an initial filter candidate set for each filter, considering that each filter coefficient can be integerized to two possible values (upper and lower values). The filter candidate set may be evaluated according to an error metric, and the filter candidate with the lowest error may be selected as the integerization result of the filter.
[0016] In one embodiment, the error metric is the squared error between the integerized filter coefficients and the scaled floating-point coefficients. In another embodiment, the error metric is calculated by approximating the integral of the spectral error between the frequency response of the filter candidate and the frequency response of an ideal interpolator. In yet another embodiment, the error metric in any one of the above embodiments is modified by applying an importance weight to the error. In yet another embodiment, rather than selecting the filter candidate with the lowest error for a particular error metric, a reduced set of filter candidates can be generated and tested by the video encoding system to determine the final filter candidate.
[0017] As described in this application, in some embodiments, the video coding efficiency is improved by integerizing the interpolation filter coefficients. By generating a filter candidate set based on multiple possible integerization values and selecting the filter that minimizes the error metric, the error introduced by the integerization can be reduced and the characteristics of the interpolation filter can be maintained as much as possible. Thus, the predictions generated using the interpolation filters can be more accurate and the coding efficiency can be improved. This technique can be an effective coding tool in future video coding standards.
[0018] Referring now to the drawings, Figure 1 is a block diagram illustrating an example of a video encoder 100 configured to implement embodiments of the present application. In the example shown in Figure 1, the video encoder 100 includes a partitioning module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, an in-loop filter module 120, an intra prediction module 126, an inter prediction module 124, a motion estimation module 122, a decoded picture buffer 130, and an entropy encoding module 116.
[0019] The input to the video encoder 100 is an input video 102 that includes a sequence of pictures (also called frames or pictures). In a block-based video encoder, for each picture, the video encoder 100 employs a partition module 112 to partition the picture into blocks 104, each block including a number of pixels. The blocks may be macroblocks, coding tree units, coding units, prediction units, and / or prediction blocks. A picture may include blocks of different sizes, and the block partitioning for different pictures of a video may be different. Each block may be coded using different predictions, such as intra prediction or inter prediction or hybrid intra and inter prediction.
[0020] Typically, the first picture of a video signal is an intra-coded picture, which is coded using only intra-prediction. In the intra-prediction mode, blocks of the picture are predicted using only previously coded data of the same picture. An intra-coded picture can be decoded without information of other pictures. To perform intra-prediction, the video encoder 100 shown in FIG. 1 can use an intra-prediction module 126. The intra-prediction module 126 is configured to generate an intra-predicted block (prediction block 134) using reconstructed samples in a reconstructed block 136 of a neighboring block of the same picture. The intra-prediction is performed according to the intra-prediction mode selected for the block. Then, the video encoder 100 calculates the difference between the block 104 and the intra-predicted block 134. The difference is called a residual block 106.
[0021] To further remove redundancy from the block, the transform module 114 transforms the residual block 106 into a transform domain by performing a transform on the samples in the block. Examples of transforms may include, but are not limited to, a discrete cosine transform (DCT) or a discrete sine transform (DST). The transformed values are also called transform coefficients, which represent the residual block in the transform domain. In some examples, the residual block can be quantized directly without going through the transform of the transform module 114. This mode is called a transform skip mode.
[0022] The video encoder 100 may further use a quantization module 115 to quantize the transform coefficients to obtain quantized coefficients. Quantization involves dividing a sample by a quantization step size followed by rounding, and inverse quantization involves multiplying the quantized value by the quantization step size. Such a quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of a video sample (transformed or untransformed) to represent the video sample with fewer bits.
[0023] Quantization of coefficients / samples within a block can be performed independently, and such quantization methods are used in several existing video compression standards, such as H.264 and high efficiency video coding (HEVC). For an N*M (N-by-M) block, the 2D coefficients of the block can be converted to a 1-D array in some scan order for coefficient quantization and encoding. Quantization of coefficients within a block can utilize scan order information. For example, the quantization of a given coefficient within a block can be determined by the state of one previous quantization value in the scan order. To further improve the coding efficiency, multiple quantizers can be used. Which quantizer is used to quantize the current coefficient depends on the previous information of the current coefficient in the encoding / decoding scan order. Such a quantization method is called dependent quantization.
[0024] The quantization step size can be used to adjust the degree of quantization. For example, in the case of scalar quantization, different quantization step sizes can be applied to achieve finer or coarser quantization. The smaller the quantization step size, the finer the quantization, and the larger the quantization step size, the coarser the quantization. The quantization step size can be indicated by a quantization parameter (QP). By providing the quantization parameter in the encoded bitstream of the video, a video decoder can access and apply the quantization parameter for decoding.
[0025] The entropy encoding module 116 then encodes the quantized samples to further reduce the size of the video signal. The entropy encoding module 116 is configured to apply an entropy encoding algorithm to the quantized samples. In some examples, the quantized samples are binarized into binary bins, and the encoding algorithm further compresses the binary bins into bits. Examples of binarization methods include, but are not limited to, a combination of combined truncated Rice (TR) and k-th order Exp-Golomb (EGk) binarization, and k-th order Exp-Golomb binarization. Examples of entropy coding algorithms include, but are not limited to, variable length coding (VLC) schemes, context adaptive VLC schemes (CAVLC), arithmetic coding schemes, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The entropy coded data is added to a bitstream that outputs the coded video 132.
[0026] As mentioned above, the reconstructed block 136 from the neighboring blocks is used for intra prediction of the block of the picture. The generation of the reconstructed block 136 of a block includes calculating the reconstructed residual of the block. The reconstructed residual can be determined by applying inverse quantization and inverse transform to the quantized residual of the block. The inverse quantization module 118 is configured to apply inverse quantization to the quantized samples to obtain inverse quantized coefficients. The inverse quantization module 118 applies the inverse scheme of the quantization scheme used in the quantization module 115 by using the same quantization step size as the quantization module 115. The inverse transform module 119 is configured to apply the inverse transform (such as inverse DCT or inverse DST) of the transform applied to the transform module 114 to the inverse quantized samples. The output of the inverse transform module 119 is the reconstructed residual of the block in the pixel domain. The reconstructed block 136 in the pixel domain can be obtained by adding the reconstructed residual to the prediction block 134 of the block. For blocks whose transform has been skipped, the inverse transform module 119 is not applied to these blocks. The dequantized samples are the reconstructed residual of the block.
[0027] Blocks of subsequent pictures after the first intra-predicted picture can be coded using inter prediction or intra prediction. In inter prediction, prediction of blocks within a picture is based on one or more previously coded video pictures. To perform inter prediction, video encoder 100 uses inter prediction module 124. Inter prediction module 124 is configured to perform motion compensation on the blocks based on motion estimates provided by motion estimation module 122.
[0028] The motion estimation module 122 performs motion estimation by comparing the current block 104 of the current picture with the decoded reference picture 108. The decoded reference picture 108 is stored in the decoded picture buffer 130. The motion estimation module 122 selects a reference block from the decoded reference picture 108 that best matches the current block. The motion estimation module 122 further identifies an offset between the location (e.g., x, y coordinates) of the reference block and the location of the current block. The offset is called a motion vector (MV) and is provided to the inter prediction module 124 together with the selected reference block. In some cases, multiple reference blocks are identified for the current block in multiple decoded reference pictures 108. Thus, multiple motion vectors are generated and provided to the inter prediction module 124 together with the corresponding reference blocks.
[0029] The inter prediction module 124 performs motion compensation using one or more motion vectors and other inter prediction parameters to generate a prediction of the current block (i.e., inter prediction block 134). For example, the inter prediction module 124 may identify one or more prediction blocks pointed to by the one or more motion vectors in the corresponding one or more reference pictures based on the one or more motion vectors. When there are multiple prediction blocks, the prediction blocks are combined with a certain weight to generate the prediction block 134 of the current block.
[0030] For an inter-predicted block, video encoder 100 may subtract inter-predicted block 134 from block 104 to generate residual block 106. Residual block 106 may be transformed, quantized, and entropy coded in the same manner as the residual for an intra-predicted block described above. Similarly, the residual may be inverse quantized, inverse transformed, and combined with the corresponding prediction block 134 to obtain a reconstructed block 136 for the inter-predicted block.
[0031] To obtain the decoded picture 108 used for motion estimation, the reconstruction block 136 is processed by the in-loop filter module 120. The in-loop filter module 120 is configured to smooth pixel transients, thereby improving video quality. The in-loop filter module 120 can be configured to implement one or more in-loop filters, such as a deblocking filter, or a sample-adaptive offset (SAO) filter, or an adaptive loop filter (ALF).
[0032] 2 illustrates an example of a video decoder 200 configured to implement embodiments of the present application. The video decoder 200 processes encoded video 202 in a bitstream to generate decoded pictures 208. In the example illustrated in FIG. 2, the video decoder 200 includes an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, an in-loop filter module 220, an intra prediction module 226, an inter prediction module 224, and a decoded picture buffer 230.
[0033] The entropy decoding module 216 is configured to perform entropy decoding on the encoded video 202. The entropy decoding module 216 decodes coding parameters, including quantized coefficients, intra-prediction parameters and inter-prediction parameters, and other information. In some examples, the entropy decoding module 216 decodes the bitstream of the encoded video 202 into a binary representation and then converts the binary representation into quantized levels of the coefficients. The entropy decoded coefficient levels are inverse quantized by the inverse quantization module 218 and then inverse transformed to the pixel domain by the inverse transform module 219. The functions of the inverse quantization module 218 and the inverse transform module 219 are similar to the inverse quantization module 118 and the inverse transform module 119 described with respect to FIG. 1 above, respectively. The inverse transformed residual blocks can be added to the corresponding prediction blocks 234 to generate the reconstruction blocks 236. For blocks whose transformations have been skipped, the inverse transform module 219 is not applied to these blocks. The dequantized samples generated by the dequantization module 118 are used to generate a reconstruction block 236 .
[0034] A prediction block 234 for a particular block is generated based on the prediction mode of the block. If the coding parameters of the block indicate that the block is intra predicted, a reconstructed block 236 of a reference block in the same picture can be input to the intra prediction module 226 to generate the prediction block 234 for the block. If the coding parameters of the block indicate that the block is intra predicted, the prediction block 234 is generated by the inter prediction module 224. The functions of the intra prediction module 226 and the inter prediction module 224 are similar to the intra prediction module 126 and the inter prediction module 124 of FIG. 1, respectively.
[0035] 1 above, inter prediction involves one or more reference pictures. The video decoder 200 generates a decoded picture 208 of the reference picture by applying an in-loop filter module 220 to the reconstructed blocks of the reference picture. The decoded picture 208 is stored in a decoded picture buffer 230 and provided to an inter prediction module 224 for output.
[0036] Now, referring to FIG. 3, FIG. 3 illustrates an example of coding tree unit division of a picture in a video according to some embodiments of the present disclosure. As described with reference to FIG. 1 and FIG. 2, to encode a picture of a video, the picture is divided into blocks such as coding tree units (CTUs) 302 in versatile video coding (VVC), as shown in FIG. 3. For example, the CTUs 302 may be blocks of 128×128 pixels. The CTUs are processed according to the order shown in FIG. 3. In some examples, as shown in FIG. 4, each CTU 302 in a picture may be divided into one or more coding units (CUs) 402, and the one or more CUs may be further divided into prediction units or transform units (TUs) for prediction and transformation. Depending on the encoding scheme, the CTUs 302 may be divided into multiple CUs 402 in different manners. For example, in VVC, CUs 402 may be rectangular or square, and may be coded without further division into prediction units or transform units. Each CU 402 may be the same size as its root CTU 302, or may be a subdivision as small as a 4x4 block of the root CTU 302. As shown in Figure 4, methods for dividing a CTU 302 into CUs 402 in VVC include quad-tree division, binary tree division, and ternary tree division. In Figure 4, solid lines indicate quad-tree division, and dashed lines indicate binary or ternary tree division.
[0037] Motion Compensation The tool employed in hybrid video coding systems (such as VVC and HEVC) is to predict video pixels or samples of a current frame waiting to be decoded using pixels or samples of other frames that have already been reconstructed. Coding tools following this architecture are usually called "inter prediction" tools, and the reconstructed frames are sometimes called "reference frames". For stationary video scenes, inter prediction of pixels or samples of a current frame can be achieved by decoding and using coexisting pixels or samples of a reference frame. However, for video scenes with motion, inter prediction tools with motion compensation must be used. For example, a current block of samples of a current frame can be predicted from a "prediction block" of samples of a reference frame, which is determined by first decoding a "motion vector", which signals the location of the prediction block in the reference frame relative to the current block of the current frame. More complex inter prediction tools are used to exploit video scenes with complex motion (such as occlusion and affine motion).
[0038] interpolation When the position of a prediction block relative to a current block is expressed by integer samples, the samples of the prediction block can be obtained directly from the corresponding sample positions in a reference frame. However, in general, the actual motion in a scene may correspond to non-integer samples. In this case, the prediction block is determined using fractional-pel motion compensation. To determine the samples of the prediction block, the sample values at the desired fractional-pel positions are interpolated from the available samples at the integer-pel positions. The interpolation method is selected by balancing design requirements such as complexity, accuracy of the motion vector, interpolation error, and robustness to noise. Despite these trade-offs, prediction of the interpolated prediction block using fractional-pel motion compensation proves to be more advantageous compared to the prediction block using only integer-pel motion compensation.
[0039] To simplify the computation, most interpolation methods are realized by convolution of the available reference frame samples with a linear, shift-invariant set of coefficients. Such an operation is also called filtering. Video coding standards usually realize the interpolation of two-dimensional prediction blocks by applying separate one-dimensional filtering in the vertical and horizontal directions. To be able to signal the motion vector information, the motion vectors are usually restricted to multiples of fractional pixel precision. For example, the motion vectors of luma prediction may be restricted to multiples of 1 / 16 pixel precision.
[0040] JPEG2025508535000002.jpg91170
[0041] JPEG2025508535000003.jpg142170
[0042] The filter design method chosen depends on the tradeoffs considered for a particular video standard. For example, in the research leading to the design of the AVC standard, it was empirically found that typical video content at the time contained noise due to imperfect video capture equipment. Such noise usually has high spatial frequencies and may be due to lack of sensor sensitivity in low lighting conditions or lack of sensor resolution. Video content with large noisy components makes the prediction task difficult, so robustness to noise is a key factor in AVC interpolation filter design. The half-pixel interpolation filters employed in AVC are based on a Wiener filter design, which assumes that the video signal contains noise components due to aliasing during video capture.
[0043] Modern video content is generally of much higher quality, and current filter designs for motion compensation assume that noise from video capture is invisible. Furthermore, modern video codecs employ more complex inter-prediction tools, which significantly reduce the "noise" that results from inaccurate motion modeling. Thus, in current video codecs, the tradeoff in filter design is more focused on minimizing the interpolation error than on robustness to noise. For example, post-VVC exploratory experiments use interpolation filters based on a windowed sinc filter design.
[0044] JPEG2025508535000004.jpg86170
[0045] However, because the sinc function has infinite support, an ideal interpolator cannot be implemented in a practical video codec. To generate a finite support filter, the windowed sinc method multiplies the sinc function by a window function w(s), which is zero outside the desired finite support. For example, the cosine window function JPEG2025508535000005.jpg30170
[0046] JPEG2025508535000006.jpg40170
[0047] Integerization Filter design methods typically produce filters with floating-point filter coefficients. For example, using the windowed sinc design method described above, we can generate filter coefficients as follows: JPEG2025508535000007.jpg64170
[0048] In practical video encoders, floating-point calculations are undesirable because floating-point results may be generated differently on different computing architectures. This instability of floating-point arithmetic limits the interoperability of video coding standards. Floating-point arithmetic is more computationally expensive than integer multiplication. Therefore, to implement a filter set designed in this way in a practical video coding standard, it is necessary to convert the floating-point filter coefficients to a fixed-point representation with the desired number of bits F. Filtering with fixed-point precision filters can be equivalently performed on hardware using integer multiplication and bit-shifting operations, as shown below. JPEG2025508535000008.jpg13170
[0049] JPEG2025508535000009.jpg60170
[0050] However, this approximate integerization leads to degradation of coding performance, since some desired properties of the filter may be lost in the integerization process. For example, one desired property of a filter is that when the interpolation filter acts on a constant signal x[n]=u, the output of the interpolation filter is u. Such a property corresponds to requiring that the frequency response of the interpolation filter have a "DC gain" of 1. The DC gain is equal to the sum of the fixed-point filter coefficients. In the integerized representation, the DC gain constraint can be equivalently satisfied by requiring the following condition: JPEG2025508535000010.jpg11170
[0051] JPEG2025508535000011.jpg69170
[0052] To solve these problems, the present application proposes an integerization method with the smallest interpolation error, which can be applied to any filter design method (e.g., the windowed sinc method described above) that generates a set of P filters, where the P filters have length N and have floating-point filter coefficients. Figure 5 illustrates an example of an integerization interpolation filter generation process 500 according to some embodiments of the present disclosure. One or more computing devices (e.g., a computing device implementing the video encoder 100 or another computing device) implement the operations illustrated in Figure 5 by executing appropriate program code.
[0053] JPEG2025508535000012.jpg73170
[0054] JPEG2025508535000013.jpg89170
[0055] JPEG2025508535000014.jpg173170
[0056] For typical interpolation filter lengths, the number of filter candidates generated is large, but not unmanageable. For example, if N=12, then each floating-point filter generates an initial set of 4096 filter candidates, where N can be any integer between 4 and 16.
[0057] JPEG2025508535000015.jpg29170
[0058] At block 508, the process 500 includes calculating an error metric for each filter candidate in the filter candidate set, specifically as follows: At block 510, the process 500 includes selecting an integerized interpolation filter from the filter candidate set with the lowest error metric as the integerization of filter h_k, each filter repeating the process until all filters of phases 1, 2, ..., P / 2 have selected an integerization. At block 512, the process outputs the integerized interpolation filter set for use in video encoding and decoding.
[0059] JPEG2025508535000016.jpg47170
[0060] In one example of this embodiment, for a windowed sinc filter design using a cosine window, where P=32, N=6, and F=8, the integer filter coefficients are: [Table 1] JPEG2025508535000018.jpg92165
[0061] In another embodiment, the error metric is calculated by approximating the integral of the spectral error between the frequency response of the candidate filter and the frequency response of an ideal interpolator. The frequency response of the candidate filter can be determined by taking a Fourier transform as follows (with some adjustments to the indexing convention used in this disclosure): JPEG2025508535000019.jpg17156
[0062] The error of the ideal interpolator is calculated as follows: JPEG2025508535000020.jpg13156
[0063] The error metric can be estimated as the sum of the spectral power of a discrete sampling of spatial frequency ω. Since we only need the overall size of the error metric between the candidate filters, we can ignore the normalization to the spatial frequency bin width. JPEG2025508535000021.jpg12156
[0064] In one example of this embodiment, for a windowed sinc filter design using a cosine window, where P=32, N=6, and F=8, the integer filter coefficients are: [Table 2] JPEG2025508535000023.jpg91161
[0065] JPEG2025508535000024.jpg40170
[0066] The importance weighting function can model the average power spectral density of the video signal or the average power spectral density of the prediction block selected for motion compensation. One typical feature of natural video signals is the presence of a "deadzone" at high spatial frequencies, i.e., the spectral power at high spatial frequencies is significantly reduced due to the use of anti-aliasing filters in the video acquisition process. One example of a simple weighting function that models this feature is a constant function modified with the "deadzone" at high frequencies. JPEG2025508535000025.jpg14158
[0067] In one example of this embodiment, for a windowed sinc filter design using a cosine window, P=32, N=6, and F=8, and when the weighting function m(ω) in equation (15) is used, the integer filter coefficients are as follows: [Table 3] JPEG2025508535000027.jpg91158
[0068] In yet another embodiment, rather than selecting the filter candidate with the lowest error metric value for a particular error metric, the filter candidate set can be reduced to a reduced filter candidate set with a lower error than the discarded filter candidate. The reduced set can be selected based on any of the error metrics described above. The reduced filter candidate set for each phase is tested with a full hybrid video coding system, and the filter candidate with the best rate-distortion result at each phase is selected. The rate-distortion performance can be measured by a Bjontegaard metric calculated over a set of operating points of the quantization parameters between a reference video coding system with an unchanged filter and a test video coding system with the candidate filters.
[0069] FIG. 6 illustrates an example of a video encoding process 600 using an integerized interpolation filter according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video encoder 100) can implement the operations illustrated in FIG. 6 by executing appropriate program code (including the inter-prediction module 124 and other modules, etc.). For illustrative purposes, the process 600 will be described with reference to the examples illustrated in the figure, although other implementations are possible. However, other implementations are possible.
[0070] At block 602, the process 600 includes accessing a set of frames or pictures of a video signal. As described with respect to FIG. 1, the set of frames of a video can be divided into blocks, such as the encoding unit 402 described with respect to FIG. 4, or any type of block that is processed by a video encoder as a unit when performing inter prediction. At block 604, the process 600 includes performing inter prediction on the set of frames using an integerized interpolation filter set to generate prediction residuals for a plurality of frames. In some examples, the integerized interpolation filter set is generated by the process 500 described with respect to FIG. 5. As described above, the video encoder can use the integerized interpolation filter set to calculate an inter prediction value for a block and calculate a residual by subtracting the inter prediction from the samples of the block. At block 606, the process 600 includes encoding the prediction residuals for the set of frames into a bitstream representing the video. As described in detail above, encoding can include operations such as transforming, quantizing, and entropy encoding the prediction residuals. The coded bits of the prediction residual are included, along with other data, in the bitstream for the video.
[0071] FIG. 7 illustrates an example of a video decoding process 700 according to some embodiments of the present disclosure. One or more computing devices may implement the operations illustrated in FIG. 7 by executing appropriate program code. For example, a computing device implementing the video decoder 200 (e.g., including the inter prediction module 224) may implement the operations illustrated in FIG. 7 by executing the program code. For illustrative purposes, the process 700 will be described with reference to some examples shown in the figure, although other implementations are possible. However, other implementations are possible.
[0072] At block 702, the process 700 includes decoding one or more frames from a video bitstream (e.g., the encoded video 202). As described above, the decoding may include entropy decoding, inverse quantization, inverse transform, and reconstruction of blocks of the frame based on the inter-predicted blocks or intra-predicted blocks. At block 704, the process 700 includes performing inter prediction using an integerized interpolation filter set based on the one or more frames to decode another frame of the video. In some examples, the integerized interpolation filter set is generated according to the process 500 described above with respect to FIG. 5. The inter prediction may be performed using the decoded frame or frames as reference frames and motion vectors decoded from the video bitstream, as described above. At block 706, the process 700 includes decoding the remaining frames in the video into pictures. In some examples, the decoding is performed according to the process described above with respect to FIG. 2. The decoded video may be output for display.
[0073] Examples of Computing Systems Any suitable computing system may be used to perform the operations described herein. For example, FIG. 8 illustrates an example of a computing device 800 that may implement the video encoder 100 of FIG. 1 or the video decoder 200 of FIG. 2. In some embodiments, the computing device 800 may include a processor 812 that is communicatively coupled to a memory 814 and that executes computer-executable program code and / or accesses information stored in the memory 814. The processor 812 may include a microprocessor, an application-specific integrated circuit (ASIC), a state machine, or other processing device. The processor 812 may include any one of a number of processing devices. Such a processor may include or be in communication with a computer-readable medium that stores instructions that, when executed by the processor 812, cause the processor to perform the steps described herein.
[0074] The memory 814 may include any suitable non-transitory computer-readable medium. The computer-readable medium may include any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include magnetic disks, memory chips, ROM, RAM, ASICs, configured processors, optical memory, magnetic tape or other magnetic storage, or any other medium from which a computer processor can read instructions. The instructions may include processor-specific instructions generated by a compiler and / or interpreter based on code written in any suitable computer programming language (including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.).
[0075] The computing device 800 may further include a bus 816. The bus 816 may communicatively couple one or more components of the computing device 800. The computing device 800 may further include a number of external or internal devices, such as input or output devices, e.g., the computing device 800 is shown to include an input / output (I / O) interface 818, which may receive input from one or more input devices 820 or provide output to one or more output devices 822. The one or more input devices 820 and the one or more output devices 822 may be communicatively coupled to the I / O interface 818. The communicative coupling may be realized in any suitable manner (e.g., connection via a printed circuit board, connection via a cable, communication via wireless transmission, etc.). Non-limiting examples of input devices 820 include a touchscreen (e.g., one or more cameras for photographing the touched area, or a pressure sensor for detecting pressure changes caused by a touch), a mouse, a keyboard, or any other device for generating input events in response to physical movements of a user of the computing device. Non-limiting examples of output devices 822 include an LCD screen, an external monitor, speakers, or any other device for displaying or otherwise presenting output generated by the computing device.
[0076] The computing device 800 may execute program code that causes the processor 812 to perform one or more of the operations described above with respect to Figures 1-7. The program code may include the video encoder 100 or the video decoder 200. The program code may reside in the memory 814 or any suitable computer-readable medium and may be executed by the processor 812 or any other suitable processor.
[0077] The computing device 800 may further include at least one network interface device 824. The network interface devices 824 may include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks 828. Non-limiting examples of the network interface devices 824 include Ethernet network adapters, wireless modems, and / or similar devices. The computing device 800 may transmit messages via the network interface devices 824 as electronic or optical signals.
[0078] General Considerations Numerous details are described in this application to provide a thorough understanding of the claimed subject matter. However, as one of ordinary skill in the art will understand, the claimed subject matter may be practiced without these details. In other instances, methods, devices, or systems known to those skilled in the art have not been described in detail so as not to obscure the claimed subject matter.
[0079] Unless otherwise indicated, terms such as "processing," "computing," "calculating," "determining," "identifying," and the like are used herein to describe operations or processes of a computing device (e.g., one or more computers or similar electronic computing devices or apparatus) that manipulate or transform data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage, transmission, or display devices of the computing platform.
[0080] The system or systems described in this application are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include general-purpose computer systems based on microprocessors that access stored software that programs or configures a computing system from a general-purpose computing device to a special-purpose computing device that implements one or more embodiments of the present subject matter. Any suitable programming, scripting, or other type of language or combination of languages can be used to implement the teachings contained in this application in the software used to program or configure a computing device.
[0081] The method embodiments disclosed in this application may be performed by operation of such a computing device. The order of the blocks shown in the above examples may be changed, e.g., the blocks may be reordered, combined, or divided into sub-blocks. Some blocks or processes may be performed in parallel.
[0082] As used in this application, "applied to" or "configured to" is intended to be open and inclusive language and does not exclude equipment adapted or configured to perform additional tasks or steps. Additionally, the use of "based on" is open and inclusive because a process, step, calculation, or other operation "based on" one or more recited conditions or values may in fact be based on additional conditions or values beyond the recited conditions or values. Headings, lists, and numbering contained in this application are for ease of description only and are not intended to be limiting.
[0083] Although the present subject matter has been described in detail with respect to specific embodiments thereof, it should be understood that those skilled in the art, after understanding the foregoing, may readily effect modifications, variations, and equivalents to these embodiments. It should therefore be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter herein that may be readily effected by those skilled in the art.
Claims
1. 1. A method for encoding video, comprising: accessing a plurality of pictures of said video; performing inter prediction on the plurality of pictures using a set of integerized interpolation filters to generate prediction residuals for the plurality of pictures; encoding the prediction residuals of the plurality of pictures into a bitstream representing the video; The integerized interpolation filter set is generated by: accessing a set of interpolation filters, each of the interpolation filters having floating-point filter coefficients; For each interpolation filter in the set of interpolation filters: generating two integer filter coefficient values for each filter coefficient of the interpolation filter; generating a filter candidate set based on the two integerized filter coefficient values for each filter coefficient; calculating an error metric for each filter candidate in the set of filter candidates; selecting an integer interpolation filter for the interpolation filter from the filter candidate set, the selected integer interpolation filter having the lowest error metric in the filter candidate set.
2. the error metric is defined as the squared error between the integerized filter coefficients of the filter candidate and the corresponding floating-point filter coefficients scaled by a particular value; 10. The video encoding method of claim 1.
3. the error metric is calculated by approximating the integral of the spectral error between a first frequency response of the filter candidate and a second frequency response of an ideal interpolator; 10. The video encoding method of claim 1.
4. the error metric is further defined by applying a weight to the spectral error.
4. The video encoding method of claim 3.
5. Selecting the integer interpolation filter from the filter candidate set includes: selecting a reduced filter candidate set from the filter candidate set, wherein each filter candidate in the reduced filter candidate set has an error metric that is lower than the error metrics of the remaining filter candidates in the filter candidate set; applying each filter candidate of the reduced filter candidate set to a video coding system to determine a rate-distortion result for the corresponding filter candidate, wherein the integerized interpolation filter is selected from the reduced filter candidate set based on the rate-distortion result.
10. The video encoding method of claim 1.
6. the two integerized filter coefficient values for each filter coefficient include a maximum integer value and a minimum integer value, the maximum integer value being smaller than the filter coefficient scaled by a particular value, and the minimum integer value being larger than the filter coefficient scaled by the particular value; 10. The video encoding method of claim 1.
7. the filter candidate set includes filter candidates having filter coefficients selected from the two integerized filter coefficient values of corresponding filter coefficients; 7. The video encoding method of claim 6.
8. 1. A computer-readable storage medium having stored thereon program code and a bitstream, the program code, when executed by one or more processing devices, performing the following operations to generate the bitstream, the operations comprising: accessing multiple pictures of a video; performing inter prediction on the plurality of pictures using a set of integerized interpolation filters to generate prediction residuals for the plurality of pictures; encoding the prediction residuals of the plurality of pictures into the bitstream representing the video; The integerized interpolation filter set is generated by: accessing a set of interpolation filters, each of the interpolation filters having floating-point filter coefficients; For each interpolation filter in the set of interpolation filters: generating two integer filter coefficient values for each filter coefficient of the interpolation filter; generating a filter candidate set based on the two integerized filter coefficient values for each filter coefficient; calculating an error metric for each filter candidate in the set of filter candidates; selecting an integer interpolation filter for the interpolation filter from the filter candidate set, the selected integer interpolation filter having the lowest error metric in the filter candidate set.
9. 1. A method for decoding video from a video bitstream, comprising: decoding one or more pictures of the video from the video bitstream; performing inter prediction using an integerized interpolation filter set based on the one or more pictures to decode another picture of the video; and displaying the decoded one or more pictures and another decoded picture; The integerized interpolation filter set is generated by: accessing a set of interpolation filters, each of the interpolation filters having floating-point filter coefficients; For each interpolation filter in the set of interpolation filters: generating two integer filter coefficient values for each filter coefficient of the interpolation filter; generating a filter candidate set based on the two integerized filter coefficient values for each filter coefficient; calculating an error metric for each filter candidate in the set of filter candidates; selecting an integer interpolation filter for the interpolation filter from the filter candidate set, the selected integer interpolation filter having the lowest error metric in the filter candidate set.
10. the error metric is defined as the squared error between the integerized filter coefficients of the filter candidate and the corresponding floating-point filter coefficients scaled by a particular value; 10. The video decoding method of claim 9.
11. the error metric is calculated by approximating the integral of the spectral error between a first frequency response of the filter candidate and a second frequency response of an ideal interpolator; 10. The video decoding method of claim 9.
12. the error metric is further defined by applying a weight to the spectral error. The video decoding method of claim 11.
13. Selecting the integer interpolation filter from the filter candidate set includes: selecting a reduced filter candidate set from the filter candidate set, wherein each filter candidate in the reduced filter candidate set has an error metric that is lower than the error metrics of the remaining filter candidates in the filter candidate set; applying each filter candidate of the reduced filter candidate set to a video coding system to determine a rate-distortion result for the corresponding filter candidate, wherein the integerized interpolation filter is selected from the reduced filter candidate set based on the rate-distortion result.
10. The video decoding method of claim 9.
14. the two integerized filter coefficient values for each filter coefficient include a maximum integer value and a minimum integer value, the maximum integer value being smaller than the filter coefficient scaled by a particular value, and the minimum integer value being larger than the filter coefficient scaled by the particular value; 10. The video decoding method of claim 9.
15. the filter candidate set includes filter candidates having filter coefficients selected from the two integerized filter coefficient values of corresponding filter coefficients; 15. The video decoding method of claim 14.