Adaptive resolution change in video processing
Picture-level resampling and cosine windowed sinc filters in video coding enable efficient adaptive resolution changes, addressing complexity and performance issues in existing methods, ensuring seamless transitions and maintaining coding efficiency.
Patent Information
- Application Number
- JP2025193973
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-13
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-24
AI Technical Summary
Existing video coding standards face challenges in efficiently adapting to changes in video resolution without compromising coding efficiency, particularly in environments with network and device diversity, and existing block-based resampling methods complicate encoder design and degrade performance.
Implementing picture-level resampling methods that adaptively change video resolution by resampling reference pictures to match the current picture's resolution, allowing motion estimation and compensation to be performed independently of resolution changes, and using cosine windowed sinc filters for motion-compensated interpolation.
Enables seamless and efficient resolution changes in video encoding and decoding, maintaining coding efficiency and compatibility with other coding tools, reducing complexity and information loss.
Smart Images

Figure 2026031570000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 865,927, filed June 24, 2019, and U.S. Provisional Patent Application No. 62 / 900,439, filed September 13, 2019, both of which are incorporated herein by reference in their entireties.
[0002] Technical Field FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to video processing, and more particularly to methods and systems for adaptive resolution changes in video coding. [Background technology]
[0003] background
[0003] A video is a series of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video may be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC) / H.265 standard, the Versatile Video Coding (VVC) / H.266 standard, and the AVS standard, that specify specific video coding formats are developed by standardization organizations. As increasingly advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention [Means for solving the problem]
[0004] Disclosure Overview
[0004] Embodiments of the present disclosure provide a method for adaptive resolution change during video encoding or decoding. In one example embodiment, the method includes comparing resolutions of a target picture and a first reference picture, resampling the first reference picture to generate a second reference picture in response to the target picture and the first reference picture having different resolutions, and encoding or decoding the target picture using the second reference picture.
[0005]
[0005] Embodiments of the present disclosure also provide a device for adaptive resolution change during video encoding or decoding. In one example embodiment, the device includes one or more memories that store computer instructions, and one or more processors configured to execute the computer instructions to cause the device to: compare resolutions of a target picture and a first reference picture; in response to the target picture and the first reference picture having different resolutions, resample the first reference picture to generate a second reference picture; and encode or decode the target picture using the second reference picture.
[0006]
[0006] Embodiments of the present disclosure provide a non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer system to cause the computer system to perform a method for adaptive resolution change. In an example embodiment, the method includes comparing resolutions of a target picture and a first reference picture, resampling the first reference picture to generate a second reference picture in response to the target picture and the first reference picture having different resolutions, and encoding or decoding the target picture using the second reference picture.
[0007] BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments and various aspects of the present disclosure are set forth in the following detailed description and accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram illustrating an exemplary video encoder consistent with disclosed embodiments. [Figure 2]
[0009] FIG. 1 is a schematic diagram illustrating an exemplary video decoder consistent with embodiments of the disclosure. [Figure 3]
[0010] 10 illustrates an example where the resolution of a reference picture is different from the current picture, consistent with embodiments of the disclosure. [Figure 4]
[0011] 1 is a table illustrating sub-pel motion compensation interpolation filters used for the luma component in Versatile Video Coding (VVC), consistent with disclosed embodiments. [Figure 5]
[0012] 10 is a table illustrating sub-pel motion compensation interpolation filters used for chroma components in VVC, consistent with disclosed embodiments. [Figure 6]
[0013] 1 illustrates an exemplary reference picture buffer consistent with disclosed embodiments. [Figure 7]
[0014] 1 is a table illustrating an example of a supported resolution set including three different resolutions, consistent with disclosed embodiments. [Figure 8]
[0015] 1 illustrates an exemplary decoded picture buffer (DPB) where both resampled and original reference pictures are stored, consistent with disclosed embodiments. [Figure 9]
[0016] 1 illustrates progressive downsampling consistent with disclosed embodiments. [Figure 10]
[0017] 1 illustrates an exemplary video encoding process with resolution conversion consistent with disclosed embodiments. [Figure 11]
[0018] 1 is a table illustrating exemplary downsampling filters consistent with disclosed embodiments. [Figure 12]
[0019] 1 illustrates an exemplary video encoding process using a peak signal-to-noise ratio (PSNR) calculation, consistent with disclosed embodiments. [Figure 13]
[0020] 1 illustrates the frequency response of an exemplary low pass filter, consistent with disclosed embodiments. [Figure 14]
[0021] 1 is a table illustrating a 6-tap filter, consistent with an embodiment of the disclosure. [Figure 15]
[0022] 1 is a table illustrating an 8-tap filter, consistent with an embodiment of the disclosure. [Figure 16]
[0023] 1 is a table illustrating a 4-tap filter, consistent with an embodiment of the disclosure. [Figure 17]
[0024] 10 is a table illustrating interpolation filter coefficients for 2:1 downsampling, consistent with disclosed embodiments. [Figure 18]
[0025] 10 is a table illustrating interpolation filter coefficients for 1.5:1 downsampling, consistent with disclosed embodiments. [Figure 19]
[0026] 1 illustrates an exemplary luma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 20]
[0027] 1 illustrates an exemplary chroma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 21]
[0028] 1 illustrates an exemplary luma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 22]
[0029] 1 illustrates an exemplary chroma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 23]
[0030] 1 illustrates an exemplary luma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 24]
[0031] 1 illustrates an exemplary chroma sample interpolation filtering process for reference downsampling, consistent with disclosed embodiments. [Figure 25]
[0032] 10 illustrates an exemplary chroma fractional sample position calculation for reference downsampling, consistent with disclosed embodiments. [Figure 26] 1 is a table illustrating an 8-tap filter for MC interpolation using 2:1 ratio reference downsampling, consistent with disclosed embodiments. [Figure 27] 2 is a table illustrating an 8-tap filter for MC interpolation using a 1.5:1 ratio reference downsampling, consistent with disclosed embodiments. [Figure 28] 3 is a table illustrating an 8-tap filter for MC interpolation using a 2:1 ratio reference downsampling, consistent with disclosed embodiments. [Figure 29] 4 is a table illustrating an 8-tap filter for MC interpolation using reference downsampling at a 1.5:1 ratio, consistent with an embodiment of the disclosure. [Figure 30] 5 is a table illustrating 6-tap filter coefficients for luma 4×4 block MC interpolation using 2:1 ratio reference downsampling, consistent with disclosed embodiments. [Figure 31] 6 is a table illustrating 6-tap filter coefficients for luma 4×4 block MC interpolation using 1.5:1 ratio reference downsampling, consistent with disclosed embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] Detailed Description
[0007] Reference will now be made in detail to exemplary embodiments illustrated in the accompanying drawings. The following description will refer to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present invention as set forth in the appended claims. Certain aspects of the present disclosure will now be described in more detail. In the event of a conflict with incorporated terms and / or definitions, the terms and definitions provided herein will control.
[0010]
[0008] A video is a series of still pictures (or "frames") arranged in time sequence to preserve visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display capability) can be used to display such pictures in time sequence. In addition, in some applications, such as for surveillance, conference hosting, or live broadcasting, the video capture device can transmit the captured video to a video playback device (e.g., a computer with a monitor) in real time.
[0011] To reduce the storage space and transmission bandwidth required for such applications, video may be compressed. For example, video may be compressed before storage and transmission and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. A compression module is commonly referred to as an "encoder," and a decompression module is commonly referred to as a "decoder." Encoders and decoders may be collectively referred to as a "codec." Encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression may be performed by a variety of algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x family, etc. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard, in which case the codec is sometimes called a "transcoder."
[0012]
[0010] The video encoding process can identify and retain useful information that can be used for picture reconstruction. If the information ignored in the video encoding process cannot be completely reconstructed, the encoding process may be called "lossy." Otherwise, it may be called "lossless." Most encoding processes are lossy; this is a trade-off to reduce the required storage space and transmission bandwidth.
[0013]
[0011] In many cases, useful information about a picture being encoded (called the "current picture") may include changes relative to a reference picture (e.g., a previously encoded or reconstructed picture). Such changes may include pixel position changes, luminance changes, or color changes, among which position changes are of greatest concern. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.
[0014]
[0012] To achieve the same subjective quality as HEVC / H.265 at half the bandwidth, JVET is developing technology that exceeds HEVC using the joint exploration model ("JEM") reference software. Because the coding technology has been incorporated into JEM, JEM has achieved significantly higher coding performance than HEVC. VCEG and MPEG have also officially begun development of a next-generation video compression standard that exceeds HEVC: the Versatile Video Coding (VVC / H.266) standard.
[0015] The VVC standard continues to add additional encoding techniques that provide better compression performance. VVC can be implemented in the same video encoding systems that have been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263. FIG. 1 is a schematic diagram illustrating an exemplary video encoder 100 consistent with disclosed embodiments. For example, video encoder 100 may perform intra- or inter-coding of blocks within video frames, including video blocks or partitions or subpartitions of video blocks. Intra-coding may rely on spatial prediction to reduce or remove spatial redundancy in video within a given video frame. Inter-coding may rely on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames of a video sequence. Intra-modes may refer to several spatial-based compression modes, and inter-modes (e.g., uni-predictive or bi-predictive) may refer to several temporal-based compression modes.
[0016]
[0014] Referring to FIG. 1, the input video signal 102 may be processed block by block. For example, a video block unit may be a 16x16 pixel block (e.g., a macroblock (MB)). In HEVC, extended block sizes (e.g., coding units (CUs)) can be used to compress video signals of resolutions such as 1080p and higher. In HEVC, a CU may include up to 64x64 luma samples and corresponding chroma samples. In VVC, the size of a CU can be further increased to include 128x128 luma samples and corresponding chroma samples. The CU may be partitioned into prediction units (PUs) to which separate prediction methods can be applied. Each input video block (e.g., MB, CU, PU, etc.) may be processed by using a spatial prediction unit 160 or a temporal prediction unit 162.
[0017] Spatial prediction unit 160 performs spatial prediction (e.g., intra prediction) on the current CU using information about the same picture / slice that contains the current CU. Spatial prediction may use pixels from already-encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction can reduce spatial redundancy inherent in video signals. Temporal prediction (e.g., inter prediction or motion-compensated prediction) may use samples from already-encoded video pictures to predict the current video block. Temporal prediction can reduce temporal redundancy inherent in video signals.
[0018] Temporal prediction unit 162 performs temporal prediction (e.g., inter-prediction) on the current CU using information from one or more pictures / slices different from the picture / slice containing the current CU. Temporal prediction for a video block may be signaled by one or more motion vectors. A motion vector may indicate the amount and direction of motion between the current block and one or more of its predictive blocks in a reference frame. If multiple reference pictures are supported, one or more reference picture indexes may be sent for the video block. The one or more reference indexes may be used to identify which reference picture or pictures in decoded picture buffer (DPB) 164 (also referred to as reference picture store 164) the temporal prediction signal may come from. After spatial prediction or temporal prediction, mode decision and encoder control unit 180 in the encoder may choose a prediction mode based on, for example, a rate-distortion optimization method. The predictive block may be subtracted from the current video block at adder 116. The prediction residual may be transformed by transform unit 104 and quantized by quantization unit 106. The quantized residual coefficients may be inverse quantized in inverse quantization unit 110 and inverse transformed in inverse transform unit 112 to form a reconstructed residual. The reconstructed block may be added to the prediction block in summer 126 to form a reconstructed video block. In-loop filtering, such as a deblocking filter and adaptive loop filter 166, may be applied to the reconstructed video block before it is added to reference picture store 164 and used to encode future video blocks. To form output video bitstream 120, the coding mode (e.g., inter or intra), prediction mode information, motion information, and the quantized residual coefficients may be sent to entropy coding unit 108 to be compressed and packed to form bitstream 120.
[0019]
[0017] Consistent with the disclosed embodiments, the above-mentioned units of the video encoder 100 may be implemented as software modules (e.g., computer programs that implement different functions), hardware components (e.g., different circuit blocks for performing respective functions), or a hybrid of software and hardware.
[0020]
[0018] Figure 2 is a schematic diagram illustrating an exemplary video decoder 200 consistent with embodiments of the disclosure. Referring to Figure 2, a video bitstream 202 may be unpacked or entropy decoded in an entropy decoding unit 208. Coding mode or prediction information may be sent to a spatial prediction unit 260 (e.g., if intra-coded) or a temporal prediction unit 262 (e.g., if inter-coded) to form a prediction block. If inter-coded, the prediction information may include a prediction block size, one or more motion vectors (e.g., which may indicate the direction and amount of motion), or one or more reference indices (e.g., which may indicate which reference picture the prediction signal is obtained from).
[0021]
[0019] Motion compensated prediction may be applied by the temporal prediction unit 262 to form a temporal prediction block. The residual transform coefficients may be sent to the inverse quantization unit 210 and the inverse transform unit 212 to reconstruct a residual block. The prediction block and the residual block may be summed at 226. The reconstructed block may undergo in-loop filtering (by a loop filter 266) before being stored in a decoded picture buffer (DPB) 264 (also called reference picture store 264). The reconstructed image in the DPB 264 can be used to drive a display device or to predict future video blocks. The decoded image 220 may be displayed on a display.
[0022]
[0020] Consistent with the disclosed embodiments, the above units of the video decoder 200 may be implemented as software modules (e.g., computer programs that implement different functions), hardware components (e.g., different circuit blocks for performing respective functions), or a hybrid of software and hardware.
[0023]
[0021] One of the goals of the VVC standard is to provide videoconferencing applications with the ability to tolerate network and device diversity. In particular, the VVC standard must provide the ability to rapidly adapt to changing network environments, including rapidly reducing the encoding bit rate when network conditions deteriorate and rapidly increasing video quality when network conditions improve. In addition, for adaptive streaming services that provide multiple representations of the same content, each of the multiple representations may have different characteristics (e.g., spatial resolution or sample bit depth), and the video quality may vary from low to high. Therefore, the VVC standard must support fast representation switching for adaptive streaming services. During switching from one representation to another (e.g., switching from one resolution to another), the VVC standard must enable the use of an efficient prediction structure without compromising fast and seamless switching capabilities.
[0024] In some embodiments consistent with this disclosure, to change resolution, an encoder (e.g., encoder 100 of FIG. 1) sends an instantaneous-decoder-refresh (IDR) coded picture to clear the contents of a reference picture buffer (e.g., DPB 164 of FIG. 1 and DPB 264 of FIG. 2). Upon receiving the IDR coded picture, a decoder (e.g., decoder 200 of FIG. 2) marks all pictures in the reference buffer as "not used for reference." All subsequent transmitted pictures can be decoded without referencing any frames decoded before the IDR picture. The first picture in a coded video sequence is always an IDR picture.
[0025]
[0023] In some embodiments consistent with this disclosure, adaptive resolution change (ARC) techniques may be used to enable a video stream to change spatial resolution between coded pictures within the same video sequence without requiring new IDR pictures and without requiring multiple layers as in the case of scalable video codecs. According to the ARC technique, at a resolution switch point, the currently coded picture is predicted from a reference picture of the same resolution (if available) or from a reference picture of a different resolution by resampling the reference picture. A diagram of adaptive resolution change is shown in Figure 3, where the resolution of reference picture "Ref0" is the same as the resolution of the currently coded picture. However, the resolutions of reference pictures "Ref1" and "Ref2" are different from the resolution of the current picture. To generate a motion-compensated prediction signal for the current picture, both "Ref1" and "Ref2" are resampled to the resolution of the current picture.
[0026]
[0024] Consistent with the present disclosure, when the resolution of a reference frame is different from that of a current frame, one method for generating a motion-compensated prediction signal is picture-based resampling, in which the reference picture is first resampled to the same resolution as the current picture, and an existing motion compensation process using motion vectors can be applied. The motion vectors may be scaled (if the motion vectors are sent in units before resampling is applied) or not (if the motion vectors are sent in units after resampling is applied). When using picture-based resampling, particularly with regard to reference picture downsampling (i.e., when the resolution of the reference picture is greater than that of the current picture), information may be lost in the reference resampling step before motion-compensated interpolation, because downsampling is usually achieved by low-pass filtering followed by decimation.
[0027] Another method is block-based resampling, where resampling is performed at the block level by examining one or more reference pictures used by the current block, and if one or both of them have a different resolution than the current picture, resampling is performed in combination with a sub-pel motion compensation interpolation process.
[0028]
[0026] This disclosure provides both picture-based and block-based resampling methods for use with ARC. The following description first refers to the disclosed picture-based resampling method, and then refers to the disclosed block-based resampling method.
[0029]
[0027] The disclosed picture-based resampling method can solve several problems caused by conventional block-based resampling methods. First, in conventional block-level resampling methods, on-the-fly block-level resampling may be performed when it is determined that the resolution of a reference picture is different from that of a current picture. However, the block-level resampling process may complicate encoder design because the encoder may need to resample blocks at each search point during motion search. Motion search is generally a time-consuming process for an encoder, and therefore, the on-the-fly requirement during motion search may complicate the motion search process. As an alternative to on-the-fly block-level resampling, the encoder may resample the reference picture in advance so that the resampled reference picture has the same resolution as the current picture. However, this may cause prediction signals during motion estimation and motion compensation to differ, thereby reducing coding efficiency.
[0030]
[0028] Second, in conventional block-based resampling methods, the block-level resampling design is incompatible with some other useful coding tools, such as subblock-based temporal motion vector prediction (SbTMVP), affine motion compensation prediction, and decoder side motion vector refinement (DMVR). When ARC is enabled, these coding tools can be disabled. However, such disabling significantly degrades coding performance.
[0031]
[0029] Third, any sub-pel motion compensation interpolation filter known in the art can be used for block resampling, but it may not be equally applicable to block upsampling and block downsampling. For example, if the reference picture has a lower resolution than the current picture, the interpolation filter used for motion compensation can be used for upsampling. However, if the reference picture has a higher resolution than the current picture, the interpolation filter is not suitable for downsampling because it cannot filter integer positions and therefore may cause aliasing. For example, Table 1 (FIG. 4) shows exemplary interpolation filter coefficients used for luma component integer positions, along with various values of fractional sample positions. Table 2 (FIG. 5) shows exemplary interpolation filter coefficients used for chroma components, along with various values of fractional sample positions. As shown in Table 1 (FIG. 4) and Table 2 (FIG. 5), no interpolation filter is applied at fractional sample position 0 (i.e., integer position).
[0032] To avoid the above problems associated with conventional block-based resampling methods, the present disclosure provides a picture-level resampling method that can be used for ARC. The picture resampling process may involve upsampling or downsampling. Upsampling is increasing the spatial resolution while maintaining a two-dimensional (2D) representation of an image. In the upsampling process, the resolution of a reference picture is increased by interpolating unavailable samples from neighboring available samples. In the downsampling process, the resolution of a reference picture is decreased.
[0033]
[0031] According to the disclosed embodiments, resampling can be performed at the picture level. In picture-level resampling, if the resolution of a reference picture is different from that of the current picture, the reference picture is resampled to the resolution of the current picture. Motion estimation and / or compensation for the current picture can be performed based on the resampled reference picture. In this way, ARC can be implemented in a "transparent" manner in the encoder and decoder, since the block-level operations are independent of the resolution change.
[0034]
[0032] Picture-level resampling can be performed on the fly, i.e., while the current picture is being predicted. In some exemplary embodiments, only original (non-resampled) reference pictures are stored in a decoded picture buffer (DPB), e.g., DPB 164 in encoder 100 (FIG. 1) and DBP 264 in decoder 200 (FIG. 2). The DPB can be managed in the same manner as in current versions of VVC. In the disclosed picture-level resampling method, before encoding or decoding a picture, the encoder or decoder resamples the reference picture if the resolution of the reference picture in the DPB is different from the resolution of the current picture. In some embodiments, a resampling picture buffer may be used to store all resampled reference pictures for the current picture. The resampled pictures of the reference pictures are stored in the resampling picture buffer, and motion estimation / compensation for the current picture is performed using pictures from the resampling picture buffer. After encoding or decoding is completed, the resampling picture buffer is removed. 6 shows an example of on-the-fly picture-level resampling, in which low-resolution reference pictures are resampled and stored in a resampled picture buffer. As shown in FIG. 6, the DPB of an encoder or decoder contains three reference pictures. The resolutions of reference pictures "Ref0" and "Ref2" are the same as the resolution of the current picture. Therefore, "Ref0" and "Ref2" do not need to be resampled. However, the resolution of reference picture "Ref1" is different from the resolution of the current picture and therefore needs to be resampled. Therefore, only the resampled "Ref1" is stored in the reference picture buffer, and "Ref0" and "Ref2" are not stored in the reference picture buffer.
[0035]
[0033] Consistent with some exemplary embodiments, in on-the-fly picture-level resampling, the number of resampled reference pictures a picture can use is constrained. For example, the maximum number of resampled reference pictures for a given current picture may be preset. For example, the maximum number may be set to be 1. In this case, a bitstream constraint may be imposed such that an encoder or decoder can tolerate at most one of the reference pictures having a different resolution than the resolution of the current picture, and all other reference pictures must have the same resolution. This maximum number is directly related to the size of the resampled picture buffer and the worst-case decoder complexity, as it indicates the maximum number of resamplings that can be performed on any picture in the current video sequence. Therefore, this maximum number may be signaled as part of a sequence parameter set (SPS) or a picture parameter set (PPS), or may be specified as part of a profile / level definition.
[0036]
[0034] In some exemplary embodiments, two versions of a reference picture are stored in the DPB. One version has the original resolution, and the other version has the full resolution. If the resolution of the current picture is different from the original resolution or the full resolution, the encoder or decoder may perform on-the-fly downsampling from the stored full resolution picture. For picture output, the DPB always outputs the original (non-resampled) reference picture.
[0037] In some exemplary embodiments, the resampling ratio can be selected arbitrarily, and the vertical and horizontal scaling ratios may be different. Because picture-level resampling is applied, the block-level operations of the encoding / decoder are made independent of the resolution of the reference picture, thereby allowing for arbitrary resampling ratios without further complicating the block-level design logic of the encoder / decoder.
[0038]
[0036] In some exemplary embodiments, the signaling of the coding resolution and the maximum resolution can be performed as follows: The maximum resolution of any picture of the video sequence is signaled in the SPS. The coding resolution of a picture can be signaled in the PPS or in the slice header. In both cases, one flag is signaled indicating whether the coding resolution is the same as the maximum resolution. If the coding resolution is not the same as the maximum resolution, the coding width and height of the current picture are additionally signaled. If PPS signaling is used, the signaled coding resolution applies to all pictures that reference this PPS. If slice header signaling is used, the signaled coding resolution applies only to the current picture itself. In some embodiments, the difference between the current coding resolution and the signaled maximum resolution can be signaled in the SPS.
[0039]
[0037] In some exemplary embodiments, instead of using any resampling ratio, the resolution is limited to a predefined set of N supported resolutions (N is the number of supported resolutions in the sequence). The values of N and the supported resolutions can be signaled in the SPS. The coding resolution of a picture is signaled using the PPS or slice header. In both cases, instead of signaling the actual coding resolution, a corresponding resolution index is signaled. Table 3 (Fig. 7) shows an example of a supported resolution set where a coding sequence allows three different resolutions. If the coding resolution of the current picture is 1440x816, the corresponding index value (=1) is signaled using the PPS or slice header.
[0040]
[0038] In some exemplary embodiments, both resampled and non-resampling reference pictures are stored in the DPB. For each reference picture, one original (i.e., non-resampling) picture and N-1 resampled copies are stored. If the resolution of the reference picture is different from that of the current picture, the resampled picture in the DPB is used to encode the current picture. For picture output, the DPB outputs the original (non-resampling) reference picture. As an example, Figure 8 shows the occupancy rate of the DPB for the supported resolution set shown in Table 3 (Figure 7). As shown in Figure 8, the DPB contains N (e.g., N=3) copies of each picture used as a reference picture.
[0041]
[0039] The choice of the value of N depends on the application. Larger values of N provide more flexibility in resolution selection, but increase complexity and memory requirements. Smaller values of N are suitable for less complex devices, but limit the resolutions that can be selected. Thus, in some embodiments, flexibility may be given to the encoder to determine the value of N signaled by the SPS based on the application and device capabilities.
[0042]
[0040] Consistent with this disclosure, downsampling of reference pictures may be performed progressively (i.e., gradually). In conventional (i.e., direct) downsampling, downsampling of an input image is performed in a single step. If the downsampling ratio is high, single-step downsampling requires a longer tap downsampling filter to avoid severe aliasing problems. However, longer tap filters are computationally expensive. In some exemplary embodiments, to maintain the quality of the downsampled image, progressive downsampling is used when the downsampling ratio is greater than a threshold (e.g., greater than 2:1 downsampling), in which downsampling is performed gradually. For example, to achieve a downsampling ratio greater than 2:1, a downsampling filter sufficient for 2:1 downsampling may be repeatedly applied. Figure 9 shows an example of progressive downsampling in which a factor of four downsampling is performed across two pictures in both the horizontal and vertical dimensions. The first picture is downsampled by a factor of two (in both directions), and the second picture is downsampled by a factor of two (in both directions).
[0043]
[0041] Consistent with this disclosure, the disclosed picture-level resampling method can be used with other coding tools. For example, in some exemplary embodiments, scaled motion vectors based on resampled reference pictures can be used in a temporal motion vector predictor (TMVP) or an advanced temporal motion vector predictor (ATMVP).
[0044]
[0042] Next, a test methodology for evaluating encoding tools for ARC will be described. Since adaptive resolution change is mainly used to adapt to network bandwidth, the following test conditions can be considered to compare the encoding efficiency of different ARC schemes during network bandwidth change:
[0045] In some embodiments, the resolution is changed to half (both vertically and horizontally) at a particular time instance, and then restored to the original resolution after a period of time. Figure 10 shows an example of resolution change. At time t1, the resolution is changed to half, and then restored to full resolution at time t2.
[0046] In some embodiments, when the resolution is reduced, if the downsampling ratio is too large (e.g., the downsampling ratio for a given dimension is greater than 2:1), then progressive downsampling is used over a period of time. However, when upsampling is used, progressive upsampling is not used.
[0047] In some embodiments, a Scalable HEVC Test Model (SMH) downsampling filter may be used for source resampling. Table 4 (FIG. 11) shows the detailed filter coefficients along with different fractional sample positions.
[0048] In some exemplary embodiments, two peak signal-to-noise ratios (PSNRs) are calculated to measure video quality. A diagram of the PSNR calculation is shown in Figure 12. The first PSNR is calculated between the resampled source and the decoded picture. The second PSNR is calculated between the original source and the upsampled decoded source.
[0049] In some embodiments, the number of pictures encoded is the same as the current test condition for VVC.
[0050]
[0048] Next, the disclosed block-based resampling method will be described. Block-based resampling can reduce the above-mentioned information loss by integrating resampling and motion-compensated interpolation into one filtering operation. Take the following case as an example: the motion vector of the current block has half-pel precision in one dimension (e.g., in the horizontal dimension), and the width of the reference picture is twice that of the current picture. In this case, compared to picture-level resampling, which reduces the width of the reference picture by half to match the width of the current picture and then performs half-pel motion interpolation, the block-based resampling method directly fetches odd positions of the reference picture as reference blocks with half-pel precision. At the 15th JVET meeting, a block-based resampling method for ARC, in which motion-compensated (MC) interpolation and reference resampling are integrated and performed with a single-step filter, was adopted for VVC. In VVC Draft 6, the existing filter for MC interpolation without reference resampling is reused for MC interpolation with reference resampling. The same filter is used for both reference upsampling and downsampling. The details of filter selection are described below.
[0051] For the luma component, if half-pel AMVR mode is selected and the interpolation position is half-pel, a 6-tap filter [3,9,20,20,9,3] is used, and if the motion compensation block size is 4x4, the following 6-tap filter as shown in Table 5 (Figure 13) is used; otherwise, an 8-tap filter as shown in Table 6 (Figure 14) is used. For the chroma component, a 4-tap filter as shown in Table 7 (Figure 15) is used.
[0052]
[0050] In VVC, the same filter is used for MC interpolation without reference resampling and MC interpolation with reference resampling. Although the MCIF of VVC is designed based on DCT upsampling, it may not be appropriate to use it as a single-step filter that integrates reference downsampling and MC interpolation. For example, in the case of phase 0 filtering (i.e., scaled mv is an integer), the VVC 8-tap MCIF coefficient is [0,0,0,64,0,0,0,0], which means that the predicted sample is directly copied from the reference sample. This may not be a problem for MC interpolation without reference downsampling or with reference upsampling, but in the case of reference downsampling, it may cause aliasing artifacts due to the lack of a low-pass filter before decimation.
[0053]
[0051] This disclosure provides a method using a cosine windowed sinc filter for MC interpolation using reference downsampling.
[0054]
[0052] A windowed sinc filter is a bandpass filter that separates one band of frequencies from another. A windowed sinc filter is a lowpass filter with a frequency response that passes all frequencies below the cutoff frequency with an amplitude of 1 and cuts all frequencies above the cutoff frequency with an amplitude of zero, as shown in Figure 16.
[0055]
[0053] The filter kernel, also known as the filter's impulse response, is obtained by taking the inverse Fourier transform of the frequency response of an ideal low-pass filter: The impulse response of a low-pass filter is the general form of a sinc function:
number
number
[0056]
[0054] The sinc function is infinite. To make the filter kernel finite length, a window function is applied to truncate the filter kernel to L points. To obtain a smooth tapered curve, a cosine windowing function is used, which is given by:
number
[0057]
[0055] The kernel of a cosine windowed sinc filter is the product of an ideal response function h(n) and a cosine window function w(n).
number
[0058]
[0056] Two parameters chosen for the windowed sink kernel are the cutoff frequency fc and the kernel length L. For the downsampling filter used in the Scalable HEVC Test Model (SHM), fc=0.9 and L=13.
[0059]
[0057] The filter coefficients obtained in (4) are real numbers. Applying the filter is equivalent to calculating a weighted average of reference samples, with the weights being the filter coefficients. For efficient calculation in a digital computer or hardware, the coefficients are normalized, multiplied by a scaler, and rounded to integers so that the sum of the coefficients is equal to 2 to the Nth power (N is an integer). The filtered sample is divided by 2 to the Nth power (equivalent to a right shift of N bits). For example, in VVC Draft 6, the sum of the interpolation filter coefficients is 64.
[0060] In some disclosed embodiments, a downsampling filter is used in SHM for VVC motion compensated interpolation using reference downsampling for both luma and chroma components, and an existing MCIF is used for motion compensated interpolation using reference upsampling. If the kernel length L=13 and the first coefficient is small and rounded to zero, the filter length can be reduced to 12 without affecting the filter performance.
[0061] As an example, the filter coefficients for 2:1 downsampling and 1.5:1 downsampling are shown in Table 8 (FIG. 17) and Table 9 (FIG. 18), respectively.
[0062] Besides the coefficient values, there are several other differences between the SHM filter design and existing MCIF designs.
[0063]
[0061] The first difference is that the SHM filter requires filtering at integer and fractional sample positions, while the MCIF only requires filtering at fractional sample positions. An example of modifications to the luma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is listed in Table 10 (Figure 19).
[0064]
[0062] An example of modifications to the chroma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is listed in Table 11 (Figure 20).
[0065]
[0063] A second difference is that for the SHM filter, the total number of filter coefficients is 128, while for the existing MCIF, the total number of filter coefficients is 64. In VVC Draft 6, to reduce losses caused by rounding errors, the intermediate prediction signal is maintained at a higher precision (represented by a larger bit depth) than the output signal. The precision of the intermediate signal is called the internal precision. In one embodiment, to maintain the same internal precision as VVC Draft 6, the output of the SHM filter needs to be right-shifted by one more bit compared to when the existing MCIF is used. An example of changes to the luma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is shown in Table 12 (Figure 21).
[0066] An example of modifications to the chroma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is shown in Table 13 (FIG. 22).
[0067] According to some embodiments, the internal precision can be increased by 1 bit and an additional 1-bit right shift can be used to convert the internal precision to the output precision. An example of modifications to the luma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is shown in Table 14 (Figure 23).
[0068]
[0066] An example of modifications to the chroma sample interpolation filtering process for the reference downsampling case of VVC Draft 6 is shown in Table 15 (Figure 24).
[0069]
[0067] As a third difference, the SHM filter has 12 taps. Therefore, 11 adjacent samples (5 to the left, 6 to the right, or 5 above, 6 below) are required to generate an interpolated sample. Compared to MCIF, additional adjacent samples are fetched. In VVC Draft 6, the chroma mv precision is 1 / 32. However, the SHM filter only has 16 phases. Therefore, the chroma mv can be rounded to 1 / 16 for reference downsampling. This can be done by right-shifting the last 5 bits of chroma mv by 1 bit. An example of changes to the chroma fractional sample position calculation for the reference downsampling case in VVC Draft 6 is shown in Table 16 (Figure 25).
[0070] In some embodiments, to match the existing MCIF design of the VVC draft, we propose using an 8-tap cosine windowed sinc filter. The filter coefficients can be derived by setting L=9 in the cosine windowed sinc function of equation (4). To further match the existing MCIF filter, the total filter coefficients can be set to 64. Example filter coefficients for ratios of 2:1 and 1.5:1 are shown in Table 17 (FIG. 26) and Table 18 (FIG. 27), respectively.
[0071] According to some embodiments, a 32-phase cosine windowed sinc filter set may be used in chroma motion compensated interpolation using reference downsampling to accommodate 1 / 32 sample precision for the chroma components. Example filter coefficients for ratios of 2:1 and 1.5:1 are shown in Table 19 (Figure 28) and Table 20 (Figure 29), respectively.
[0072] According to some embodiments, for a 4x4 luma block, a 6-tap cosine windowed sinc filter may be used for MC interpolation with reference downsampling. Example filter coefficients for ratios of 2:1 and 1.5:1 are shown below in Table 21 (Figure 30) and Table 22 (Figure 31), respectively.
[0073] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which can be executed by a device (such as the disclosed encoder and decoder) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or other magnetic data storage media, CD-ROMs, other optical data storage media, any physical media with a pattern of holes, RAM, PROMs, and EPROMs, FLASH-EPROMs or other flash memories, NVRAMs, caches, registers, other memory chips or cartridges, and networked versions of the above. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0074]
[0072] It should be noted that relational terms herein, such as "first" and "second," are used only to distinguish one entity or operation from another, and do not require or imply an actual relationship or order between those entities or operations. Also, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent in meaning and to be open-ended in that the term or terms following any one of these words is not an exhaustive enumeration of such term or terms, or limited to only the enumerated term or terms.
[0075]
[0073] As used herein, unless specifically stated otherwise, the term "or" encompasses all possible combinations unless it is impracticable. For example, if it is stated that a database may include A or B, then the database may include A, or B, or A and B, unless specifically stated otherwise or it is impracticable. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless specifically stated otherwise or it is impracticable.
[0076]
[0074] It is understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that more than one of the above modules / units can be integrated into one module / unit, and that each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0077]
[0075] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the above specification and examples be considered exemplary only, with the true scope and spirit of the invention being indicated by the following claims. Additionally, the order of steps depicted in the figures is intended for illustrative purposes only and is not intended to be limited to any particular order of steps. Thus, one skilled in the art will recognize that these steps may be performed in different orders while performing the same method.
[0078]
[0076] The embodiments can be further described using the following clauses. 1. A computer-implemented image processing method, comprising: comparing the resolutions of the target picture and the first reference picture; resampling the first reference picture to generate a second reference picture in response to the target picture and the first reference picture having different resolutions; encoding or decoding the target picture using the second reference picture; A method comprising: 2. storing a second reference picture in a first buffer, the first buffer being different from a second buffer storing a decoded picture used to predict a future picture; removing the second reference picture from the first buffer after completing encoding or decoding of the target picture; 2. The method of clause 1, further comprising: 3. Encoding or decoding the target picture is 3. The method of any one of clauses 1 and 2, comprising encoding or decoding the target picture by using no more than a predetermined number of resampled reference pictures. 4. Encoding or decoding the target picture is 4. The method of clause 3, comprising signaling the predetermined number as part of a sequence parameter set or a picture parameter set. 5. Resampling the first reference picture to generate the second reference picture; storing a first version and a second version of the first reference picture, the first version having an original resolution and the second version having a maximum resolution to which the reference picture can be resampled; resampling the maximum version to generate a second reference picture; 2. The method according to clause 1, comprising: 6. Storing the first version and the second version of the first reference picture in a decoded picture buffer; outputting the first version for encoding future pictures; 6. The method of clause 5, further comprising: 7. Resampling the first reference picture to generate the second reference picture, 10. The method of claim 1, comprising generating a second reference picture at a support resolution. 8. Number of supported resolutions and Pixel dimensions corresponding to the supported resolutions, 8. The method of claim 7, further comprising signaling information indicative of: 9. Information 9. The method of claim 8, including an index corresponding to at least one of the supported resolutions. 10. The method of any one of clauses 8 and 9, further comprising setting the number of supported resolutions based on the configuration of the video application or video device. 11. Resampling the first reference picture to generate the second reference picture; 11. The method of any one of clauses 1 to 10, comprising progressive downsampling of the first reference picture to generate the second reference picture. 12. A device comprising: one or more memories for storing computer instructions; Execute computer instructions to cause the device to: comparing a resolution of the target picture with a resolution of the first reference picture; resampling the first reference picture to generate a second reference picture in response to the target picture and the first reference picture having different resolutions; encoding or decoding the target picture using the second reference picture; one or more processors configured to cause Including, the device. 13. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: comparing the resolutions of the target picture and the first reference picture; resampling the first reference picture to generate a second reference picture in response to the target picture and the first reference picture having different resolutions; encoding or decoding the target picture using the second reference picture; 1. A non-transitory computer-readable medium comprising: 14. Responsive to the target picture and the reference picture having different resolutions, applying a bandpass filter to the reference picture to perform motion compensated interpolation and to generate the reference block; encoding or decoding a block of the target picture using the reference block; A computer-implemented image processing method, comprising: 15. The method of clause 14, wherein the bandpass filter is a cosine windowed sinc filter. 16. Cosine windowed sinc filter with kernel function
number
number
[0079]
[0077] In the drawings and specification, illustrative embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A computer-implemented image processing method, comprising: comparing the resolutions of the target picture and the first reference picture; in response to the target picture and the first reference picture having different resolutions, resampling the first reference picture to generate a second reference picture; encoding or decoding the target picture using the second reference picture; A method comprising:
2. storing the second reference picture in a first buffer, the first buffer being different from a second buffer that stores decoded pictures used to predict future pictures; removing the second reference picture from the first buffer after the encoding or decoding of the target picture is completed; The method of claim 1 further comprising:
3. encoding or decoding the target picture, 10. The method of claim 1, comprising encoding or decoding the target picture by using no more than a predetermined number of resampled reference pictures.
4. encoding or decoding the target picture, 4. The method of claim 3, comprising signaling the predetermined number as part of a sequence parameter set or a picture parameter set.
5. resampling the first reference picture to generate the second reference picture; storing a first version and a second version of the first reference picture, the first version having an original resolution and the second version having a maximum resolution to which the reference picture can be resampled; resampling the maximum version to generate the second reference picture; The method of claim 1 , comprising:
6. storing the first version and the second version of the first reference picture in a decoded picture buffer; outputting the first version for encoding future pictures; The method of claim 5 further comprising:
7. resampling the first reference picture to generate the second reference picture; The method of claim 1 , comprising generating the second reference picture at a support resolution.
8. The number of supported resolutions and pixel dimensions corresponding to the supported resolution; The method of claim 7 , further comprising signaling information indicative of:
9. The information is The method of claim 8 , including an index corresponding to at least one of the supported resolutions.
10. The method of claim 8 , further comprising setting the number of supported resolutions based on a configuration of a video application or a video device.
11. resampling the first reference picture to generate the second reference picture; The method of claim 1 , comprising progressively downsampling the first reference picture to generate the second reference picture.
12. A device, one or more memories for storing computer instructions; Executing the computer instructions causes the device to: comparing a resolution of the target picture with a resolution of the first reference picture; in response to the target picture and the first reference picture having different resolutions, resampling the first reference picture to generate a second reference picture; encoding or decoding the target picture using the second reference picture; one or more processors configured to cause Including, the device.
13. The one or more processors execute the computer instructions to cause the device to: storing the second reference picture in a first buffer, the first buffer being different from a second buffer that stores decoded pictures used to predict future pictures; removing the second reference picture from the first buffer after the encoding or decoding of the target picture is completed; The device of claim 12 , further configured to:
14. The one or more processors execute the computer instructions to cause the device to: The device of claim 12 , further configured to cause encoding or decoding of the target picture by using no more than a predetermined number of resampled reference pictures.
15. The one or more processors execute the computer instructions to cause the device to: The device of claim 14 , further configured to cause signaling of the predetermined number as part of a sequence parameter set or a picture parameter set.
16. The one or more processors execute the computer instructions to cause the device to: storing a first version and a second version of the first reference picture in the one or more memories, the first version having an original resolution and the second version having a maximum resolution to which the reference picture can be resampled; resampling the maximum version to generate the second reference picture; The device of claim 12 , further configured to:
17. The one or more processors execute the computer instructions to cause the device to: storing the first version and the second version of the first reference picture in a decoded picture buffer; outputting the first version for encoding future pictures; 17. The device of claim 16, further configured to:
18. The one or more processors execute the computer instructions to cause the device to: The device of claim 12 , further configured to generate the second reference picture at a supported resolution.
19. The one or more processors execute the computer instructions to cause the device to: The number of supported resolutions and pixel dimensions corresponding to the supported resolution; 20. The device of claim 18, further configured to signal information indicative of: as part of a sequence parameter set or a picture parameter set.
20. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: comparing the resolutions of the target picture and the first reference picture; in response to the target picture and the first reference picture having different resolutions, resampling the first reference picture to generate a second reference picture; encoding or decoding the target picture using the second reference picture; 1. A non-transitory computer-readable medium comprising:
Citation Information
Patent Citations
Method and apparatus for video coding and decoding
US20140219346A1
Support of adaptive resolution change in video coding
WO2020142382A1
Signaling of adaptive picture size in video bitstream
WO2020185814A1
Signaling for reference picture resampling
WO2020228691A1
Signaling for reference picture resampling
WO2020263665A1