Temporal filtering strengh for reference picture resampling
By fine-tuning temporal filtering strength and refining motion estimation based on the RPR mode's resolution ratio, the method addresses the inefficiencies in video encoding when RPR is applied, resulting in improved coding efficiency and compression performance.
Patent Information
- Application Number
- PCT/EP2024/086447
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-26
AI Technical Summary
Existing video encoding technologies face challenges in achieving optimal coding efficiency when Reference Picture Resampling (RPR) is applied, as the temporal filtering process is not adequately adapted to the changing resolution ratios.
The proposed solution involves fine-tuning the temporal filtering strength based on the resolution ratio of the RPR mode, either by applying adaptive filtering before encoding or by postponing filtering until after rescaling, and refining motion estimation by adjusting the block size according to the resolution ratio.
This approach enhances coding efficiency by reducing noise and improving motion estimation accuracy, leading to better compression results when RPR is employed.
Smart Images

Figure EP2024086447_26062025_PF_FP_ABST
Abstract
Description
TEMPORAL FILTERING STRENGH FOR REFERENCE PICTURE RESAMPLINGCROSS REFERENCE TO RELATED APPLICATION[1] This application claims the benefit of European Patent Application No. 23307319.6, filed on December 21 , 2023, which is incorporated herein by reference in their entirety.TECHNICAL FIELD[2] The present embodiments generally relate to a method and an apparatus for adapting the temporal filtering process to reference picture resampling in video encoding.BACKGROUND[3] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction.SUMMARY[4] In various implementations, methods and devices are disclosed that adapt a temporal filtering of input video that reduces noise before encoding according to a resolution ratio when a Reference Picture Resampling is applied to achieve better coding efficiency. According to one embodiment, the temporal filtering strength value is fine-tuned according to the resolution ratio of the RPR mode at preprocessing module. According to another embodiment, the temporal filtering is postponed after rescaling when the RPR mode is enabled, and the filtering strength value are fine-tuned according to the resolution ratio of the RPR mode. A third embodiment proposes to refine motion estimation in an encoder by setting the block size of motion estimation in the temporal filtering according to the resolution ratio of the RPR mode.[5] According to a first aspect, a method of video encoding is disclosed that comprises obtaining a plurality of resampling ratio corresponding to a down-sampling of a picture from a first resolution to a second resolution of a picture, wherein a first resolution is greater than a second resolution of the picture; applying to a picture to encode, an adaptive temporal filtering to form a plurality of filtered pictures, each filtered picture being associated to a resampling ratio, wherein a strength of the adaptive temporal filtering is based on the resampling ratio, and wherein the plurality of filtered images have a first resolution; determining an indication whether Reference Picture Resampling is to be applied to encode the picture; responsive to the indication that Reference Picture Resampling is to be applied to encode the picture, determining a resampling ratio to be applied to encode the picture and down-sampling thefiltered picture associated to the determined resampling ratio from the first resolution to the second resolution of the picture to form a downsampled picture; and encoding the downsampled picture.[6] According to a second aspect, a method of video encoding is disclosed that comprises determining an indication whether Reference Picture Resampling is to be applied to encode a picture; responsive to the indication that Reference Picture Resampling is to be applied to encode a picture, determining a resampling ratio to be applied to encode the picture and downsampling the picture from the first resolution to a second resolution of the picture to form a downsampled picture, wherein a first resolution is greater than a second resolution of the picture; applying to the downsampled picture an adaptive temporal filtering to form a filtered picture associated to a resampling ratio, wherein a strength of the adaptive temporal filtering is based on the resampling ratio; and encoding the downsampled picture.[7] According to another embodiment, an apparatus for encoding video is provided, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform the encoding method according to any of the embodiments described herein.[8] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the encoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding a video according to the methods described herein.[9] One or more embodiments also provide a computer readable storage medium having stored thereon video data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the video data generated according to the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented.
[0011] FIG. 2 illustrates a block diagram of an embodiment of a video encoder.
[0012] FIG. 3 illustrates a block diagram of an embodiment of a video decoder.
[0013] FIG. 4 illustrates a block diagram of an embodiment of a video encoder using RPR.
[0014] FIG. 5 illustrates a block diagram of an embodiment of a video decoder using RPR.
[0015] FIG. 6 illustrates picture resolution changes with usage of RPR for inter-prediction.
[0016] FIG. 7a and 7b illustrate picture ordering within a GOP example.
[0017] FIG. 8 illustrate RD-curves (Anchor vs RPR).
[0018] FIG. 9 illustrates a method with pre-processing temporal filtering to which at least one embodiment may apply.
[0019] FIG. 10 illustrates a method adapting temporal filtering to RPR mode according to an embodiment.
[0020] FIG. 1 1 illustrates a method adapting temporal filtering to RPR mode according to an embodiment.
[0021] FIG. 12 illustrates a method adapting temporal filtering to RPR mode according to an embodiment.
[0022] FIG. 13 illustrates a method adapting temporal filtering to RPR mode according to an embodiment.
[0023] FIG. 14 illustrates resolution layer for motion estimation according to an embodiment. DETAILED DESCRIPTION
[0024] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.
[0025] The system 100 includes at least one processor 1 10 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device, and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.
[0026] System 100 includes an encoder / decoder module 130 configured, for example, toprocess data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
[0027] Program code to be loaded onto processor 110 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 1 10, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0028] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 1 10 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, HEVC, or VVC.
[0029] The input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0030] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (forexample) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0031] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 1 10 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0032] Various elements of system 100 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 1 15, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0033] The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0034] Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802. 11. The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
[0035] The system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device- to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0036] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0037] FIG. 2 illustrates an example video encoder 200, such as a VVC (Versatile Video Coding) encoder. FIG. 2 may also illustrate an encoder in which improvements are made to the VVC standard or an encoder employing technologies similar to VVC.
[0038] In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, and the terms “image,” “picture” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side.
[0039] Before being encoded, the video sequence may go through pre-encoding processing(201 ), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression. Pre-encoding processing (201 ) may also comprise temporal filtering to reduce noise present in the video sequences and by the way improves the encoding efficiency. Metadata can be associated with the preprocessing, and attached to the bitstream.
[0040] In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, CUs. Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction (260). In an inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which one of the intra mode or inter mode to use for encoding the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. After prediction, prediction enhancement (285) is applied to the prediction block. Prediction residuals are calculated, for example, by subtracting (210) the predicted block from the original image block.
[0041] The prediction residuals are then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes.
[0042] The encoder decodes an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (240) and inverse transformed (250) to decode prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (265) are applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (280).
[0043] FIG. 3 illustrates a block diagram of an example video decoder 300. In the decoder 300, a bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG. 2. The encoder 200 also generally performs video decoding as part of encoding video data.
[0044] In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (335)the picture according to the decoded picture partitioning information. The transform coefficients are de-quantized (340) and inverse transformed (350) to decode the prediction residuals. Combining (355) the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). After prediction, prediction enhancement (390) is applied to the prediction block. In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380).
[0045] The decoded picture can further go through post-decoding processing (385), for example, an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the preencoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0046] Reference Picture Resampling (RPR)
[0047] In the Versatile Video Coding (VVC) / H.266 standard, the picture-based re-scaling feature for video coding is named Reference Picture Resampling (RPR). Given an original video sequence composed of pictures of size (width x height), the encoder may choose for each frame which resolution (picture size) to use for coding the frame. Different picture parameter sets (PPS) are coded in the bit-stream with the possible sizes of the pictures and the slice / picture header indicates which PPS to use to decode the current video coding layer (VCL) network abstraction layer (NAL) unit.
[0048] FIG. 4 is a functional block diagram of an example encoder (400) using RPR. The encoder (400) includes a down-sampler (440), a core encoder (410), a re-sampler (430, a component of motion compensator 450), and a decoded picture buffer (DPB, 420). The encoder (400) is employed to encode an original video into a coded video bitstream. In an aspect, the core encoder (410) represents the encoder (200) of FIG. 2, except for the reference picture buffer (280) and the motion compensator (270) that are shown in FIG. 4 as decoded picture buffer (420) and motion compensator (450), respectively.
[0049] In the example of FIG. 4, for each input frame of the original video, the encoder (400) can select whether to encode the frame at the original frame size or at a reduced frame size. If the encoder selects to encode at a reduced frame size, the input frame is down-sampled, by the down-sampler (440), before the frame is encoded by the core encoder (410). For example, the frame may be sub-sampled (440) by a resolution rate of 1 / 2, resulting in frame size of width W / 2 and of height H / 2. The decision whether to encode a frame at the original frame size or at a reduced frame size can be done by the encoder (400), for example, by comparing coding results of the input frame at various resolutions or based on spatial andtemporal activity related to the coding of input frames of the original video. Consequently, when image content from a currently encoded frame and a corresponding content from a reference frame (needed for inter-prediction of the content from the currently encoded frame) are not at the same resolution, a resampling operation is required. In VVC, the resampling operation and the motion compensation are combined into one single step. This can be done by the re-sampler (430) that can be configured to change the resolution of image content from reference frames to a resolution that matches the resolution of content from the current frame.
[0050] FIG. 5 is a functional block diagram of an example decoder (500) using RPR. The decoder (500) includes a core decoder (510), a re-sampler (530, a component of motion compensator 550), a decoded picture buffer (520), and an up-sampler (540). The decoder (500) is employed to decode the original video from a given coded video bitstream, resulting in the output video. In an aspect, the core decoder (510) represents the decoder (300) of FIG. 3, except for the reference picture buffer (380) and the motion compensator (375) that are shown in FIG. 5 as decoded picture buffer (520) and motion compensator (550), respectively.
[0051] In the example of FIG. 5, image content from one or more frames (stored in the coded picture buffer 520) that is needed as reference to inter-predict image content from a currently decoded frame is re-sampled, by the re-sampler (530), into a resolution that matches the resolution of the currently decoded image content. To output the video frames that are decoded (500) in a reduced resolution are up-sampled to the resolution of the original video, by the up-sampler (540), resulting in the output video. Note that the down-sampler (440) of FIG. 4 (e.g., employed in a pre-processing stage) and the up-sampler (540) of FIG. 5 (e.g., employed in a post-processing stage) are not specified by the VCC standard.
[0052] For each frame, the encoder chooses whether to encode at the original or down-sized resolution (e.g., picture width / height divided by 2). The choice can be made with two-pass encoding or considering spatial and temporal activity in the original pictures for example. Consequently, the decoded picture buffer (DPB, 420, 520) can contain pictures with different sizes than the current picture size.
[0053] In case one reference picture in the DPB has a size different from the current picture, the resampling (430, 530) (up-scaling or down-scaling) of the reference block to build the prediction block is made implicitly during the motion compensation process (270, 375).
[0054] FIG. 6 illustrates the application of RPR in a coded stream with successive pictures having different picture resolutions. This feature is for instance useful for adaptive streaming, to adapt the bitrate to the network constraints. Even if it is not normatively specified, the down- sampled pictures, that are not at the resolution indicated in the HLS, need to be resampled to show all decoded pictures on the end-display at a same (display or target) resolution, as illustrated in FIG. 5 (540).
[0055] Temporal Filtering
[0056] In the state of the art, video encoders such as VVC Test model encoder, Temporal Filtering is applied as pre-processing step at the encoder side. It brings benefits in terms of encoding efficiency by reducing the video noise. Some prior art approaches also introduce a Motion Compensated Temporal Filtering (MCTF) using a bilateral filter performed within a group of pictures (GOP) structure of hierarchical video frames encoding.
[0057] FIG. 7a and 7b illustrates picture ordering within an GOP example. In the Random- Access configuration mode (RA), video sequence is encoded using a hierarchical group of picture structure (GOP). An example of 16 frames per GOP is depicted in FIG. 7a and 7b. Pictures in the bitstream are coded according to the GOP structure as shown on FIG. 7a rather than Picture Order Counting (POC) or display order as shown on FIG. 7b. For instance, pictures with 0 and 16 have a temporal ID (TID) of 0 as shown in FIG. 7b and coded respectively at 0 (first) and 1 (second) in the coding order as shown in FIG. 7a. Picture of POC 8 has TID of 1 and coded at 2 (third) in the coding order.
[0058] A motion estimation (ex: block-based motion compensation) may be used to calculate the motion compensation vectors. On that purpose, the luma component of the two neighboring frames before and after the current frame to encode are down sampled two times to get three resolution layers. The motion estimation may be performed using a full-pel search of NxN pixel block. A sum of squared difference of each block at different resolution layer and the displaced block in the neighboring frames is achieved to find the best motion estimation vectors. Motion compensation is then applied to create motion compensated neighboring pictures.
[0059] Temporal filtering is processed only on pictures located at low coding hierarchy (low TID value). These pictures are referenced the most and encoded with higher fidelity. As depicted in FIG. 7a and 7b, it typically corresponds to temporal ID (TID) 0 and 1 of GOP hierarchy. A filter strength value may be set as follows: 1,5 n = 0,16,32 etc s0(n)=0,95 n = 8,24,40 etc (.0 otherwise
[0060] Bilateral filtering may be applied on each pixel of source image Io as follows:
[0061] In is the filtered sample value, Io is the original sample value, lr(i) is the co-located sample value in the neighboring picture I after motion compensation and wr(i,a) is the weight of neighboring picture i when the number of available neighboring pictures is equal to a.
[0062] For the luma samples, the weight, wr(i,a), may be calculated as follows:wr(i, a) = 0.4Where: i(Q ) = 3 * (QP - 10) 0 1 23For all other values of I and a: sr(i,1 ) = 0.3.
[0063] For chroma samples, the weight, wr(i,a) is calculated as follows: wr(i, d) = 0.55<jc= 30
[0064] Evaluation of video coding performance
[0065] Rate-distortion and complexity are two criteria that are usually used to compare video codec performance.
[0066] Rate-distortion measures the compression efficiency and gives the relationship between the bitrate (Eq. 1 ) and the quality of the reconstructed video. Peak Signal to Noise Ratio (PSNR) (Eq. 2) is often used to evaluate the quality, where / (j,j) andrespectively represent pixel at (i, j) position in the original image I and pixel at (i, j) position in the reconstructed image K.
[0067] Generally, a tradeoff between the bitrate and the distortion is controlled by a quantization parameter (QP) input. Video codec performance is evaluated by performing the bitrate and PSNR measurement at several QP values, constituting the rate distortion curve (RD-curve). Traditionally, if one wants to compare the performance of two versions of the same video encoder for a range of QP, Bjontegaard delta rate (BD-rate) measurement is used to measure the average bitrate and the quality difference between RD-curve of each version of the encoder.
[0068] FIG. 8 illustrates such RD-curves. In the following we denote by “RPR mode” the resizing (down-sampling or rescaling) of the frame. In case of RPR vs. non-RPR encoderversion (ie Anchor), the use of RPR usually improves the BD-rate for high QP values. It may be determined at which location the RD curves cross as shown in FIG. 8. The crossing point, also known as a QP switch, is video content dependent. It does not have the same value and may be different from frame to frame or from group of frames to another. The QP switch may be used to decide whether to apply or not the RPR mode, in this example based on a PSNR performance metric. It has been demonstrated that most of the time RPR brings gain for low bitrate encoding.
[0069] Thus, when RPR mode is activated in a video encoder, video frames are rescaled, before encoding, with respect to a resolution rate (or downscaling ratio) decided by subsampling decision process at the encoder side. Downscaling applies a low pass filtering additionally to the temporal filtering already achieved during the pre-processing module. These 2 processes are applied independently resulting in a strong effect of filtering before the encoding process. For instance, too much high frequency may be removed by the filtering which may have an impact on the encoding / decoding process in terms of efficiency.
[0070] The present document introduces at least one method adapting the temporal filtering process according to the resolution ratio when RPR mode is activated in order to achieve better coding efficiency. According to one embodiment, a method is disclosed that fine-tune the temporal filtering strength value according to the resolution ratio of the RPR mode at preprocessing module. According to another embodiment, a method is disclosed that postpone the temporal filtering after rescaling when the RPR mode is activated and fine-tune the filtering strength value according to the resolution ratio of the RPR mode. A third embodiment proposes to refine motion estimation in an encoder by setting the block size of motion estimation in the temporal filtering according to the resolution ratio of the RPR mode.
[0071] Generic encoding method with a pre-processing temporal filtering adapted to RPR
[0072] In a first embodiment, the temporal filtering strength at preprocessing module is adapted to the resolution ratio of the RPR mode.
[0073] FIG. 9 illustrates a method (900) with pre-processing temporal filtering to which at least one embodiment may apply. The frames or pictures of the original input video (910) are first pre-processed by a temporal filtering (920) before being sent to the encoder as explained above. The pre-processing temporal filtering has the advantage to lower the effect of the noise present in the video frames and consequently improves the video encoding efficiency. Besides, a decision is made (930) inside the encoder whether to apply the image subsampling or not, the rescaler (940) downscales the filtered frames with a downscaling ratio (also referred as the subsampling ratio, down-sampling ratio, resampling ratio, downscaling ratio or resolutionratio) set by the subsampling decision module (930).
[0074] FIG. 10 illustrates a method (1000) adapting temporal filtering to RPR mode according to an embodiment. The method proposes to optimize the temporal filtering process by sharing the predefined downscaling ratios values between the subsampling decision and the temporal filtering modules. In a first step, a plurality of subsampling ratio r0to rNcorresponding to a subsampling of a picture from a first (original) resolution to a second (reduced) resolution of a picture are obtained. For instance, for a frame having a size of width W and of height H corresponding to the first resolution, a subsampling ratio r0of 1 / 2 results in frame of size of width W / 2 and of height H / 2 corresponding to the second resolution. In a variant, one may consider a subsampling ratio of 1 meaning that the resolution of the frame does not change (ie RPR mode is not activated). Thus, the first (original) resolution is greater than the second (reduced) resolution of the picture but a variant may consider that the first resolution of the picture is greater than or equal to the second resolution of the picture. The overall strength value So(n) (where n denotes the frame number in the display order) is weighted by each subsampling ratio value resulting in 1 to N temporally filtered image where N is the number of predefined subsampling ratios.
[0075] In a step 1020, an adaptive temporal filtering is applied to the input picture to form a plurality of filtered pictures (1030), each of the filtered pictures being associated to a given subsampling ratio r. A strength So, s’oof the adaptive temporal filtering is based on a subsampling ratio r0, rNand provided to the temporal filtering to adapt the filtering to the subsampling ratio. The plurality of filtered frames (1030) is input to the encoder. The encoder may determine (1040) an indication whether Reference Picture Resampling is to be applied to encode the picture along with a subsampling ratio to be applied to encode the picture. Responsive to the indication that Reference Picture Resampling is to be applied to encode the picture subsampling ratio to be applied to encode the picture, the filtered picture associated to the determined subsampling ratio is selected (switch 2) to be subsampled by the rescaler module (1050). Responsive to the indication that Reference Picture Resampling is to be applied to encode the picture subsampling ratio to be applied to encode the picture, the rescaled filtered picture at the second resolution or the filtered picture at the first resolution is selected (1060) to be further processed by the encoding loop.
[0076] As detailed above, the switch 2 module selects the filtered images corresponding to the downscaling ratio decided by the subsampling decision module. When the subsampling decision is activated, the selected filtered frames are rescaled and outputted by switch 1 to the encoder before encoding gets started. Alternatively, when subsampling decision is disabled filtered frames with the initial strength so are selected by switch 1 and presented to the encoder input. The new strength parameter may be a function depending on So anddownscaling ratio n (equation 4). The determining the strength parameter from so and n may be implemented by a look up table as shown on FIG. 10 where So and n are the inputs parameters of the LUT and Si‘ is the output of the LUT corresponding to the new strength weighted by the predefined downscaling ratios. Equations (5) and (6) are examples of mapping functions that may be implemented in the look up table.6 [r0, ri, . Tv] where N is the number of predefined downscaling ratios, n is the downscaling ratio of index i and So is the initial strength as defined in section 1.1. The goal of such solution is to have a fine tune control on the filtering process performed by the temporal filtering in accordance with the image downscaling achieved by the rescaler driven by the subsampling decision.
[0077] FIG. 11 illustrates a method adapting temporal filtering to RPR mode according to a variant embodiment. In this variant where, the temporal filtering is turned off when RPR is enabled, the strength value may be set to 0 for any n, thus the original video frames 1 110 are sent to the Rescaler instead of the temporally filtered frames when RPR mode is activated. When RPR mode decision is deactivated, original images are filtered using strength sO as shown in FIG. 1 1 , and the filtered frames are sent to the encoding through the selection via the switch 1 .
[0078] FIG. 12 illustrates a method adapting temporal filtering to RPR mode according to another variant embodiment. According to this variant, when subsampling mode is activated, whether the weighted strength So‘ or So is applied in the temporal filtering module according to the frames TID. For instance, frames in TID 0 and / or 1 may be temporally filtered with So‘ and while the other frames with So. Indeed, the frames in the lowest sublayers TID are referenced the most, and this variant allows having a fine tuning of the temporal strength on such frames which is impactful. As shown on FIG.12, POC counter values are shared between temporal filtering and Subsampling decision module as well.
[0079] As an example, for a GOP of 16 frames and TID value equal to 0 and 1 , frames of POC value 0, 8 and 16 are temporally filtered with So’ and frames within TID greater than 1 are temporally filtered with so. It is also possible to disable the temporal filtering for images of specific TID by setting So’ to 0 accordingly. Then, original images of that TID are redirected to the rescaler instead of temporally filtered frames when subsampling decision is activated.
[0080] In variants, the concept of adapting the filter strength depending on the ratio, as described above, may be easily and directly generalized to other parameters used to controlthe action of the filter. For instance, the filter length (e.g. number of successive pictures used to perform the temporal filtering), or the temporal distance between the current pictures and the pictures used to filter it, may be depending on the downscaling ratio. Typically, the filter length or temporal distance is decreased while the downscaling ratio increased (ie from 1 / 2 to 1 / 4). According to yet another variant, an adaptation of the motion interpolation or an adaptation of the motion vector accuracy may depend on the resolution ratio. Typically, motion vector accuracy is increased while the downscaling ratio increases. Besides, when the motion vector accuracy is increased, the set of motion interpolation filters may support more interpolation phases. Another example is the adaption of the parameters used for deriving the weight wr(i, a) mentioned above. For instance, in the formula below, Sbase may depend on the downscaling ratio (decreased when downscaling ratio increases, and vice-versa). Si can also depend on the downscaling ratio (increased when downscaling ratio increases).
[0081] Generic encoding method with temporal filtering applied after rescaling and adapted to RPR
[0082] In a second embodiment, the temporal filtering strength is adapted to the resolution ratio of the RPR mode and applied after the rescaler module.
[0083] FIG. 13 illustrates a method adapting temporal filtering to RPR mode according to a second embodiment. This embodiment proposes to implement the temporal filtering after the rescaler when subsampling is activated. This has the advantage to adapt the temporal filtering strength on the fly as the downscaling ratio value is known and the subsampling decision already performed at this stage of the processing.
[0084] According to this embodiment, in a pre-processing step, a temporal filtering is applied to an input video picture to form a filtered picture. Both the filtered picture and the original picture are input to the encoding. The encoder may determine an indication whether Reference Picture Resampling is to be applied to encode the picture. The encoder may further determine a resampling ratio to be applied to encode the picture when RPR mode is enabled. Then, responsive to the indication that Reference Picture Resampling is to be applied to encode the picture, the original picture is downsampled by the rescaler from the first resolution to a second (reduced) resolution of the picture to form a downsampled picture. In a following step, the downsampled picture is filtered using an adaptive temporal filter to form a filtered picture. Any of the variants of the previous embodiment relative to adapting the temporal filter are compatible with the current design. For instance, a strength of the adaptive temporal filtering may based on the resampling ratio. Thus, temporal filtering strength can be fine-tuned according to the downscaling ratio selected by the subsampling decision. For instance, theadaptive temporal filtering strength may be obtained from a LUT taking as input the resampling ratio and original filter strength So. The fine-tuned filtering strength So’ is outputted by the Look Up Table that is driven by the downscaling ratio and the initial strength So. ln yet another variant, the temporal filtering strength may be adapted to the POC of the frame or to TID of the frame. When RPR mode is activated, for the variant corresponding to turning off the temporal filtering (setting strength to 0), the downscaled original frames are redirected to the encoder instead of the temporally filtered frames.
[0085] Besides, responsive to the indication that Reference Picture Resampling is not to be applied to encode the picture, the filtered picture at the first resolution is selected to be further processed by the encoding loop. Thus, in the other way around, when subsampling decision is disabled, original images are filtered using strength so and feed the encoder input as shown in FIG. 13.
[0086] According to yet another example, for the variant that use different temporal filtering strengths according to the image TID, rescaled images within TID 0 and / or 1 are temporally filtered with So' and the other frames with So. A fine tuning of the temporal strength on such frames is impactful because the frames in the lowest sublayers TID are referenced the most. As depicted in FIG. 13, POC counter values may be used by the temporal filtering on that purpose.
[0087] Generic encoding method with motion estimation adapted to RPR
[0088] In another embodiment, a motion prediction of a block of a picture to encode is determined by applying motion estimation to the block. Responsive to the indication that Reference Picture Resampling is to be applied to encode a picture, a size of the block used in the block matching search for motion estimation is weighted by the subsampling ratio; and the block is encoded based on the motion prediction. Advantageously, this embodiment may be combined with any of the variant of the first and second implementation of an adaptive temporal filtering described previously.
[0089] A hierarchical estimation of motion vector field is calculated for motion compensation used by the temporal filtering. It has the advantage to reduce the noise present in the video sequences and by the way improves the encoding efficiency. Besides, the use of hierarchical motion estimation is more robust and allows to find large motion easier.
[0090] FIG. 14 illustrates resolution layer for motion estimation according to an embodiment. This embodiment proposes to weight the size of the blocks matching search used to calculate the motion estimation vectors field by the downscaling ratio defined by the subsampling decision. Some prior art approach proposes a fast hierarchical motion vector estimation algorithm using mean pyramid. Several layers (usually two L1 and L2, figure 14) of lowresolution are created for each image at frame F-1 and F, motion estimation is performed by searching the best NxN pixels block matching in the lowest resolution corresponding to the smallest Mean Absolute Difference luminance. The final estimation vectors field is refined through the consecutive layers of upper resolution. Refinement consists of initializing the block search at level LN by motion vectors calculated previously at level LN+1. The process ends when level LO is reached corresponding to the initial image resolution.
[0091] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0092] Various methods and other aspects described in this application can be used to modify modules, for example, the motion compensation module (270) of a video encoder 200 as shown in FIG. 2. Moreover, the present aspects are not limited to ECM and VVC, and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination.
[0093] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0094] Various implementations involve decoding. “Decoding,” as used in this application, may encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0095] Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application may encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream.
[0096] Note that the syntax elements as used herein are descriptive terms. As such, they donot preclude the use of other syntax element names.
[0097] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
[0098] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[0099] Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0100] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0101] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0102] It is to be appreciated that the use of any of the following“and / or”, and “at least oneof”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0103] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a quantization matrix for de-quantization. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0104] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0105] We describe a number of embodiments. Features of these embodiments can be provided alone or in any combination, across various claim categories and types. Further,embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:
Claims
CLAIMS1 . A method of video encoding, comprising: obtaining a plurality of subsampling ratio corresponding to a subsampling of a picture from a first resolution to a second resolution of a picture, wherein a first resolution is greater than a second resolution of the picture; applying to a picture to encode, an adaptive temporal filtering to form a plurality of filtered pictures, each filtered picture being associated to a subsampling ratio, wherein a strength of the adaptive temporal filtering is based on the subsampling ratio, and wherein the plurality of filtered pictures has a first resolution; determining an indication whether Reference Picture Resampling is to be applied to encode the picture; responsive to the indication that Reference Picture Resampling is to be applied to encode the picture, determining a subsampling ratio to be applied to encode the picture and subsampling the filtered picture associated to the determined subsampling ratio from the first resolution to the second resolution of the picture to form a subsampled picture; and encoding the subsampled picture.
2. The method of claim 1 , wherein responsive to the indication that Reference Picture Resampling is to be applied to encode the picture, the method comprises disabling the adaptive temporal filtering and subsampling the picture to encode from the first resolution to the second resolution to form a subsampled picture.
3. The method of claim 2, wherein a strength of the adaptive temporal filtering is further based on a coding hierarchy level of the picture to encode in a group of pictures.
4. The method of any one of claims 1 -3, wherein responsive to a coding hierarchy level of the picture to encode in a group of pictures being larger than a level, the method comprises disabling the adaptive temporal filtering for the picture to encode and subsampling the picture to encode from the first resolution to the second resolution to form a subsampled picture.
5. The method of any one of claims 1 -4, wherein a length of the adaptive temporal filtering is further based on the subsampling ratio.
6. The method of any one of claims 1 -5, wherein a length of the adaptive temporal filtering is decreased with an increase of the subsampling ratio.
7. A method of video encoding, comprising: determining an indication whether Reference Picture Resampling is to be applied to encode a picture; responsive to the indication that Reference Picture Resampling is to be applied to encode a picture, determining a subsampling ratio to be applied to encode the picture and subsampling the picture from a first resolution to a second resolution of the picture to form a subsampled picture, wherein the first resolution is greater than the second resolution of the picture; applying to the subsampled picture an adaptive temporal filtering to form a filtered pictures associated to a subsampling ratio, wherein a strength of the adaptive temporal filtering is based on the subsampling ratio; and encoding the subsampled picture.
8. The method of claim 7, wherein responsive to the indication that Reference Picture Resampling is to be applied to encode the picture, the method comprises disabling the adaptive temporal filtering for the subsampled picture.
9. The method of any one of claims 7-8, wherein a strength of the adaptive temporal filtering is further based on a coding hierarchy level of the picture to encode in a group of pictures.
10. The method of any one of claims 7-9, wherein responsive to a coding hierarchy level of the picture to encode in a group of pictures being larger than a level, the method comprises disabling the adaptive temporal filtering for the picture to encode.1 1 . The method of any one of claims 7-10, wherein a length of the adaptive temporal filtering is further based on the subsampling ratio.
12. The method of any one of claims 7-11 , further comprising: determining a motion prediction of a block of the picture to encode by applying motion estimation to the block, wherein responsive to an indication that Reference Picture Resampling is to be applied to encode a picture, a size of the block used in block matching search for motion estimation is weighted by a subsampling ratio; and encoding the block of the picture based on the motion prediction for the block.
13. An apparatus, comprising one or more processors, wherein said one or more processors are configured to perform the method of any of claims 1 -12.
14. A signal comprising video data, formed by performing the method of any one of claims 1 -12.
15. A computer readable storage medium having stored thereon instructions for video encoding according to the method of any one of claims 1 -12.
Citation Information
Patent Citations
Buffer management in subpicture decoding
US11553177B2
Color Component Processing In Down-Sample Video Coding
US20230047271A1