Temporal filter strength for reference picture resampling

By fine-tuning the temporal filter intensity and adapting the motion estimation block size according to the resolution ratio of the RPR mode during video encoding, the problem of insufficient reference image resampling adaptation in the prior art is solved, thus improving encoding efficiency and quality.

CN122460065APending Publication Date: 2026-07-24INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-12-16
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to effectively utilize temporal filtering processes to adapt to reference image resampling while achieving high compression efficiency, resulting in poor coding efficiency.

Method used

Adaptive temporal filtering is achieved by fine-tuning the temporal filter strength value according to the resolution ratio of the RPR mode during video encoding, and delaying the temporal filter when appropriate, or adapting the temporal filter strength after rescaling, in order to optimize the block size of motion estimation.

Benefits of technology

It improves the encoding and decoding efficiency of video, especially by significantly reducing the bit rate at high QP values, thereby improving encoding quality and compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122460065A_ABST
    Figure CN122460065A_ABST
Patent Text Reader

Abstract

In various embodiments, when reference picture resampling is applied to achieve better coding efficiency, methods and devices adapt temporal filtering of a noise-reduced input video prior to encoding according to a resolution ratio. According to one embodiment, temporal filtering strength values are fine-tuned at a pre-processing module according to the resolution ratio of the RPR mode. According to another embodiment, temporal filtering is deferred after rescaling when the RPR mode is enabled, and filtering strength values are fine-tuned according to the resolution ratio of the RPR mode. A third embodiment proposes to refine motion estimation in the encoder by setting the block size of motion estimation in temporal filtering according to the resolution ratio of the RPR mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims the benefit of European Patent Application No. 23307319.6, filed on 21 December 2023, which is incorporated herein by reference in its entirety. Technical Field

[0002] This embodiment generally relates to a method and apparatus for adapting a temporal filtering process to reference image resampling in video encoding. Background Technology

[0003] To achieve high compression efficiency, image and video codec schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Intra-frame or inter-frame prediction is usually used to leverage intra-frame or inter-frame image correlations, followed by transform, quantization, and entropy encoding / decoding of the differences between the original and predicted blocks (typically represented as prediction error or prediction residual). To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy encoding / decoding, quantization, transform, and prediction. Summary of the Invention

[0004] In various embodiments, methods and apparatus are disclosed for adapting temporal filtering of noise-reducing input video according to the resolution ratio before encoding when applying reference image resampling to achieve better encoding / decoding efficiency. According to one embodiment, the temporal filter intensity value is fine-tuned at the preprocessing module according to the resolution ratio of the RPR mode. According to another embodiment, when the RPR mode is enabled, temporal filtering is delayed after rescaling, and the filter intensity value is fine-tuned according to the resolution ratio of the RPR mode. A third embodiment proposes refining motion estimation in the encoder by setting the block size of motion estimation in temporal filtering according to the resolution ratio of the RPR mode.

[0005] According to a first aspect, a video coding method is disclosed, the method comprising: obtaining a plurality of resampling rates corresponding to downsampling an image from a first resolution to a second resolution, wherein the first resolution of the image is greater than the second resolution; applying an adaptive temporal filter to the image to be encoded to form a plurality of filtered images, each filtered image being associated with a resampling rate, wherein the intensity of the adaptive temporal filter is based on the resampling rate, and wherein the plurality of filtered images have a first resolution; determining an indication of whether to apply reference image resampling to encode the image; in response to the indication of applying reference image resampling to encode the image, determining a resampling rate to be applied to encode the image, and downsampling the filtered images associated with the determined resampling rate from the first resolution of the image to the second resolution to form a downsampled image; and encoding the downsampled image.

[0006] According to a second aspect, a video coding method is disclosed, comprising: determining an indication to apply reference image resampling to encode an image; in response to the indication to apply reference image resampling to encode the image, determining a resampling rate to be applied to encode the image, and downsampling the image from a first resolution to a second resolution to form a downsampled image, wherein the first resolution of the image is greater than the second resolution; applying an adaptive temporal filter to the downsampled image to form a filtered image associated with the resampling rate, wherein the strength of the adaptive temporal filter is based on the resampling rate; and encoding the downsampled image.

[0007] According to another embodiment, an apparatus for encoding video is provided, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform an encoding method according to any embodiment described herein.

[0008] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding method according to any embodiment described herein. One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for encoding or decoding video according to the methods described herein.

[0009] One or more embodiments also provide a computer-readable storage medium storing video data generated according to the method described above. One or more embodiments also provide methods and apparatus for transmitting or receiving video data generated according to the methods described herein. Attached Figure Description

[0010] Figure 1 The diagram illustrates a block diagram of a system in which various aspects of this embodiment can be implemented.

[0011] Figure 2 A block diagram illustrating an embodiment of a video encoder is shown.

[0012] Figure 3 A block diagram illustrating an embodiment of a video decoder is shown.

[0013] Figure 4 A block diagram illustrating an embodiment of a video encoder using RPR is shown.

[0014] Figure 5 A block diagram illustrating an embodiment of a video decoder using RPR is shown.

[0015] Figure 6 The illustration shows the change in image resolution when using RPR for inter-frame prediction.

[0016] Figure 7a and Figure 7b The image sorting in the GOP example is illustrated.

[0017] Figure 8 The diagram illustrates the RD curve (anchor relative to RPR).

[0018] Figure 9 The illustration shows a method with preprocessing time filtering that can be applied to at least one embodiment.

[0019] Figure 10 The illustration shows a method for adapting time filtering to RPR mode according to an embodiment.

[0020] Figure 11 The illustration shows a method for adapting time filtering to RPR mode according to an embodiment.

[0021] Figure 12 The illustration shows a method for adapting time filtering to RPR mode according to an embodiment.

[0022] Figure 13 The illustration shows a method for adapting time filtering to RPR mode according to an embodiment.

[0023] Figure 14 The illustration shows a resolution layer for motion estimation according to an embodiment. Detailed Implementation

[0024] Figure 1 The diagram illustrates an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0025] System 100 includes at least one processor 110 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0026] System 100 includes an encoder / decoder module 130, which is configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. Encoder / decoder module 130 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100, or it may be incorporated within processor 110 as a combination of hardware and software known to those skilled in the art.

[0027] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 11. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0028] In several embodiments, the memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external volatile memory, such as RAM, is used as working memory for video encoding / decoding operations, such as for MPEG-2, HEVC, or VVC.

[0029] As shown in block 105, inputs can be provided to the components of system 100 through various input devices. Such input devices include, but are not limited to, (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0030] In various embodiments, the input device of block 105 has associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) band-limiting the narrower band again to select, for example, a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0031] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, in a separate input processing IC or within processor 110. Similarly, as needed, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.

[0032] Various components of system 100 can be provided within an integrated housing, in which the various components can be interconnected and transmit data between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit board.

[0033] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0034] In various embodiments, a Wi-Fi network, such as IEEE 802.11, is used to stream data to system 100. The Wi-Fi signals in these embodiments are received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, with the set-top box transmitting data via an HDMI connection to input block 105. Still other embodiments use an RF connection to input block 105 to provide streaming data to system 100.

[0035] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of embodiments, other peripheral devices 185 include one or more of a standalone DVR, disc player, stereo system and / or lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, signaling is used to transmit control signals between system 100 and the display 165, speaker 175, or other peripheral devices 185. Signaling may include AV.Link, CEC, or other communication protocols capable of enabling device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. The display 165 and speaker 175 can be integrated into a single unit within an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0036] For example, if the RF portion of input 105 is part of a standalone set-top box, then display 165 and speaker 175 can alternatively be separated from one or more other components. In various embodiments where display 165 and speaker 175 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0037] Figure 2 The illustration shows an example video encoder 200, such as a VVC (Various Video Codec) encoder. Figure 2 The diagram can also show encoders that have improved upon the VVC standard or encoders that use similar VVC technology.

[0038] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "encoding" and "encoding / decoding," and the terms "image," "picture," and "frame" are used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0039] Before encoding, the video sequence may undergo precoding (201), for example, applying color transformations to the input color image (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a signal distribution more resilient to compression. Precoding (201) may also include temporal filtering to reduce noise present in the video sequence and simultaneously improve coding efficiency. Metadata may be associated with the preprocessing and attached to the bitstream.

[0040] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed in units, for example, CUs. Each unit is encoded using, for example, intra-frame or inter-frame modes. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. After prediction, prediction enhancement (285) is applied to the prediction block. For example, the prediction residual is calculated by subtracting (220) the prediction block from the original image block.

[0041] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded / decoded (245) to output a bitstream. The encoder can skip the transform and directly quantize the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode and decode the residual without applying the transform or quantization process.

[0042] The encoder decodes the coded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (280).

[0043] Figure 3 A block diagram of an example video decoder 300 is shown. In decoder 300, the bitstream is decoded by decoder elements, as described below. Video decoder 300 typically performs operations similar to... Figure 2 The encoding rounds described herein are the opposite of the decoding rounds. Encoder 200 typically also performs video decoding as part of the encoded video data.

[0044] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoding / decoding information. Picture partitioning information indicates how the picture is partitioned. Therefore, the decoder can partition (335) the picture based on the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. The prediction blocks can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375) (370). After prediction, prediction enhancement (390) is applied to the prediction blocks. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0045] The decoded image can undergo further post-decoding processing (385), such as inverse color transformation (e.g., from YcbCr 4:3:0 to RGB 4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding processing can use metadata derived in the pre-encoding process and signaled in the bitstream.

[0046] Reference image resampling (RPR) In the Universal Video Coding (VVC) / H.266 standard, the image-based rescaling feature used for video encoding and decoding is called Reference Picture Resampling (RPR). Given an original video sequence consisting of images of a certain size (width × height), the encoder can select which resolution (image size) to use to encode and decode each frame. Different Picture Parameter Sets (PPS) are encoded and decoded in the bitstream at possible image sizes, and the slice / image header indicates which PPS is used to decode the current Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit.

[0047] Figure 4 This is a functional block diagram of an example encoder (400) using RPR. The encoder (400) includes a downsampler (440), a core encoder (410), a resampler (430, a component of the motion compensator 450), and a decoder picture buffer (DPB, 420). The encoder (400) is used to encode the raw video into a codec video bitstream. In one aspect, in addition to the reference picture buffer (280) and the motion compensator (270) in... Figure 4 Apart from the decoded image buffer (420) and motion compensator (450) shown respectively, the core encoder (410) represents Figure 2 The encoder (200).

[0048] exist Figure 4In the example, for each input frame of the original video, the encoder (400) can choose to encode the frame at the original frame size or at a reduced frame size. If the encoder chooses to encode at a reduced frame size, the downsampler (440) downsamples the input frame before the core encoder (410) encodes it. For example, it can be downsampled to... The frame is subsampled (440) according to the resolution ratio to obtain the width. and height The frame size. The encoder (400) can decide whether to encode the frame at the original frame size or at a reduced frame size, for example, by comparing the encoding and decoding results of input frames at various resolutions, or based on the spatial and temporal activities associated with the encoding and decoding of the input frames of the original video. Therefore, a resampling operation is required when the image content from the current encoded frame and the corresponding content from the reference frame (required for inter-frame prediction of the content from the current encoded frame) are not at the same resolution. In VVC, the resampling operation and motion compensation are combined into a single step. This can be done by a resampler (430), which can be configured to change the resolution of the image content from the reference frame to a resolution that matches the resolution of the content from the current frame.

[0049] Figure 5 This is a functional block diagram of an example decoder (500) using RPR. The decoder (500) includes a core decoder (510), a resampler (530, a component of the motion compensator 550), a decoder picture buffer (520), and an upsampler (540). The decoder (500) is used to decode the raw video from a given codec video bitstream to produce an output video. In one aspect, in addition to the reference picture buffer (380) and the motion compensator (375) in... Figure 5 Apart from the decoded image buffer (520) and motion compensator (550) shown respectively, the core decoder (510) represents Figure 3 The decoder (300).

[0050] exist Figure 5 In the example, image content from one or more frames (stored in the codec image buffer 520), which is needed as a reference for the inter-frame predicted image content from the currently decoded frame, is resampled by the resampler (530) to a resolution matching the resolution of the currently decoded image content. To output a video frame decoded (500) at a reduced resolution, the video frame is upsampled by the upsampler (540) to the resolution of the original video, thus producing the output video. Note that... Figure 4 The downsampler (440) (e.g., used in the preprocessing stage) and Figure 5The upsampler (540) (for example, used in the post-processing stage) is not specified by the VCC standard.

[0051] For each frame, the encoder chooses whether to encode at the original resolution or a reduced resolution (e.g., image width / height divided by 2). This choice can be made, for example, by using two rounds of encoding or by considering spatial and temporal activity in the original image. Therefore, the decoded image buffer (DPB, 420, 520) can contain images with a different size than the current image.

[0052] In cases where a reference image in the DPB has a different size than the current image, the reference block is implicitly resampled (430, 530) (scaled up or scaled down) during the motion compensation process (270, 375) to construct the prediction block.

[0053] Figure 6 The illustration shows the application of RPR in encoding / decoding streams of consecutive images with different image resolutions. This feature is useful, for example, for adaptive streaming to adapt the bitrate to network constraints. Even if not specified in the specification, it is necessary to resample downsampled images at resolutions not indicated in the HLS to display all decoded images at the same (display or target) resolution on the terminal display, such as... Figure 5 As shown in the diagram (540).

[0054] Time filtering In existing technologies, video encoders such as the VVC test model encoder apply temporal filtering as a preprocessing step on the encoder side. This provides benefits in terms of coding efficiency by reducing video noise. Some existing methods also introduce motion-compensated temporal filtering (MCTF) using bilateral filters, which are performed within the group of pictures (GOP) structure of hierarchical video frame coding.

[0055] Figure 7a and Figure 7b The illustration shows the image ordering in a GOP example. In Random Access Configuration (RA) mode, a hierarchical group of pictures (GOP) structure is used to encode video sequences. Figure 7a and Figure 7b The image depicts an example of 16 frames per GOP. The images in the bitstream are based on... Figure 7a The GOP structure shown above, instead of based on Figure 7b The images shown above are encoded and decoded using either a point of view (POC) or display order. For example, as... Figure 7b As shown, images with 0 and 16 have a Time ID (TID) of 0, and as... Figure 7aAs shown, the encoding and decoding are performed with 0 (first) and 1 (second) respectively in the encoding and decoding order. The image with POC 8 has a TID of 1 and is encoded and decoded with 2 (third) in the encoding and decoding order.

[0056] Motion estimation (e.g., block-based motion compensation) can be used to compute motion compensation vectors. For this purpose, the luminance components of the two adjacent frames before and after the current frame to be encoded are downsampled twice to obtain three resolution layers. Motion estimation can be performed using an integer-pixel search of NxN pixel blocks. The sum of the squared differences between each block at different resolution layers and the displacement blocks in adjacent frames is obtained to find the optimal motion estimation vector. Motion compensation is then applied to create motion-compensated neighboring images.

[0057] Temporal filtering is applied only to images located at low encoding / decoding levels (low TID values). These images are referenced most frequently and encoded with higher fidelity. For example... Figure 7a and Figure 7b As depicted, it typically corresponds to the Time ID (TID) 0 and 1 at the GOP level. The filter strength values ​​can be set as follows: .

[0058] Bilateral filtering can be applied to each pixel of the source image I0 as follows: .

[0059] I n These are the filtered sample values, I0 is the original sample value, and I... r (i) represents the co-localization sample value in the adjacent image I after motion compensation, and w r (i,a) is the weight of neighboring image i when the number of available neighboring images is equal to a.

[0060] For a brightness sample, the weight w r (i, a) can be calculated as follows: in: .

[0061] For all other values ​​of i and a: .

[0062] For chroma samples, the weight w r (i, a) is calculated as follows: .

[0063] Video encoding and decoding performance evaluation Rate-distortion and complexity are two commonly used metrics for comparing the performance of video codecs.

[0064] Rate-distortion is used to measure compression efficiency and to show the relationship between bit rate (Equation 1) and the quality of the reconstructed video. Peak signal-to-noise ratio (PSNR) (Equation 2) is often used to evaluate quality, where... and Let (i, j) represent the pixel at position (i, j) in the original image I and the pixel at position (i, j) in the reconstructed image K, respectively. in .

[0065] Typically, the trade-off between bitrate and distortion is controlled by the quantization parameter (QP) input. Video codec performance is evaluated by performing bitrate and PSNR metrics at several QP values, forming a rate-distortion curve (RD curve). Traditionally, if one wants to compare the performance of two versions of the same video encoder within a QP range, the Bjøntegaard incremental bitrate (BD bitrate) metric is used to measure the quality difference between the average bitrate and the RD curve for each version of the encoder.

[0066] Figure 8 The diagram illustrates such an RD curve. In the following text, "RPR mode" refers to frame resizing (downsampling or rescaling). In the case of RPR versus non-RPR encoder versions (i.e., anchors), the use of RPR generally improves BD bitrate at high QP values. The intersection of the RD curves can be determined, such as... Figure 8 As shown. The crossover point, also known as the QP switch, is content-dependent. It doesn't have a uniform value and can vary from frame to frame or from frame group to frame. The QP switch can be used to determine whether to apply RPR mode; in this example, it's based on the PSNR performance metric. It has been shown that RPR provides a gain for low bitrate coding most of the time.

[0067] Therefore, when RPR mode is activated in the video encoder, the video frames are rescaled relative to the resolution (or downscaling ratio) determined by the subsampling decision process on the encoder side before encoding. In addition to the temporal filtering already implemented during the preprocessing module, downscaling also applies a low-pass filter. These two processes are applied independently, resulting in a strong filtering effect before the encoding process. For example, excessive high frequencies may be removed through filtering, which can impact the efficiency of the encoding / decoding process.

[0068] This document describes at least one method that adapts the temporal filtering process to resolution when the RPR mode is activated, in order to achieve better encoding and decoding efficiency. According to one embodiment, a method is disclosed for fine-tuning the temporal filter intensity value at the preprocessing module based on the resolution ratio of the RPR mode. According to another embodiment, a method is disclosed that, when the RPR mode is activated, the temporal filtering is delayed after rescaling, and the filter intensity value is fine-tuned based on the resolution ratio of the RPR mode. A third embodiment proposes refining the motion estimation in the encoder by setting the block size of motion estimation in temporal filtering according to the resolution ratio of the RPR mode.

[0069] A general coding method with preprocessing time filtering suitable for RPR In the first embodiment, the temporal filtering intensity at the preprocessing module is adapted to the resolution ratio of the RPR mode.

[0070] Figure 9 The illustration shows a method (900) with preprocessing temporal filtering that can be applied in at least one embodiment. As explained above, frames or images of the original input video (910) are first preprocessed by temporal filtering (920) before being sent to the encoder. Preprocessing temporal filtering has the advantage of reducing the impact of noise present in the video frames and thus improving video coding efficiency. Furthermore, a decision (930) is made within the encoder to apply image subsampling, and the rescaler (940) reduces the filtered frames by a reduction ratio (also known as subsampling rate, downsampling rate, resampling rate, reduction ratio, or resolution ratio) set by the subsampling decision module (930).

[0071] Figure 10 The illustration depicts a method (1000) for adapting temporal filtering to an RPR mode according to an embodiment. This method proposes optimizing the temporal filtering process by sharing a predefined downsampling ratio value between the subsampling decision and temporal filtering modules. In a first step, multiple subsampling rates r0 to r10 corresponding to subsampling the image from a first (original) resolution to a second (reduced) resolution are obtained. N For example, for a width corresponding to the first resolution. and height Frame size, The subsampling rate r0 results in a width corresponding to the second resolution. and height The frame size. In the variant, a subsampling rate of 1 can be considered, meaning the frame resolution remains unchanged (i.e., RPR mode is not activated). Therefore, the first (original) resolution of the image is greater than the second (reduced) resolution, but the variant can consider the first resolution of the image to be greater than or equal to the second resolution. The total intensity value s0(n) (where n represents the frame number in the display order) is weighted by each subsampling rate value, resulting in a time-filtered image from 1 to N, where N is the number of predefined subsampling rates.

[0072] In step 1020, an adaptive temporal filter is applied to the input image to form multiple filtered images (1030), each filtered image being associated with a given subsampling rate r. The intensity s0, s'0 of the adaptive temporal filter is based on the subsampling rates r0, r N The filtered images are then provided to a time filter to adapt the filter to the subsampling rate. Multiple filtered frames (1030) are input to the encoder. The encoder can determine (1040) an indication to apply reference image resampling to encode the image, and the subsampling rate to be applied to encode the image. In response to the indication to apply reference image resampling to encode the image and the subsampling rate to be applied to encode the image, a filtered image associated with the determined subsampling rate is selected (switch 2) for subsampling by the rescaler module (1050). In response to the indication to apply reference image resampling to encode the image and the subsampling rate to be applied to encode the image, a rescaled filtered image at a second resolution or a filtered image at a first resolution is selected (1060) for further processing by the encoding loop.

[0073] As detailed above, switch 2 module selects a filtered image corresponding to the reduction ratio determined by the subsampling decision module. When subsampling decision is activated, the selected filtered frame is rescaled and output to the encoder by switch 1 before encoding begins. Alternatively, when subsampling decision is disabled, a filtered frame with an initial intensity s0 is selected by switch 1 and presented to the encoder input. The new intensity parameter can be dependent on s0 and the reduction ratio r. i The function (Equation 4). From s0 and r i Strength parameters can be determined by, for example Figure 10 The lookup table shown is used to implement this, where s0 and r i These are the input parameters of the LUT, and s i ’ This is the output of the LUT corresponding to the new intensity weighted by a predefined reduction ratio. Equations (5) and (6) are examples of mapping functions that can be implemented in a lookup table. Where N is the predefined reduction ratio, r is is the reduction ratio of index i, and s0 is the initial intensity defined in Section 1.1. The goal of this solution is to fine-tune the filtering process performed by temporal filtering based on the image reduction achieved by the rescaler driven by subsampling decisions.

[0074] Figure 11 The illustration depicts a method for adapting time filtering to RPR mode according to a variant embodiment. In this variant, time filtering is disabled when RPR is enabled, for any r i The intensity value can be set to 0, so when RPR mode is activated, the original video frame 1110 is sent to the rescaler instead of the time-filtered frame. When the RPR mode decision is deactivated, a value such as... Figure 11 The intensity s0 shown filters the original image, and the filtered frame is sent to the encoder via selection via switch 1.

[0075] Figure 12 The illustration depicts a method for adapting temporal filtering to an RPR mode according to another variant embodiment. According to this variant, when the subsampling mode is activated, a weighted intensity s0 is applied in the temporal filtering module based on the frame TID. ’ Still s0. For example, frames in TID 0 and / or 1 can be represented using s0. ’ Temporal filtering is performed on one frame, while other frames are temporally filtered using s0. In fact, frames in the lowest sublayer TID are referenced the most, and this variant allows for fine-tuning of the temporal intensity of these frames, which has an impact. Figure 12 As shown, the POC counter value is also shared between the time filtering and subsampling decision modules.

[0076] As an example, for a 16-frame GOP with TID values ​​equal to 0 and 1, frames with POC values ​​of 0, 8, and 16 are represented by s0. ’ Temporal filtering is performed, while frames with a TID greater than 1 are temporally filtered using s0. Alternatively, s0 can be adjusted accordingly. ’ Setting it to 0 disables temporal filtering for a specific TID image. Then, when the subsampling decision is activated, the original image of that TID is redirected to the rescaler instead of the temporally filtered frame.

[0077] In a variant, as described above, the concept of adapting filter strength based on bitrate can be easily and directly extended to other parameters used to control filter action. For example, filter length (e.g., the number of consecutive images used to perform temporal filtering), or the temporal distance between the current image and the image used to filter it, can depend on the downscaling ratio. Typically, as the downscaling ratio increases (i.e., from 1 / 2 to 1 / 4), the filter length or temporal distance decreases. According to yet another variant, adaptation of motion interpolation, or adaptation of motion vector accuracy, can depend on the resolution ratio. Typically, as the downscaling ratio increases, motion vector accuracy increases. Furthermore, as motion vector accuracy increases, the set of motion interpolation filters can support more interpolation phases. Another example is adaptation used to derive the aforementioned weights. The parameters. For example, in the formula below, Sbase can depend on the reduction ratio (it decreases as the reduction ratio increases, and vice versa). l It can also depend on the reduction ratio (it increases as the reduction ratio increases).

[0078] .

[0079] A general coding method that applies time filtering after rescaling and adapts to RPR. In the second embodiment, the temporal filter intensity is adapted to the resolution of the RPR mode and applied after the rescaler module.

[0080] Figure 13 The illustration depicts a method for adapting temporal filtering to an RPR mode according to a second embodiment. This embodiment proposes implementing temporal filtering after the rescaler when subsampling is activated. This has the advantage of adapting the temporal filtering strength in runtime, since the downsampling ratio value is known and the subsampling decision has already been performed at this stage of processing.

[0081] According to this embodiment, in the preprocessing step, a temporal filter is applied to the input video image to form a filtered image. Both the filtered image and the original image are input to the encoder. The encoder can determine an indication whether to apply reference image resampling to encode the image. When RPR mode is enabled, the encoder can further determine the resampling rate to be applied to encode the image. Then, in response to the indication to apply reference image resampling to encode the image, the original image is downsampled from a first resolution to a second (reduced) resolution by a rescaler to form a downsampled image. In a subsequent step, the downsampled image is filtered using an adaptive temporal filter to form the filtered image. Any variations of the previous embodiments related to the adaptive temporal filter are compatible with the current design. For example, the strength of the adaptive temporal filter can be based on the resampling rate. Therefore, the temporal filter strength can be fine-tuned according to the reduction ratio selected by the subsampling decision. For example, the adaptive temporal filter strength can be obtained from the LUT by taking the resampling rate and the original filter strength s0 as input. The fine-tuned filter strength s0 ’ The output is a lookup table driven by the reduction ratio and the initial intensity s0. In another variation, the temporal filter intensity can be adapted to either the frame's POC or the frame's TID. When RPR mode is activated, for the variation corresponding to disabling temporal filtering (setting the intensity to 0), the reduced original frame, instead of the temporally filtered frame, is redirected to the encoder.

[0082] Furthermore, in response to the indication that reference image resampling will not be applied to the encoded image, a filtered image at a first resolution is selected for further processing by the encoding loop. Therefore, conversely, when subsampling decisions are disabled, the original image is filtered using intensity s0 and fed into the encoder input, as... Figure 13 As shown.

[0083] According to yet another example, for a variant using different temporal filter intensities based on image TID, the rescaled image within TID 0 and / or 1 is used with s0 ’ Temporal filtering is performed on some frames, while other frames are temporally filtered using s0. Fine-tuning the temporal intensity of these frames has an impact because frames in the lowest sub-layer TID are referenced the most. For example... Figure 13 As described, the POC counter value can be time-filtered for this purpose.

[0084] A general coding method adapted to motion estimation for RPR In another embodiment, motion prediction for blocks of the image to be encoded is determined by applying motion estimation to those blocks. In response to an indication that reference image resampling will be applied to encode the image, the block size used in the block matching search for motion estimation is weighted by a subsampling rate; and the block is encoded based on the motion prediction. Advantageously, this embodiment can be combined with any variations of the first and second embodiments of the previously described adaptive temporal filtering.

[0085] A hierarchical estimate of the motion vector field is computed for motion compensation used by temporal filtering. Its advantages include reducing noise present in the video sequence and simultaneously improving coding efficiency. Furthermore, the use of hierarchical motion estimation is more robust and allows for easier identification of large motions.

[0086] Figure 14 The illustration shows a resolution layer for motion estimation according to an embodiment. This embodiment proposes weighting the size of the block-matching search used to compute the motion estimation vector field by a reduction ratio defined by the subsampling decision. Some prior art methods propose a fast hierarchical motion vector estimation algorithm using a mean pyramid. Several low-resolution layers (typically two L1 and L2 layers) are created for each image at frame F-1 and frame F. Figure 14 Motion estimation is performed by searching for the best N×N pixel block match in the lowest resolution corresponding to the minimum mean absolute difference luminance. The final estimated vector field is refined through successive layers of higher resolution. Refinement involves initializing the block search at level LN using the motion vector previously calculated at level LN+1. The process ends when level L0, corresponding to the initial image resolution, is reached.

[0087] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, terms such as "first" and "second" can be used in various embodiments to modify elements, components, steps, operations, etc., for example, "first decoding" and "second decoding." Unless specifically required, the use of these terms does not imply an ordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0088] The various methods and other aspects described in this application can be used to modify the module, for example, such as Figure 2The motion compensation module (270) of the video encoder 200 is shown. Furthermore, this aspect is not limited to ECM and VVC, and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.

[0089] Various numerical values ​​are used in this application. Specific values ​​are for illustrative purposes only, and the aspects described are not limited to these specific values.

[0090] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of the process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, these processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. It will be clear, and is considered well understood by those skilled in the art, that the phrase “decoding process” is intended to specifically refer to a subset of operations or to refer to a broader decoding process, based on the context of the specific description.

[0091] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding” as used in this application can encompass, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream.

[0092] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0093] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), implementations of the discussed features can be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. These methods can be implemented, for example, in apparatus, such as a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0094] References to "an embodiment" or "an embodiment" or "an implementation" or "implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation," and any other variations appearing throughout this application in various places, do not necessarily all refer to the same embodiment.

[0095] Furthermore, this application may involve "determining" various information fragments. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory.

[0096] Furthermore, this application may involve "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0097] Furthermore, this application may relate to "receiving" various pieces of information. Like "access," receiving is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more. Moreover, "receiving" is generally referred to in one way or another during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0098] It should be understood that any use of the following “ / ”, “and / or”, and “…at least one of…”, such as in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to a large number of listed items.

[0099] Furthermore, as used herein, among other things, the word “signal” refers to indicating something to the corresponding decoder. For example, in some embodiments, the encoder signals the quantization matrix used for dequantization. In this way, in embodiments, the same parameters are used at both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, then signaling (implicit signaling) can then be used without transmission to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the word “signal” was mentioned above, the word “signal” can also be used as a noun herein.

[0100] It will be apparent to those skilled in the art that implementations can generate a wide variety of signals that are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a wide variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0101] Numerous embodiments are described. Features of these embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across various claim classes and types.

Claims

1. A video encoding method, comprising: Obtain multiple subsampling rates corresponding to subsampling an image from a first resolution to a second resolution, where the first resolution of the image is greater than the second resolution; An adaptive temporal filter is applied to the image to be encoded to form multiple filtered images, each of which is associated with a subsampling rate, wherein the intensity of the adaptive temporal filter is based on the subsampling rate, and wherein the multiple filtered images have a first resolution; Indicators for determining whether to apply reference image resampling to encode the image; In response to an instruction to apply reference image resampling to encode an image, a subsampling rate to be applied to encode the image is determined, and a filtered image associated with the determined subsampling rate from a first resolution to a second resolution of the image is subsampled to form a subsampled image. as well as Encode the subsampled images.

2. The method according to claim 1, wherein, In response to an instruction to apply reference image resampling to encode an image, the method includes disabling adaptive temporal filtering and subsampling the image to encode it from a first resolution to a second resolution to form a subsampled image.

3. The method according to claim 2, wherein, The strength of the adaptive temporal filter is further based on the encoding / decoding level of the images to be encoded in the image group.

4. The method according to any one of claims 1-3, wherein, In response to the fact that the encoding / decoding level of the images to be encoded in the image group is greater than a certain level, the method includes disabling adaptive temporal filtering of the images to be encoded and subsampling the images to be encoded from a first resolution to a second resolution to form subsampled images.

5. The method according to any one of claims 1-4, wherein, The length of the adaptive time filter is further based on the subsampling rate.

6. The method according to any one of claims 1-5, wherein, The length of the adaptive time filter decreases as the subsampling rate increases.

7. A video encoding method, comprising: Indicators for determining whether to apply reference image resampling to encode the image; In response to an instruction to apply reference image resampling to encode an image, a subsampling rate to be applied to encode the image is determined, and the image is subsampled from a first resolution to a second resolution to form a subsampled image, wherein the first resolution of the image is greater than the second resolution; An adaptive temporal filter is applied to the subsampled image to form a filtered image associated with the subsampling rate, wherein the strength of the adaptive temporal filter is based on the subsampling rate; and Encode the subsampled images.

8. The method according to claim 7, wherein, In response to an instruction to apply reference image resampling to encode an image, the method includes disabling adaptive temporal filtering of the subsampled images.

9. The method according to any one of claims 7-8, wherein, The strength of the adaptive temporal filter is further based on the encoding / decoding level of the images to be encoded in the image group.

10. The method according to any one of claims 7-9, wherein, In response to the fact that the encoding / decoding level of the images to be encoded in the image group is greater than a certain level, the method includes disabling adaptive temporal filtering of the images to be encoded.

11. The method according to any one of claims 7-10, wherein, The length of the adaptive time filter is further based on the subsampling rate.

12. The method according to any one of claims 7-11, further comprising: Motion prediction of the blocks is determined by applying motion estimation to the blocks of the image to be encoded, wherein, in response to an instruction to apply reference image resampling to encode the image, the size of the blocks used in the block matching search for motion estimation is weighted by a subsampling rate. as well as Block-based motion prediction encodes blocks of an image.

13. An apparatus comprising one or more processors, wherein the one or more processors are configured to perform the method of any one of claims 1-12.

14. A signal comprising video data, formed by performing the method of any one of claims 1-12.

15. A computer-readable storage medium having stored thereon instructions for video encoding according to any one of claims 1-12.