Method and apparatus for encoding / decoding video

By dynamically rescaling and using adaptive filter selection for pictures of varying sizes, the method addresses inefficiencies in video encoding and decoding, improving coding efficiency and reducing complexity.

JP7855600B2Active Publication Date: 2026-05-08INTERDIGITALCE PATENT HLDG SAS
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERDIGITALCE PATENT HLDG SAS
Filing Date
2022-02-22
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing video encoding and decoding methods face challenges in achieving high compression efficiency when dealing with pictures of different resolutions, particularly in scenarios where reference and current pictures have varying sizes, leading to suboptimal coding efficiency and increased computational complexity.

Method used

The method involves dynamically rescaling pictures by downsampling and upsampling using interpolation filters, allowing reconstruction of a first picture from a second picture of different sizes, and incorporating adaptive filter selection based on classification processes to enhance encoding and decoding efficiency.

Benefits of technology

This approach improves coding efficiency by allowing flexible resolution handling, reducing computational complexity, and optimizing the use of interpolation filters, thereby enhancing the overall performance of video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007855600000006
    Figure 0007855600000006
  • Figure 0007855600000007
    Figure 0007855600000007
  • Figure 0007855600000008
    Figure 0007855600000008
Patent Text Reader

Abstract

A method is provided for reconstructing at least a portion of a first picture from at least a portion of a second picture, the first picture and the second picture having different sizes. Reconstructing includes decoding the second picture from a bitstream and determining at least one first sample of the at least a portion of the first picture using at least one resampling filter applied to at least one second sample of the at least a portion of the decoded second picture. Corresponding devices are provided for reconstructing at least a portion of a first picture. Methods and corresponding devices for encoding / decoding video are provided, which include reconstructing at least a portion of a first picture from at least a portion of a second picture, the first picture and the second picture having different sizes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This embodiment generally relates to a method and apparatus for encoding or decoding video. Some embodiments relate to a method and apparatus for encoding or decoding video in which the original picture and the reconstructed picture are dynamically rescaled for encoding. [Background technology]

[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to leverage spatial and temporal redundancy within video content. Generally, intra-picture or inter-picture correlation is used to utilize intra-picture or inter-picture correlation, and the difference between the original and predicted blocks—often called the prediction error or prediction residual—is then transformed, quantized, and entropicated. To reconstruct the video, the compressed data is decoded by the reverse process corresponding to entropicating, quantizing, transforming, and predicting. [Overview of the project]

[0003] According to one embodiment, a method is provided for reconstructing at least a portion of a first picture from at least a portion of a second picture, wherein the first picture and the second picture have different sizes, and the reconstruction includes decoding the second picture from a bitstream and determining at least one first sample of the first picture using at least one resampling filter applied to at least one second sample of the at least portion of the decoded second picture.

[0004] According to another embodiment, there is a device for reconstructing at least a portion of a first picture from at least a portion of a second picture, the device comprising one or more processors, the one or more processors configured to decode the second picture from a bitstream and to determine at least one first sample of the first picture using at least one resampling filter applied to at least one second sample of the at least portion of the decoded second picture, wherein the first picture and the second picture are of different sizes.

[0005] According to another embodiment, a video encoding method is provided, which includes encoding a second picture in a bitstream, the second picture being a downscaled picture from a first picture, and encoding a third picture in a bitstream, the third picture being the same size as the first picture, wherein encoding the third picture includes reconstructing at least a portion of the first picture by upsampling at least a portion of the second picture after decoding, the upsampling includes determining at least one first sample of the first picture using at least one upsampling filter applied to at least one second sample of the at least portion of the decoded second picture.

[0006] According to another embodiment, an apparatus for video encoding is provided, comprising one or more processors, the one or more processors configured to encode a second picture in a bitstream, the second picture being a downscaled picture from a first picture, and a third picture in a bitstream, the third picture being the same size as the first picture, wherein encoding the third picture includes reconstructing at least a portion of the first picture by upsampling at least a portion of the second picture after decoding, the upsampling includes determining at least one first sample of the first picture using at least one upsampling filter applied to at least one second sample of the at least portion of the decoded second picture.

[0007] According to another embodiment, a video decoding method is provided, which includes decoding a second picture in a bitstream, the second picture being a downscaled picture from a first picture, and decoding a third picture in a bitstream, the third picture being the same size as the first picture, wherein decoding the third picture includes reconstructing at least a portion of the first picture by upsampling at least a portion of the second picture after decoding, the upsampling includes determining at least one first sample of the first picture using at least one upsampling filter applied to at least one second sample of the at least portion of the decoded second picture.

[0008] According to another embodiment, a device for video decoding is provided, comprising one or more processors, the one or more processors configured to decode a second picture in a bitstream, the second picture being a downscaled picture from a first picture, and a third picture in a bitstream, the third picture being the same size as the first picture, wherein decoding the third picture includes reconstructing at least a portion of the first picture by upsampling at least a portion of the second picture after decoding, the upsampling includes determining at least one first sample of the first picture using at least one upsampling filter applied to at least one second sample of the at least portion of the decoded second picture.

[0009] In one variant, a method for encoding / decoding video includes storing at least a reconfigured portion of a first picture in a decoded picture buffer that stores a reference picture for encoding a third picture.

[0010] In another embodiment, a method for encoding video is provided, the method comprising: classifying a sample of a first picture; determining, based on the classification, a first filter for at least a portion of the first picture, which is used in a first encoding operation using the at least a portion of the first picture; providing a first modified portion of the first picture; and determining, based on the classification, a second filter for which the second filter is used in a second encoding operation using the first modified portion of the first picture.

[0011] An apparatus for encoding video is provided. The apparatus comprises one or more processors, which encode video by classifying samples of a first picture, and for at least a portion of the first picture, determine a first filter, based on the classification, which is used in a first encoding operation using that at least a portion of the first picture, and provide a first modified portion of the first picture, and determine a second filter, based on the classification, which is used in a second encoding operation using that first modified portion of the first picture.

[0012] In another embodiment, a method for decoding a video is provided, the method comprising: classifying a sample of a first picture; determining, based on the classification, a first filter for at least a portion of the first picture, which is used in a first decoding operation using the at least a portion of the first picture; providing a first modified portion of the first picture; and determining, based on the classification, a second filter for which the second filter is used in a second decoding operation using the first modified portion of the first picture.

[0013] A device for decoding video is provided. The device comprises one or more processors, which are configured to decode video, and decoding video includes classifying a sample of a first picture; determining, based on the classification, a first filter for at least a portion of the first picture, which is used in a first decoding operation using that at least a portion of the first picture; providing a first modified portion of the first picture; and determining, based on the classification, a second filter for which the second filter is used in a second decoding operation using that first modified portion of the first picture.

[0014] According to one embodiment of any one of the above aspects, the classification is stored in a decoded picture buffer that stores a reference picture, i.e., the index associated with each sample of the first picture is stored in the decoded picture buffer.

[0015] In another embodiment, there is a method for encoding video, the method comprising: classifying samples of a reference picture; determining at least a portion of the reference picture for at least one block of video using at least one motion vector of at least one block; determining at least one interpolation filter for at least a portion of the reference picture based on the classification; determining a prediction of the block based on filtering of the at least portion of the reference picture using the determined at least one interpolation filter; and encoding the block based on the prediction.

[0016] A device for encoding video is provided, comprising one or more processors configured to encode video by classifying samples of a reference picture; determine at least a portion of a reference picture for at least one block of video using at least one motion vector of at least one block; determine at least one interpolation filter for at least a portion of the reference picture based on the classification; determine a prediction of the block based on filtering of that portion of the reference picture using the determined at least one interpolation filter; and encode the block based on the prediction.

[0017] According to another aspect, there is provided another method for decoding a video, wherein decoding the video includes classifying samples of a reference picture, determining at least a part of the reference picture for at least one block of the video using at least one motion vector of the at least one block, determining at least one interpolation filter based on the classification for at least a part of the reference picture, determining a prediction of the block based on filtering of the at least a part of the reference picture using the determined at least one interpolation filter, and decoding the block based on the prediction.

[0018] An apparatus for decoding a video, comprising one or more processors configured to decode the video by classifying samples of a reference picture, determine at least a part of the reference picture for at least one block of the video using at least one motion vector of the at least one block, determine at least one interpolation filter based on the classification for at least a part of the reference picture, determine a prediction of the block based on filtering of the at least a part of the reference picture using the determined at least one interpolation filter, and decode the block based on the prediction.

[0019] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform a reconstruction method, or an encoding method or a decoding method according to any of the embodiments described herein. One or more of these embodiments also provide a computer-readable storage medium storing instructions for reconstructing a portion of a picture, encoding video data, or decoding video data according to the above method. One or more embodiments also provide a computer-readable storage medium storing a bitstream generated by the method described above. One or more embodiments also provide a method and an apparatus for transmitting or receiving a bitstream generated according to the method described above.

Brief Description of the Drawings

[0020] [Figure 1] A block diagram of a system in which aspects of this embodiment can be implemented is shown. [Figure 2] A block diagram of one embodiment of a video encoder is shown. [Figure 3] A block diagram of one embodiment of a video decoder is shown. [Figure 4] An exemplary method for encoding video according to one embodiment is shown. [Figure 5] An exemplary method for reconstructing video according to one embodiment is shown. [Figure 6] An example of motion compensation of a current block in a current picture in a reference picture when the reference picture has a different resolution from the current picture according to one embodiment is shown. [Figure 7] An example of determination of filter coefficient values as a function of sample phase according to one embodiment is shown. [Figure 8] An example of two-stage motion compensation filtering according to one embodiment is shown. [Figure 9] An example of horizontal filtering in the first stage of motion compensation filtering according to one embodiment is shown. [Figure 10]This figure shows an example of vertical filtering in the second stage of motion-compensated filtering according to one embodiment. [Figure 11] Examples of symmetric filters and filter rotations are shown. [Figure 12] An example of a method for determining an upsampling filter according to one embodiment is shown. [Figure 13] An example of a method for encoding / decoding a picture according to one embodiment is shown. [Figure 14A] An example of different phases corresponding to two types of upsampling in the horizontal and vertical directions, according to one embodiment, is shown. [Figure 14B] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14C] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14D] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14E] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14F] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14G] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14H] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 14I] Examples of different shapes of upsampling filters according to the embodiment are shown. [Figure 15] An example of a method for determining the upsampling filter coefficient according to one embodiment is shown. [Figure 16] An example of a method for encoding video according to one embodiment is shown. [Figure 17] An example of a method for decoding video according to one embodiment is shown. [Figure 18]An example of a method for encoding / decoding video according to one embodiment is shown. [Figure 19] An example of a method for encoding / decoding video according to another embodiment is shown. [Figure 20] An example of a method for encoding / decoding video according to another embodiment is shown. [Figure 21] An example of a method for decoding video according to another embodiment is shown. [Figure 22] This example illustrates two remote devices communicating via a communication network based on this principle. [Figure 23] The syntax of a signal based on this principle is shown below. [Modes for carrying out the invention]

[0021] This application describes various embodiments, including tools, features, embodiments, models, and methods. Many of these embodiments are described specifically, often in a way that may sound restrictive, at least to illustrate their individual characteristics. However, this is for the purpose of clarifying the description and not to limit the application or scope of those embodiments. In practice, all of the different embodiments can be combined and substituted to provide further embodiments. Furthermore, these embodiments can also be combined and substituted with embodiments described in previous applications.

[0022] The embodiments described and intended in this application can be implemented in many different forms. Figures 1, 2, and 3 below provide several embodiments, but other embodiments are intended, and the considerations in Figures 1, 2, and 3 are not intended to limit the scope of implementation forms. At least one of the embodiments generally relates to video encoding and decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream. These and other embodiments can be implemented as a computer-readable storage medium storing in itself instructions for encoding or decoding video data according to any of the described methods, apparatus, and / or a computer-readable storage medium storing in itself a bitstream generated according to any of the described methods.

[0023] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably.

[0024] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the proper operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. In addition, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as “first decoding” and “second decoding.” The use of such terms does not imply any ordering of the modified operations unless specifically required. Therefore, in this embodiment, the first decoding does not need to be performed before the second decoding, and may occur, for example, before the second decoding, during the second decoding, or during a time overlapping with the second decoding.

[0025] Modules of the video encoder 200 and decoder 300, such as motion compensation modules (270, 375) as shown in Figures 2 and 3, can be modified using the various methods and other embodiments described herein. Furthermore, embodiments of this disclosure are not limited to VVC or HEVC and can be applied to other standards and recommendations, whether existing or future, and to extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, embodiments described herein can be used individually or in combination.

[0026] Figure 1 shows a block diagram of an example of a system in which various embodiments and forms may be implemented. System 100 may be embodied as a device comprising various components described below and configured to perform one or more of the embodiments described herein. Examples of such a device include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 100 may be embodied individually or in combination as a single integrated circuit, a plurality of ICs, and / or individual components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of System 100 are distributed across a plurality of ICs and / or individual components. In various embodiments, System 100 is communicably coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, System 100 is configured to implement one or more of the embodiments described in this application.

[0027] System 100 includes, for example, at least one processor 110 configured to execute internally loaded instructions to implement various embodiments described in this application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits as known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include, but is not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drives, and / or optical disk drives, as well as non-volatile and / or volatile memory. The storage device 140 may, in non-limiting examples, include an internal storage device, a mounted storage device, and / or a network-accessible storage device.

[0028] System 100 includes, for example, an encoder / decoder module 130 configured to process data and provide encoded or decoded video, the encoder / decoder module 130 of which may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in the device to perform encoding and / or decoding functions. As is known, the device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of System 100, or it may be incorporated into the processor 110 as a combination of hardware and software, as is known to those skilled in the art.

[0029] Program code loaded onto the processor 110 or encoder / decoder 130 to perform various embodiments described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0030] In some embodiments, the internal memory of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory outside the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, H.265 and MPEG-H Part 2 are also known), or VVC (Variable Video Coding, a new standard under development by the Joint Video Experts Team (JVET)).

[0031] Inputs to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) a radio frequency (RF) section for receiving RF signals transmitted over the entire broadcast by a broadcaster, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Other embodiments, not shown in Figure 1, include composite video.

[0032] In various embodiments, the input device of block 105 has associated input processing elements, as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in a particular embodiment, band-limiting it again to a narrower frequency band in order to select a signal frequency band that may be referred to as a channel, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) multiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs these various functions, e.g., down-converting a received signal to a lower frequency (e.g., an intermediate frequency or adjacent baseband frequency) or to the baseband. In one embodiment of the set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, down-converting, and again filtering the RF signal transmitted over a wired (e.g., cable) medium to a desired frequency band. Various embodiments may involve rearranging the order of the elements described above (and others), removing some of these elements, and / or adding other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0033] In addition, USB and / or HDMI terminals may include their respective interface processors for connecting System 100 to other electronic devices across the entire USB and / or HDMI connection. It should be understood that various forms of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within Processor 110, as needed. Similarly, forms of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within Processor 110, as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, for example, Processor 110 and an encoder / decoder 130 that operates in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0034] Various elements of system 100 may be provided within an integrated housing, where the various elements are interconnected using internal buses known in the art, such as a suitable connection configuration 115, including an I2C bus, wiring, and a printed circuit board, and can transmit data between them.

[0035] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, transceivers configured to transmit and receive data via the communication channel 190. The communication interface 150 may also include, but is not limited to, a modem or network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0036] In various embodiments, the data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received on a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streaming data is provided to system 100 using a set-top box that distributes data via an HDMI connection on input block 105. In yet another embodiment, streaming data is provided to system 100 using an RF connection on input block 105. As shown above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0037] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various embodiments, the display 165 includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 may be for a television, tablet, laptop, mobile phone, or other device. The display 165 may also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). Other peripheral devices 185, in various embodiments of the embodiment, include one or more of a standalone digital video disc (or digital versatile disc) (DVR for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functionality based on the output of System 100. For example, the disc player performs the function of playing the output of system 100.

[0038] In various embodiments, control signals are communicated between the system 100 and the display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicably coupled to the system 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to the system 100 via communication interface 150 and communication channel 190. The display 165 and speaker 175 may be integrated into a single unit with other components of the system 100 in an electronic device such as a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0039] The display 165 and speaker 175 can, alternatively, be isolated from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0040] The embodiments can be implemented by a processor 110, by hardware, or by a combination of hardware and software, or by computer software. In a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120 can be of any type appropriate to the technical environment and, in a non-limiting example, can be implemented using any suitable data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and live-bubble memory. The processor 110 can be of any type appropriate to the technical environment and, in a non-limiting example, can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0041] Figure 2 shows the encoder 200. Although variations of this encoder 200 are also considered, for the sake of clarity, the encoder 200 will be described below without explaining all of the expected variations.

[0042] In some embodiments, Figure 2 also shows encoders that employ HEVC-like technologies, such as encoders that are improvements on the HEVC standard, or VVC (Versatile Video Coding) encoders currently under development by JVET (Joint Video Exploration Team).

[0043] Before encoding, the video sequence may undergo pre-encoding processing (201), such as applying a color conversion to the input color picture (e.g., converting from RGB4:4:4 to YCbCr4:2:0), or performing a remapping of the input picture components to obtain a signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0044] In encoder 200, the picture is encoded by encoder elements as described below. The picture to be encoded is divided into units, for example, CUs (202), and processed. Each unit is encoded using either intra-mode or inter-mode, for example. When a unit is encoded in intra-mode, it performs intra-prediction (260). In inter-mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines whether to use intra-mode or inter-mode to encode a unit (205), and indicates the intra / inter decision, for example, by a prediction mode flag. The encoder may also mix the intra-prediction results and the inter-prediction results (263), or mix results from different intra / inter-prediction methods. The prediction residual is calculated, for example, by subtracting the predicted blocks from the original image blocks (210).

[0045] The motion enhancement module (272) uses an already available reference picture to enhance the motion field of a block without referencing the original block. The motion field for a region can be thought of as the set of motion vectors for all pixels that make up that region. If the motion vectors are subblock-based, the motion field can also be represented as the set of all subblock motion vectors within the region (all pixels within a subblock have the same motion vector, although the motion vectors may differ for each subblock). If a single motion vector is used for a region, the motion field for that region can also be represented by a single motion vector (the same motion vector for all pixels within the region).

[0046] The predicted residual is then transformed (225) and quantized (230). The quantized transformation coefficients, as well as the motion vector and other syntax elements, are entropi-coded to output a bitstream (245). The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is coded directly without applying either the transformation or quantization process.

[0047] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transformation coefficients are inversely quantized (240) and inversely transformed (250) to decode the prediction residuals. The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, non-blocking / sample adaptive offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).

[0048] Figure 3 shows a block diagram of the video decoder 300. In the decoder 300, the bitstream is decoded by the decoder elements, as described below. The video decoder 300 generally performs a decoding pass which is the reverse of the encoding pass, as shown in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.

[0049] In particular, the input to the decoder includes a video bitstream, which may be generated by a video encoder 200. The bitstream is first entropy-decoded to obtain transformation coefficients, motion vectors, and other coded information (330). Picture segmentation information indicates how the picture is segmented. The decoder can therefore segment the picture according to the decoded picture segmentation information (335). The transformation coefficients are inversely quantized (340) and inversely transformed (350) to decode the predicted residuals. The image blocks are reconstructed by combining the decoded predicted residuals and the predicted blocks (355).

[0050] A prediction block can be obtained from intra-prediction (360) or motion-compensated prediction (i.e., inter-prediction) (375) (370). The decoder may mix the intra-prediction result and the inter-prediction result (373), or mix the results from multiple intra / inter-prediction methods. Before motion compensation, the motion field may be improved by using an already available reference picture (372). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0051] The decoded picture may undergo further post-decoded processing (385), such as inverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4), or inverse remapping, which performs the reverse of the remapping process performed in pre-encoded processing (201). The post-decoded processing may use metadata derived in pre-encoded processing and signaled in the bitstream.

[0052] Resampling of reference picture At low bitrates and / or when the picture has few high frequencies, for a better coding efficiency trade-off, typically for 4K or 8K frames, the picture can be encoded downsized rather than at full resolution. The decoder is then responsible for upscaling the decoded picture before display. The principle of Reference Picture Resampling (RPR) is to dynamically rescale the images in a video sequence on a picture basis for a better coding efficiency trade-off.

[0053] Figures 4 and 5 show an example of a method for encoding (400) and decoding (500) a video, respectively, according to one embodiment in which the image to be encoded can be rescaled for encoding. For example, such encoders and decoders can conform to the VVC standard.

[0054] Given an original video sequence consisting of pictures of size (picWidth × picHeight), the encoder selects a resolution (i.e., picture size) for coding a frame for each original picture. Different PPS (Picture Parameter Sets) are coded in the bitstream with the picture size, and the slice / picture header of the picture to be decoded indicates which PPS the decoder should use to decode the picture.

[0055] The downsampler (440) and upsampler (540) functions used as pre-processing and post-processing, respectively, are not specified by the standard.

[0056] For each frame, the encoder chooses whether to encode it at the original resolution or at a downsized resolution (for example, by dividing the picture's width / height by two). This choice can be made using two-pass encoding or by considering the spatial and temporal activity in the original picture.

[0057] When the encoder chooses to encode the original picture at a downsized resolution, the original picture is downscaled (440) before being input to the core encoder (410) to generate a bitstream. The picture, reconstructed at the downscaled resolution, is then stored in a decoded picture buffer (DPB) for coding subsequent pictures (420). As a result, the decoded picture buffer (DPB) may contain pictures of a different size than the current picture size.

[0058] In the decoder, the picture is decoded from the bitstream (510), and the reconstructed picture at a downscaled resolution is stored in a decoded picture buffer (DPB) for decoding subsequent pictures (520). In one embodiment, the reconstructed picture is upsampled to its original resolution (540) and transmitted, for example, to a display.

[0059] According to one embodiment, when the current picture to be encoded uses a reference picture from a DPB having a different size than the current picture, rescaling (430 / 530) (upscaling or downscaling) of the reference block for constructing the prediction block is performed (on the fly) during the motion compensation process using separable (horizontal and vertical) interpolation filters and appropriate sampling. Figure 6 shows an example of motion compensation with implicit block resampling that can be implemented in the rescaling (430 / 530) of the encoding and decoding methods discussed above. The selection of filter coefficients is based on the phase (θ x ,θ y The phase depends on the position of the sample to be interpolated in the reference picture, and in this case (Equation 1) (Figure 6), it depends on both the motion vector and the sizes of both the reference picture (620 in Figure 6) (SXref, SYref) and the current picture (610 in Figure 6) (SXcur, SYcur).

[0060] To predict the current block prediction P(610) of size (SXcur,SYcur), for each sample Xcur of P, its position (Xref,Yref) in the reference picture is determined. The value of (Xref,Yref) is a function of the current block's motion vector (MVx,MVy) and the scaling ratio between the current block size and the corresponding region (SXref,SYref) in the reference picture (620).

[0061] As shown in Figure 6, the phase, which is the non-integer part of the motion-compensated point (Xref, Yref) in the reference picture, is denoted as (θx, θy). The position (Xref, Yref) and phase (θx, θy) are given by the following equations. Xref=int(SXref×(MV X +Xcur) / SXcur) Yref = int(SYref × (MV Y +Ycur) / SYcur) (Formula 1) θ x =(SXref×(MV X +Xcur) / SXcur)-Xref θ y =(SYref×(MV Y +Ycur) / SYcur)-Yref int(x) gives the integer part of x.

[0062] In one embodiment, motion compensation (MC) uses two separate 1D filters to reduce computational complexity (Figure 7). The MC process is performed in two stages, as shown in Figures 8, 9, and 10, namely, a first horizontal (820, 900) and then a second vertical (840, 1000) motion compensation filtering, or in one variant, the vertical motion compensation filtering may be performed first, followed by the horizontal motion compensation filtering.

[0063] Figure 8 shows an example of two-stage motion compensation filtering according to one embodiment. The block position (Xref, Yref) and phase (θx, θy) in the reference picture are determined from the block position (XCur, YCur) and the motion vector (MVx, MVy) of the current block in the current picture (810). According to one embodiment, horizontal filtering using a 1D filter (shown in Figure 9) is performed to determine motion-compensated samples that have been upscaled along the horizontal direction (820, 940).

[0064] In one embodiment, since the motion vector has sub-Pel accuracy, there are as many 1D filters as there are sub-Pel positions (phases). Figure 7 shows how the filter coefficients w(i) are determined depending on the phase of the motion-compensated sample Xcur. The reconstructed sample "rec" is calculated using 1D filtering as follows.

[0065]

number

[0066] The reconstructed samples are stored in a temporary buffer of the same size (SXcur,SYref) (930 in Figure 9) (830). Then, to determine the motion-compensated samples upscaled along the vertical direction, vertical filtering is performed using a 1D filter with the temporary buffer as input, as shown in Figure 10 (840).

[0067] Note that the initial vertical filtering and the subsequent horizontal filtering are separate filters, so it is possible to perform both simultaneously.

[0068] The obtained prediction samples are stored in blocks of size (SXcur,SYcur) (1050) (850).

[0069] The above explanation assumes that the current picture and the referenced picture correspond to the same window. This means that if the motion is 0, the top-left and bottom-right samples of the two pictures correspond to the two same scene points. Otherwise, an offset window parameter should be added to (Xref, Yref).

[0070] The motion compensation with implicit resampling described above allows for the reuse of interpolation filters designed for classical motion compensation, such as those used in the VVC standard. Furthermore, this process avoids the need to store reference pictures at multiple resolutions. However, the simplicity of the upsampling filter limits the encoder's compression efficiency; therefore, improvement is needed.

[0071] In one embodiment, a method is provided for reconstructing at least a portion of a first picture from at least a portion of a second picture, wherein the first picture and the second picture have different sizes. For example, the second picture has a lower resolution than the first picture. According to this embodiment, reconstructing a portion of the first picture includes decoding the second picture from a bitstream and determining at least one first sample of the first picture from at least a portion of the decoded second picture using at least one upsampling filter applied to at least one second sample of the decoded at least portion of the second picture.

[0072] In one embodiment, the method for reconstruction includes transmitting the reconstructed at least portion of the first picture to a display. In one embodiment, the steps of the reconstruction method provided below can be carried out in the decoding method (510, 540) described with reference to Figure 5.

[0073] According to one embodiment, a method for reconstruction can be implemented in an encoding or decoding method. At least a portion of the first picture is obtained by decoding a second picture and upsampling at least a portion of the second picture, as described below. The reconstructed at least a portion of the first picture is then stored in a decoded picture buffer for future use as a reference picture when coding / decoding subsequent pictures of the same or different size as the first picture.

[0074] Several embodiments are provided below in which filter parameters are determined. The filter parameters include upsampling filter coefficients, associated tap locations (shapes), and optionally an index for identifying the filter. Any one of the embodiments provided below can be implemented alone or in combination with one or more of the other embodiments in methods for reconstructing, encoding, and / or decoding the pictures provided above.

[0075] According to one embodiment, the upsampling filter is not separable. In this embodiment, the upsampling filter cannot be processed by two-step upsampling using a 1D filter. The filter may be linear or nonlinear.

[0076] According to another embodiment, the upsampling filter coefficients are coded in the bitstream. In one variant, the upsampling filter coefficients can be coded even if the reference picture and the current picture have the same size. In the bitstream, the size of the original picture (after upsampling) is coded. The size of the original picture can be a parameter associated with the upsampling filter. The upsampling filter coefficients and / or the original size can be coded, for example, in the APS (e.g., the Adaptation Parameter Set used in the VVC standard to transmit Adaptive Loop Filter coefficients), slice header, picture header, or PPS. Default values ​​for the upsampling filter coefficients that are not coded in the bitstream may exist.

[0077] Filter coefficients can be derived for each picture, for each region within a single picture, for each group of pictures, or for each region within different pictures.

[0078] Figure 12 shows an example of a method 1200 for determining an upsampling filter according to one embodiment. Several upsampling filters may be available. The selection of the upsampling filter to be used may be controlled by the classification process.

[0079] According to one variant, when upsampling is within a motion compensation loop for predicting the current picture, upsampling of the reference picture used by the current picture is performed in response to a determination (1210) that the resolution of the reference picture is lower than that of the current picture.

[0080] The classification process determines a class index for each reference sample or group of reference samples (e.g., a group of 4x4 samples) (1220). One filter is associated with one class index. In the example in Figure 14A showing the interpolation region, the black samples represent the reference sample for which the class index has been determined and examples of the interpolation samples (1,2,3).

[0081] Each time interpolation is performed in an upsampled picture, a corresponding set of reference samples at the same location is determined. For example, Figure 14A shows an example of reference samples at the same location associated with sample 3 being interpolated (black samples in the dashed box). The class index associated with the reference samples at the same location of the sample being interpolated allows for the derivation of a single class index value for the sample being interpolated. For example, this could be the class index value of the reference sample at the closest location to the current sample being interpolated, or a predetermined relative position or mean / median of the class index values ​​of several reference samples at the same location.

[0082] For each sample to be interpolated, an upsampling filter is selected based on the class index derived for the sample to be interpolated (1230). Since classification is performed on the reference sample of the reference picture to be upsampled, or on the reference sample of the decoded picture in the case of upsampling for display, the class index value used to determine the upsampling filter for each sample to be interpolated does not need to be coded.

[0083] Next, an upsampling filter (1240) is applied to determine the value of the sample to be interpolated.

[0084] According to the embodiment, the classification process (1220) can be the same as that used in the Adaptive Loop Filter (ALF) in the VVC standard. The reconstructed sample "t(r)" is classified into K classes (K=25 for lumens samples, K=8 for chromens samples), and K different filters are determined using samples from each class. The classification is performed using directionality and activity values ​​derived using local gradients.

[0085] The above method 1200 can be applied, for example, when a picture is encoded in a downscaled version, decoded in a downscaled version, and then upsampled for output, for example for transmission to a display.

[0086] According to another embodiment, method 1200 can also be used to determine a downsampling filter that can be used to downsample a picture. For example, downsampling of a picture can be performed before encoding when the picture is to be encoded in a downscaled version.

[0087] Figure 13 shows an example of a method for encoding / decoding a picture according to one embodiment. According to this embodiment, it is determined whether to code or decode the current picture using interpretation (1305).

[0088] If the current picture is not coded / decoded using interpretation, the picture is coded / decoded using intrapretation, for example (1340).

[0089] When coding / decoding the current picture using inter prediction, it is determined (1310) whether the resolution of the reference picture is smaller than that of the current picture. If not, the current picture is coded / decoded using the reference picture stored in the DPB (1340). When the reference picture is larger in size than the current picture, downscaling is performed by the normal RPR (Reference Picture Resampling) motion interpolation process from the VVC standard during the encoding / decoding of the current picture.

[0090] When the reference picture is smaller in size than the current picture (1310), upscaling (1320) is performed with the upsampling filter determined according to any one of the embodiments proposed in this specification. Upsampling using a filter may be performed on-the-fly within the motion compensation process when encoding / decoding the current picture (1340), or the reference picture in the DPB may be upscaled (1320) before encoding / decoding the current frame (1340) and stored in the DPB (1330).

[0091] In this last case, the DPB can include several instances of reference pictures with different resolutions, and the motion compensation does not change compared to the encoding / decoding without RPR (1340).

[0092] According to one embodiment, the upsampling filter is a Wiener-based adaptive filter (WF). For example, the coefficients are determined in a similar way to the coefficients of the ALF in the VVC standard.

[0093] In VVC, the in-loop ALF filter (adaptive loop filtering) is a linear filter, and its purpose is to reduce coding artifacts for the reconstructed samples. The coefficient c nThis is determined by using a Wiener-based adaptive filtering technique to minimize the mean squared error between the original sample s(r) and the filtered sample t(r).

[0094]

number

[0095] To find the least squared sum error (SSE) between s(r) and f(r), c n We can determine the derivative of SSE with respect to and make that derivative equal to 0. Next, the coefficient value "c" is obtained by solving the following equation. [Tc].c T =v T (Formula 3) Here,

[0096]

number

[0097] In VVC, ALF coefficients can be coded in the bitstream so that they can dynamically adapt to the video content. There are also several default coefficients, and the encoder indicates which set of coefficients to use for each CTU.

[0098] In VVC, symmetric filters are used, as shown at the top of Figure 11, and some filters can be obtained from other filters by rotation, as shown at the bottom of Figure 11. Each coefficient in the filters shown at the top of Figure 11 is associated with one or two positions p(x,y). For example, the positions of c9 and c3 are represented as p9(0,0) and p3(0,-1) or p3(0,1). In the case of diagonal transformation, position p(x,y) is moved to p(y,x), in the case of vertical inversion transformation, position p(x,y) is moved to p(-x,y), and in the case of rotation, position p(x,y) is moved to p(y,-x).

[0099] According to one embodiment, the above method for determining the ALF coefficient is used to determine the upsampling filter coefficient.

[0100] According to one embodiment, there may be at least one WF for each upsampling phase. The phase of the sample to be interpolated allows for the determination of the upsampling filter to use (1230). The example shown in Figure 14A corresponds to upsampling in two directions, horizontally and vertically. The black dots are the reconstructed sample t(r) of the decoded picture (either the reference picture or the decoded picture for upsampling for display), and the white dots correspond to the interpolated sample f(r') (missing sample), where "r'" can be different from "r". In this example, there are three phases {0, 1, 2, 3}. Phase 0 has the same location as the reconstructed sample (r'=r). The WF corresponding to phase 0 may be omitted (it is assumed to be identical).

[0101] Equation (2) is modified as follows (1240).

[0102]

number

[0103] In (Equation 3), the expression for v is transformed as follows:

[0104]

number

[0105] In one variation, only missing points r(x,y) in the upscaled picture, i.e., points that are not in the same location in the downscaled picture, are interpolated. In another variation, all positions r(x,y) are interpolated, i.e., missing points and points that are in the same location in the downscaled picture are interpolated.

[0106] In one variant, some samples corresponding to certain subsets of phases are interpolated using only a WF filter, while other phases are interpolated using a standard separable 1D filter. For example, in Figure 14A, phases 0 and 1 are interpolated using WF in the first step, and the following phases 2 and 3 are interpolated using a horizontal 1D filter with filtered samples of phases 0 and 1. Conversely, phases 0 and 2 are interpolated with WF, and the following phases 1 and 3 are interpolated with a 1D vertical filter.

[0107] Figure 14A shows a 4x4 square filter shape, but it may have a different shape. Figures 14B to 14E show different shapes that can be used to interpolate a phase 3 sample, and the filter shape is indicated by a black sample representing the reconstructed sample used to interpolate the phase 3 sample.

[0108] Figures 14F and 14G show other examples of horizontal filter shapes that can be used to interpolate a sample in phase 2. Figure 14H shows another example of a vertical filter shape that can be used to interpolate a sample in phase 1. Figure 14I shows another example of a central filter shape that can be used to interpolate a sample in phase 3.

[0109] The shape may depend on the class and / or topology. Similar to ALF, the coefficients of some shapes / classes may be identical to those of other classes / shapes, but can be obtained by rotation, and the coefficients of one shape may be obtained by symmetry. For example, the coefficient of the shape in Figure 14B is 90 ° The shape after rotation may be the same as that of Figure 14C.

[0110] In one variant, a reference sample is classified (1220). A different upsampling WF is used for each class. In another variant, the classification can be the same as the classification used by the ALF.

[0111] Figure 15 shows an example of a method 1500 for determining the upsampling filter coefficient used on the encoder side, according to one embodiment.

[0112] The original picture is downscaled (1510) and encoded (1520). Reconstructed samples from the encoded picture are classified into classes (1530). A set of filter coefficients F0 is determined for a region R of the reconstructed picture, for example, CTU or groups of CTUs (1540). The set of filter coefficients F0 is given by F0 = {g 00 ,g 01 ,...,g 0M The set F0 is provided with upsampling filters for each class and phase, where M is the number of classes or phases, or the number of combinations of classes and phases, if there is one filter associated with each class and phase. The filters for set F0 are determined using (Equations 3 and 5) as described above.

[0113] The determined upsampling filter F0 is used with Equation 4 to obtain the upsampled region R of the reconstructed picture region R. up This is applied to obtain the sample f0(r') (1550).

[0114] Other upsampling filters Fi are applied in the same way, and the upsampled region R of the reconstructed picture is obtained. up The sample fi(r') is evaluated (1555), where Fi={g i0 ,g i1 ,...,g iM} and i = {1,...L}, where L is the number of possible filters per class and / or phase that have already been transmitted or are known by the decoder. Advantageously, the distortion can be directly derived from the coefficient values ​​and the original sample s(r').

[0115] The selection of filters to use for class / phase is, for example, using the rate-distortion Lagrangian cost, for each class / phase s, a new upsampling filter g 0s This involves coding the default value or the previously sent filter value g is This can be determined by finding the best trade-off between reusing i={1,...L} (1560). Distortion is the difference between the upsampled and reconstructed region and the corresponding region in the original picture (e.g., L1 norm or L2 norm).

[0116] The filter g determined for class / phase s 0s The rate distortion cost of the filter g is If the rate distortion cost is lower than any one of the following, filter g 0s The coefficients are coded in the bitstream (1570).

[0117] For each class / phase s, the index I (i=0...L) of the filter that provides the lowest rate distortion cost is coded in the bitstream of region R (1580).

[0118] In some embodiments, region R can be a region within a reconstructed picture, an entire picture, a group of several pictures, or a group of several regions within different pictures.

[0119] The method for determining the filters to be used in region R was described above for the case where there is one filter per class, class, and / or phase. A similar method can be applied when F0 and Fi each contain one single filter.

[0120] In one variant, the determination of the filter coefficients can be performed by machine learning using an iterative optimization algorithm (e.g., gradient descent). This may have the advantage of learning on a large number of samples / images without numerical constraints on Tc and v when R is large.

[0121] According to one embodiment, as shown in Figures 16 and 17, even if the coded picture corresponds to a downsampled picture, the reconstructed upsampled picture is stored in the DPB. According to this embodiment, the DPB contains only high-resolution reference pictures.

[0122] Figures 16 and 17 show a method 1600 for encoding video and a method 1700 for decoding video according to one embodiment, respectively. The original picture can be coded at low or high resolution.

[0123] The original high-resolution picture is downsampled by the encoder before coding (1610) (1660). The upsampling filter coefficients may be derived as described above (1640), and the reconstructed picture is upsampled before being stored in the DPB (1620) (1650). Then, normal RPR motion compensation is applied (the reference picture is high resolution and the current picture is low resolution) (1630).

[0124] In the decoding stage, the downscaled picture is decoded from the bitstream (1710), and if upsampling filter coefficients are present in the bitstream, the upsampling filter coefficients are decoded (1740). The low-resolution decoded picture is upsampled (1750) and stored in the DPB (1720). Then, normal RPR motion compensation is applied (the reference picture is high resolution and the current picture is low resolution) (1730). In one variant, the low-resolution decoded picture is stored in the DPB, and the upsampled decoded picture is used for display only.

[0125] If the original picture is coded at high resolution, downsampling (1660) and upsampling (1650, 1750) are bypassed.

[0126] Note that in one variant, the upsampling filter has predetermined default coefficients, and steps 1640 and 1740 are not present / not bypassed.

[0127] Post-filtering for image restoration In video standards (e.g., HEVC, VVC), reconstruction filters are applied to the reconstructed picture to reduce coding artifacts. For example, the Sample Adaptive Offset (SAO) filter is introduced in HEVC to reduce ringing and banding artifacts in the reconstructed picture, complementing the De-Blocking Filter (DBF), which particularly reduces artifacts at block boundaries. In VVC, an additional Adaptive Loop Filter (ALF) attempts to minimize the mean squared error between the original and reconstructed samples using Wiener-based adaptive filter coefficients. SAO and ALF employ a classification of the reconstructed samples to select which filter to apply.

[0128] ALF classification As discussed above, ALF is a specific post-filter for reconstructed image restoration. ALF classifies samples into K classes (for example, K=25 for lumar samples) or K regions (for example, K=8 for chroma samples), and K different filters are determined using samples from each class or region. In the case of classes, lumar sample classification is performed using directional and activity values ​​derived using local gradients.

[0129] In VVC, ALF coefficients can be coded in the bitstream so that they can dynamically adapt to the video content. These coefficients can be stored for reuse for further pictures. There are also several default coefficients, and the encoder indicates which set of coefficients to use for each CTU.

[0130] In VVC, a symmetric filter is used (as shown at the top of Figure 11), and some filter coefficients can be obtained from other filter coefficients by rotation (as shown at the bottom of Figure 11).

[0131] Motion compensation filtering and SIF In hybrid video coding, interpretation predicts the current block using motion compensation for a reference block extracted from a previously reconstructed reference picture. The difference in position between the current block and the reference block is the motion vector.

[0132] The motion vector may have sub-Pel accuracy (e.g., 1 / 16 in VVC), and the motion compensation process is performed as shown in Figure 6, with the corresponding sub-Pel position (θ) in the reference picture. x ,θ y Select an interpolation filter that has ). Traditionally, to reduce implementation complexity, motion-compensated interpolation filtering is performed using separable filters (one horizontal filter and one vertical filter).

[0133] To improve coding efficiency, for some sub-perfect positions, the encoder may select from several filters and signal them in the bitstream. For example, the VVC standard may select between two interpolation filters (normal or Gaussian filters) for 1 / 2 sub-perfect positions. Such tools are also known as switching interpolation filters (SIF tools). A Gaussian filter is a low-pass filter that smooths high frequencies compared to a normal filter.

[0134] According to ALF post-filtering, better efficiency in the filtering process is achieved when the samples (or groups of samples) to be filtered are pre-classified, and this classification is used to select a specific set of filter coefficients for each sample (or group of samples). On the encoder side, the classification can be used to determine the filter coefficients that minimize the mean squared error between the original sample "s(r)" and the filtered sample "t(r)" by using Wiener-based adaptive filtering techniques (e.g., as described in C. Tsai et al., "Adaptive Loop Filtering for Video Coding," IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL.7, NO.6, DECEMBER 2013).

[0135] However, classifying samples significantly increases the number of operations required for each sample.

[0136] In VVC, only ALF uses classification. The SIF tool signals which filter to use for motion compensation per CU, but the same filter is used to build all predictive samples in the predictive unit. In RPR, a single set of rescaling interpolation filters is selected per picture using the ratio between the reference block size and the current block size, and all samples are filtered using this single filter. The set of rescaling filters, for each phase, contains the coefficients of the filters to be used.

[0137] According to one aspect of the present principle, a method is provided for encoding / decoding video, wherein a sample classification of a reference picture is used to select at least one motion-compensated interpolation filter when predicting blocks of pictures in the video.

[0138] According to one embodiment, for each sample or group of samples from a reference picture that needs to be interpolated, the class to which the sample belongs is determined (from the classification performed on the reference picture). Then, an interpolation filter associated with this class is selected, and the samples are filtered using the coefficients of the selected filter.

[0139] According to another aspect of this principle, a method for encoding / decoding video is provided, wherein the sample classification of the reconstructed picture is shared among different encoding / decoding modules of an encoder / decoder. For example, a reference picture is classified, and this classification is then used to select at least one filter to be used during the encoding / decoding operation of a new picture using the reference picture, such as resampling filtering or motion-compensated interpolation filtering.

[0140] In another embodiment, the reconstructed picture is classified, and this classification is then used to select at least one filter to be used during encoding / decoding operations on the reconstructed picture, such as post-filtering and / or resampling for display, and / or during encoding / decoding operations on a new picture that uses the reconstructed picture as a reference picture, such as resampling filtering or motion-compensated interpolation filtering. For example, this can be done on a sample (or group of samples) basis, and the sample (or group of samples) classification allows for the selection of the filter to be used for this sample (or group of samples).

[0141] Conventionally, a filter includes several coefficients, each applied to adjacent samples of the currently filtered sample, and adjacent samples are determined according to the selected filter shape, an example of which is shown in Figure 11.

[0142] According to one embodiment, in order to share classifications among arbitrary encoding / decoding modules, the classification results are stored in a common space accessible by any one of the encoding / decoding modules, such as a decoded picture buffer (DPB) that stores the reference picture.

[0143] According to this principle, the ability of sample classification for filter selection is utilized in motion-compensated interpolation filters and resampling filters while keeping complexity relatively low. This is achieved by sharing the sample classification for several filtering purposes, such as restoration filters (e.g., ALF or bilateral filters), MC filtering, and resampling filters. In one embodiment, the classification can be stored in a DPB.

[0144] In an encoder, classifying reconstructed samples allows for the derivation of filters specialized for each sample class. This can be done, for example, by minimizing the mean squared error between the original sample and the reconstructed sample belonging to a particular class, using Wiener-based adaptive filter coefficients.

[0145] Next, on the decoder side, the selection of filters to be used is controlled by the classification process. For example, the classification process determines a class index for each sample, and one filter is associated with one class index.

[0146] In some variations, classification is performed on groups of samples rather than on individual samples. For example, a group of samples might be a 2x2 region.

[0147] Classification of interpolation filters Figure 18 shows a method 1800 for encoding or decoding video according to one embodiment. According to this embodiment, a set of interpolation filters is defined, including an interpolation filter for each class index. The interpolation filters can be determined in the same way as for the ALF filter, and the coefficients of the new interpolation filter can be sent to the decoder when they need to be adapted to the content.

[0148] A reference picture is input to the process. Samples of the reference picture are classified (1810). Then, in 1820, block motion compensation is performed to determine the prediction for the current block to encode or decode.

[0149] Motion vectors are obtained for blocks of video to be encoded or decoded. The motion vectors allow for the determination of a portion or block of a reference picture in order to predict the block.

[0150] When the motion vector points to a subsample location, as shown in Figure 6, the samples of the motion-compensated portion of the reference picture need to be interpolated to determine the block samples for prediction. According to this principle, the interpolation filter used for each subsample is determined based on classification (1830).

[0151] Therefore, the block prediction is determined as an interpolated sample of the reference picture (1840).

[0152] According to one embodiment, in order to determine the interpolation filter (1830), for each sample of the motion-compensated portion of the reference picture, a class index is determined from, for example, one or more class indices associated with one or more neighboring samples at the sample location in the reference picture. Then, an interpolation filter is selected for each subsample and interpolated using the class index determined for the subsample. Next, a block prediction is generated by interpolating each subsample of the motion-compensated portion of the reference picture with the interpolation filter selected for that subsample (1840). Finally, the block is encoded or decoded using the prediction (depending on whether the method is performed in an encoder or a decoder) (1850). During encoding, the residual between the original block and its prediction is determined and coded. During decoding, the residual is decoded and added to the prediction to reconstruct the block, and the block prediction is generated in the same process as in the encoder case.

[0153] According to another aspect of this principle, the same sample classification is shared between the encoding / decoding modules of an encoder or decoder. The set of filters is defined for each type of encoding or decoding operation that uses filters, such as motion-compensated interpolation, resampling, and ALF.

[0154] Same classification for interpolation and resampling filters Figure 19 shows an example of a method 1800 for encoding or decoding video according to another embodiment. A common classification of a reference picture (19810) can be performed and used to take advantage of the ability of sample classification for filter selection for both motion compensation (MC) interpolation (1940) filters and resampling (1930) filters.

[0155] Advantageously, classification is performed on the entire reconstructed picture, and the classification of each sample is stored so that it can be used by motion-compensated interpolation and resampling filter processes (1920). If resampling is performed implicitly within the MC process (1950), the classification is input directly into the MC.

[0156] According to one embodiment, the classification is stored in the DPB along with the reference picture so that it can be reused by other processes.

[0157] The same classification applies to interpolation, resampling filters, and post-filtering. Figure 20 shows an example of a method 2000 for encoding or decoding video according to another embodiment. In this variant, classification (2030) is performed on the reconstructed picture before applying a restoration filter (also known as a post-filter (PF)) (e.g., ALF) (2050). The classification can then be used by the encoder to derive the filter coefficients of the post-filter (e.g., ALF) (2040). The classification is used to select the filter to be used by post-filtering (2050). Advantageously, this classification is also used by resampling filtering or motion-compensated interpolation filtering so that only a single classification step (2030) is performed. Note that in this variant, the other processes (e.g., resampling filtering or motion-compensated interpolation filtering) use the classification performed before applying the restoration filter (post-filtering), while the other processes use the reconstructed picture samples (after post-filtering has been applied).

[0158] According to one embodiment, the classification may be stored in DPB(2020) so that it can be reused by other processes. In one variant, storage in DPB is performed when the picture is used only as a reference (2060).

[0159] Out-loop resampling In the case of RPR, the resampling process of the decoded picture is optional (Figure 5:540). Figure 21 shows an example of a method 2100 for decoding video according to another embodiment. The picture is decoded (2110), and samples of the decoded picture are classified (2130). A post-filter is applied based on the classification (2150), and the classification is finally stored in a DPB and made available to other processes that use the decoded picture as a reference picture.

[0160] The selection of the resampling filter to use (e.g., upsampling) may be controlled by the classification process (2130). The classification process determines a class index for each sample (or group of samples), and one filter is associated with one class index. The filter index allows for the selection of a resampling filter (2160).

[0161] It should be understood that the encoding and decoding methods described above can be implemented in the encoder 200 and decoder 300, respectively, as described in relation to Figures 2 and 3, for encoding video in a bitstream and for decoding video from a bitstream.

[0162] In one embodiment shown in Figure 22, in a transmission context between two remote devices A and B via a communication network NET, device A includes a processor associated with memory RAM and ROM configured to perform a method of encoding video according to any one of the embodiments described with respect to Figures 1 to 21, and device B includes a processor associated with memory RAM and ROM configured to perform a method of decoding video according to any one of the embodiments described with respect to Figures 1 to 21.

[0163] According to one embodiment, the network is a broadcast network adapted to broadcast / transmit encoded data representing video from device A to decoding devices including device B.

[0164] The signal intended to be transmitted by device A carries at least one bitstream containing coded data representing video. The bitstream may be generated from any embodiment of the present principle.

[0165] Figure 23 shows an example of the syntax of such a signal transmitted over a packet-based transmission protocol. Each transmitted packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include coded video data encoded according to any one of the embodiments described above. In some embodiments, the signal includes the filter (upsampling, interpolation) coefficients determined above.

[0166] Various implementations include decoding. As used in this application, “decoding” can encompass all or part of the processes performed on a received encoded sequence to produce a final output suitable for a display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include, or alternatively, processes performed by various implementations of decoders for decoding upsampling filter coefficients that upsample a decoded picture, as described in this application.

[0167] As further examples, in one embodiment, “decoding” refers only to entropy decoding; in another embodiment, “decoding” refers only to differential decoding; and in yet another embodiment, “decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase “encoding process” is intended to refer specifically to a working subset or to refer to a broader encoding process as a whole will become clear from the context of the specific explanation and will be well understood by those skilled in the art.

[0168] Various implementations involve encoding. Similar to the above consideration of "decoding," "encoding" as used in this application can encompass all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as splitting, differential coding, transformation, quantization, and entropy coding. In various embodiments, such processes also, or alternatively, include processes performed by various implementations of the encoder, such as determining the upsampling filter coefficients for upsampling the decoded picture, as described in this application.

[0169] As further examples, in one embodiment, “encoding” refers only to entropy coding; in another embodiment, “encoding” refers only to differential coding; and in yet another embodiment, “encoding” refers to a combination of differential coding and entropy coding. Whether the phrase “encoding process” is intended to refer specifically to a working subset or to refer to a broader encoding process as a whole will become clear from the context of the specific explanation and will be well understood by those skilled in the art.

[0170] Please note that the syntax elements used in this specification are descriptive terms; therefore, they do not preclude the use of other syntax element names.

[0171] This disclosure has described various types of information, such as syntax, that can be transmitted or stored. This information can be packaged or arranged in various formats, including formats common in video standards, such as including the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other formats are also available, including formats common in system-level or application-level standards, such as including the information in one or more of the following:

[0172] a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for the purpose of session announcement and session invitation, for example, as described in an RFC and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. For example, a DASH MPD (Media Presentation Description) descriptor, such as one used in DASH and transmitted over HTTP, is associated with a representation or set of representations to provide additional characteristics to the content representation. c. For example, an RTP header extension used during RTP streaming. d. For example, an ISO-based media file format that uses boxes, which are object-oriented construction blocks defined by a unique type identifier and length, and which are used in OMAF and are also known as "atoms" in some specifications. e. An HLS (HTTP Live Streaming) manifest sent via HTTP. The manifest can be associated with a version or set of versions of content to provide, for example, version or set of versions characteristics.

[0173] If a diagram is presented as a flowchart, it should be understood that the diagram also provides a block diagram of the corresponding device. Similarly, if a diagram is presented as a block diagram, it should be understood that the diagram also provides a flowchart of the corresponding method / process.

[0174] Various embodiments refer to rate-distortion optimization. Specifically, during the coding process, the balance or trade-off between rate and distortion is usually considered to impose computational complexity constraints. Rate-distortion optimization is typically formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving rate-distortion optimization problems. For example, these approaches are obtained based on extensive testing of all coding options, including all considered mode or coding parameter values, and involve a complete evaluation of their coding costs, as well as the associated distortions of the reconstructed signals after coding and decoding. To reduce coding complexity, faster methods can also be used, in particular, along with the calculation of approximate distortions based on prediction or prediction residual signals rather than the reconstructed signals. A mixture of these two methods can also be used, such as using approximate distortions for only some of the possible coding options and using full distortions for the others. Other methods evaluate only a subset of the possible coding options. More generally, many approaches employ one of various techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding costs and associated distortions.

[0175] The implementations and embodiments described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if considered only in the context of a single implementation (for example, considered only as a method), the implementations of the considered features may also be implemented in other forms (for example, apparatus or programs). For example, an apparatus may be implemented in appropriate hardware, software, and firmware. This method may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, mobile phones, and personal digital assistants ("PDAs"), which facilitate the communication of information between end users.

[0176] References to “one embodiment” or “a certain embodiment,” or “one implementation” or “a certain implementation,” or to other variations thereof, mean that the specific features, structures, characteristics, etc. described in relation to that embodiment are included in at least one embodiment. Therefore, when the phrases “in one embodiment” or “in a certain embodiment,” or “in one implementation” or “in a certain implementation,” or other variations appear in various places throughout this application, they do not necessarily all refer to the same embodiment.

[0177] In addition, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0178] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0179] In addition, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory). Furthermore, "receiving" generally involves in some way operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0180] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the following " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first listed option (A), only the second listed option (B), or both options (A and B). In further embodiments, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such expressions are intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to the number of listed items, as will be apparent to those skilled in the art in this and related fields.

[0181] Furthermore, as used herein, the term “signaling” specifically means indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals one particular upsampling filter coefficient among several. Thus, in certain embodiments, the same parameter is used on both the encoder and decoder sides. Therefore, for example, an encoder can transmit a particular parameter to a decoder so that the decoder can use the same particular parameter (explicit signaling). In contrast, if the decoder already has other parameters along with that particular parameter, it can use non-transmitting signaling (implicit signaling) so that the decoder simply knows and can select that particular parameter. Bit saving is achieved in various embodiments by avoiding the transmission of any actual function. It will be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc., are used in various embodiments to signal information to a corresponding decoder. The above relates to the verb form of the word “signal,” which may also be used as a noun herein.

[0182] As will be obvious to those skilled in the art, the implementation can bring about a variety of signals formatted to carry information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the implementations described. For example, a signal can be formatted to carry a bitstream of the embodiment described. For example, such a signal can be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog information or digital information. As is known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored in a processor-readable medium.

[0183] Several embodiments are described. Features of these embodiments may be provided individually or in any combination across various claims and types. Furthermore, embodiments may include, individually or in any combination, one or more of the following features, devices, or aspects across various claims and types. Encoding / decoding video in any of the embodiments described, which allows the original picture to be encoded at a higher or lower resolution. • Reconstructing a picture from a downscaled, decoded picture using any of the embodiments described. A bitstream or signal containing one or more of the described syntax elements, or variations thereof. A bitstream or signal containing a syntax that carries information generated according to any of the embodiments described. • To create and / or transmit and / or receive and / or decode a bitstream or signal that contains one or more of the described syntax elements or variations thereof. Creating and / or transmitting and / or receiving and / or decrypting by any of the embodiments described. A method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that performs picture reconstruction by upsampling according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that performs picture reconstruction by upsampling according to any of the embodiments described and displays the resulting image (for example, using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that selects a channel (for example, using a tuner) to receive a signal containing an encoded image and performs a reconstruction of the picture by upsampling according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that receives a wireless signal containing an encoded image (for example, using an antenna) and performs a reconstruction of the picture by upsampling according to any of the embodiments described. Encoding / decoding video, according to any of the embodiments described, such that the same classification of pictures is shared between the encoding or decoding processes. Encoding / decoding video, according to any of the embodiments described, in which classification is used to select an interpolation filter when subsamples are interpolated. A bitstream or signal containing one or more of the syntax elements described, or variations thereof. A bitstream or signal containing a syntax that carries information generated according to any of the embodiments described. • To create and / or transmit and / or receive and / or decode a bitstream or signal that contains one or more of the described syntax elements or variations thereof. Creating and / or transmitting and / or receiving and / or decrypting by any of the embodiments described. A method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that performs the reconstruction of a picture according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that performs a picture reconstruction according to any of the embodiments described and displays the resulting image (for example, using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that selects a channel (for example, using a tuner) to receive a signal containing an encoded image and performs picture reconstruction according to any of the embodiments described. A TV, set-top box, mobile phone, tablet, or other electronic device that receives a wireless signal containing an encoded image (for example, using an antenna) and performs a reconstruction of the picture according to any of the embodiments described.

Claims

1. It is a method, Decrypting the first picture, Resampling at least a portion of the first picture in order to reconstruct at least a portion of the second picture, wherein resampling at least a portion of the first picture is Classifying a sample of at least a portion of the first picture, Obtain the subpixel position of at least a portion of the first picture for at least one sample of at least a portion of the second picture, In classifying at least a portion of the first picture, the class index for the subpixel position is determined from at least one class index assigned to at least one adjacent pixel in at least a portion of the first picture, wherein the at least one adjacent pixel is a pixel adjacent to the subpixel position. Using the class index associated with the subpixel position, a resampling filter is selected for at least one sample of at least a portion of the second picture. This includes, Methods that include...

2. The method according to claim 1, further comprising transmitting the reconfigured at least portion of the second picture to a display.

3. The method according to claim 1 or 2, further comprising storing at least a reconfigured portion of the second picture in a decoded picture buffer that stores a reference picture.

4. Further including encoding a third picture, The above encoding is, Using at least a reconfigured portion of the second picture, a prediction is made for at least one block of the third picture. Using the prediction, encode at least one block of the third picture, The method according to claim 3, including the method described in claim 3.

5. Further including decrypting a third picture, The aforementioned decoding is Using at least a reconfigured portion of the second picture, a prediction is made for at least one block of the third picture. Using the prediction, decode at least one block of the third picture, The method according to claim 3, including the method described in claim 3.

6. The method according to any one of claims 1 to 5, comprising decoding the coefficients of the resampling filter from the bitstream.

7. The method according to any one of claims 1 to 6, wherein the resampling filter is a non-separable filter.

8. The method according to claim 1, wherein different resampling filters are associated with each class.

9. The method according to any one of claims 1 to 8, wherein the resampling filter is determined based on a rate distortion cost determined between at least a portion of the second picture and at least a reconfigured portion of the second picture obtained from the first picture.

10. A device comprising one or more processors, The one or more processors described above are: Decrypting the first picture, Resampling at least a portion of the first picture in order to reconstruct at least a portion of the second picture, wherein resampling at least a portion of the first picture is Classifying a sample of at least a portion of the first picture, Obtain the subpixel position of at least a portion of the first picture for at least one sample of at least a portion of the second picture, In classifying at least a portion of the first picture, the class index for the subpixel position is determined from at least one class index assigned to at least one adjacent pixel in at least a portion of the first picture, wherein the at least one adjacent pixel is a pixel adjacent to the subpixel position. Using the class index associated with the subpixel position, a resampling filter is selected for at least one sample of at least a portion of the second picture. This includes, A device configured to perform the following actions.

11. The apparatus according to claim 10, wherein one or more processors are further configured to transmit at least a reconfigured portion of the first picture to a display.

12. The apparatus according to claim 10, wherein one or more processors are further configured to store at least a reconfigured portion of the second picture in a decoded picture buffer that stores a reference picture.

13. The apparatus according to claim 10, wherein the one or more processors are further configured to encode the third picture by using at least a reconfigured portion of the first picture to determine a prediction for at least one block of the third picture, and by encoding the at least one block of the third picture using the prediction.

14. The apparatus according to claim 10, wherein one or more processors are further configured to decode the third picture by using at least a reconfigured portion of the first picture to determine a prediction for at least one block of the third picture, and by using the prediction to decode the at least one block of the third picture.

15. The apparatus according to any one of claims 10 to 14, wherein the one or more processors are further configured to decode the coefficients of the resampling filter from the bitstream.

16. The apparatus according to any one of claims 10 to 15, wherein the resampling filter is a non-separable filter.

17. The apparatus according to claim 10, wherein different resampling filters are associated with each class.

18. The apparatus according to any one of claims 10 to 17, wherein the resampling filter is determined based on a rate distortion cost determined between at least a portion of the second picture and at least a reconfigured portion of the second picture obtained from the first picture.

19. A computer-readable storage medium storing instructions for one or more processors to perform the method according to any one of claims 1 to 9.

20. A computer program including instructions, wherein, when the computer program is executed by one or more processors, the instructions cause one or more processors to perform the method according to any one of claims 1 to 9.

21. It is a device, The apparatus according to any one of claims 10 to 18, A device comprising: (i) an antenna configured to receive a signal containing data representing a video including the first picture; (ii) a band limiter configured to restrict the signal to a frequency band containing the data representing the video; or (iii) a display configured to display at least a portion of at least one of the first images.

22. The device according to claim 21, including a TV, mobile phone, tablet, or set-top box.

23. The apparatus according to claim 10, wherein the class index determined for the subpixel position is the class index of the adjacent pixel closest to the subpixel position.

24. The apparatus according to claim 10, wherein the classification is performed using directional values ​​and activity values ​​derived using local gradients in at least a portion of the first picture.

25. The apparatus according to claim 10, wherein the shape of the resampling filter depends on the subpixel position.

Citation Information

Patent Citations

  • A Switched Filter Upsampling Mechanism for Scalable Video Coding

    JP2009522971A

  • Moving image decoding method, moving image decoding apparatus, and moving image decoding program

    JP2010288182A

  • Interlayer prediction method and apparatus utilizing the same

    JP2015512216A