Method and apparatus for encoding and decoding images or videos
The use of position-dependent pixel combinations in intra-prediction methods addresses the challenge of high compression efficiency in video compression, resulting in improved video quality and reduced redundancy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-03-26
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video compression technologies face challenges in achieving high compression efficiency due to limitations in intra-prediction methods, particularly in handling spatial and temporal redundancy in video content.
The method and apparatus enhance intra-prediction by using position-dependent pixel combinations, including weighted combinations of secondary reference samples, to improve prediction accuracy and reduce blocking artifacts.
This approach leads to improved video compression efficiency by reducing redundancy and enhancing prediction accuracy, thereby improving the quality of reconstructed video.
Smart Images

Figure 2026511392000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to European Patent Application No. 23305462.6, filed on 31 March 2023, and European Patent Application No. 23305548.2, filed on 12 April 2023, both of which are incorporated herein by reference in their entirety.
[0002] (Field of Invention) This embodiment relates, in general, to video compression. This embodiment relates to a method and apparatus for encoding or decoding images or videos. More specifically, this embodiment relates to improving intra-prediction using position-dependent pixel combinations. [Background technology]
[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation that leverage the spatial and temporal redundancy of video content. Generally, intra or inter-prediction is used to leverage intra or inter-picture correlation, and then the difference between the original block and the predicted block, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy-coded. In inter-prediction, the motion vectors used in motion compensation are often predicted from motion vector predictors. To reconstruct the video, the compressed data is decoded by the reverse process corresponding to entropy coding, quantization, transformation, and prediction. [Overview of the project]
[0004] According to one embodiment, a method for encoding video is provided. The method includes: obtaining a predictor block for a video block based on an angular intra-prediction mode using at least one first primary reference sample for at least one pixel of the predictor block; modifying the predictor block for at least one pixel of the predictor block using a weighted combination of values determined from at least two secondary reference samples; and encoding the video block based at least on the modified predictor block.
[0005] In another embodiment, an apparatus for encoding video is provided. The apparatus comprises one or more processors, which are operable to acquire predictor blocks for video blocks based on an angular intra-prediction mode using at least one first primary reference sample for at least one pixel of the predictor block, modify the predictor block using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the predictor block, and encode the video block based at least on the modified predictor block.
[0006] In another embodiment, a method for decoding a video is provided. This method includes: obtaining a predictor block for a video block based on an angular intra-prediction mode using at least one first primary reference sample for at least one pixel of the predictor block; modifying the predictor block for at least one pixel of the predictor block using a weighted combination of values determined from at least two secondary reference samples; and decoding the video block based at least on the modified predictor block.
[0007] In another embodiment, a device for decoding video is provided. The device comprises one or more processors, which are operable to acquire predictor blocks for video blocks based on an angular intra-prediction mode using at least one first primary reference sample for at least one pixel of the predictor block, modify the predictor block using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the predictor block, and decode the video block based at least on the modified predictor block.
[0008] Further embodiments that may be used individually or in combination are described herein.
[0009] One or more embodiments also provide a computer program that, when executed by one or more processors, causes one or more processors to perform a method for encoding / decoding video according to any of the embodiments described herein. One or more embodiments also provide a non-temporary computer-readable medium and / or computer-readable storage medium storing instructions for encoding / decoding video according to the method described herein.
[0010] In another embodiment, a method is provided for encoding or decoding video, which includes: obtaining a predicted value for at least one pixel of a video block based on an intra-prediction mode using at least one first primary reference sample; enabling or disabling a modification of the predicted value of at least one pixel by at least one secondary reference sample based on the distance of the at least one secondary reference sample to the origin of the video block; and encoding or decoding the video block.
[0011] In another embodiment, a method for encoding or decoding video is provided, which includes: obtaining a predicted value for at least one pixel of a video block based on an intra-prediction mode using at least one first primary reference sample; enabling or disabling a modification of the predicted value of at least one pixel by at least one secondary reference sample based on the availability of at least one secondary reference sample; and encoding or decoding the video block.
[0012] A device is also provided that includes one or more processors capable of operating to carry out the above method.
[0013] One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the methods described above. [Brief explanation of the drawing]
[0014] [Figure 1] A block diagram of a system in which an embodiment of this model may be implemented is shown. [Figure 2] A block diagram of an embodiment of a video encoder in which an aspect of this embodiment may be implemented is shown. [Figure 3] A block diagram of an embodiment of a video decoder in which an aspect of this embodiment may be implemented is shown. [Figure 4A] An example of the angle intra-prediction mode in VVC is shown. [Figure 4B] An example of the angle intra-prediction mode in VVC is shown. [Figure 4C] This shows the relationship between the expansion of the set of decoded reference samples surrounding the W×H block to be predicted and the range of acceptable intra-prediction angles. [Figure 4D] This shows the relationship between the expansion of the set of decoded reference samples surrounding the W×H block to be predicted and the range of acceptable intra-prediction angles. [Figure 4E] In VVC and ECM, examples of angle modes replaced by wide-angle modes are shown for non-square blocks where the width is strictly greater than the height. [Figure 5] An example of a PDPC process in VVC or ECM is shown, where one secondary reference sample is used to correct the first predicted value in the positive prediction direction. [Figure 6] An example of a gradient PDPC process in VVC or EMC is shown, where the secondary reference sample lies on the left (upper) reference array in the same row (column) as the target pixel, and the gradient determined at the secondary reference sample is added to the first predicted value using a weight that is a decreasing function of the distance from the left (upper) reference array. [Figure 7] An example flowchart is shown for determining whether to apply a standard PDPC or a gradient PDPC to a target block. [Figure 8] An example flowchart of a method for encoding video according to one embodiment is shown. [Figure 9] An example flowchart of a method for decoding video according to one embodiment is shown. [Figure 10] An example of an integrated PDPC process according to one embodiment is shown, in which two secondary reference samples are used: a secondary reference sample determined to be for "normal" PDPC and a secondary reference sample determined to be for gradient PDPC. [Figure 11] An example of an integrated PDPC process according to another embodiment is shown when at least one of the secondary reference samples is unavailable. [Figure 12] A block diagram of a system in which an aspect of this embodiment may be implemented is shown according to another embodiment. [Figure 13] This example illustrates two remote devices communicating via a communication network based on this principle. [Figure 14] The syntax of a signal is shown using an example of this principle. [Figure 15]This shows a template for the block to be encoded or decoded, and an example of a decoded reference sample for that template. [Figure 16] An example of how to disable PDPC according to an embodiment is shown. [Figure 17] An example of disabling PDPC according to another embodiment is shown. [Figure 18] An example of adapting a set of template reference samples for a block to be predicted using template-based intra-prediction mode derivation is shown according to one embodiment. [Figure 19] Another example of adapting a set of template reference samples for blocks to be predicted using template-based intra-prediction mode derivation, according to one embodiment, is shown. [Figure 20] This example shows an adaptation of a set of template reference samples for a block to be predicted using template-based intra-predictive mode derivation, according to one embodiment, when the left-hand portion of the template is unavailable. [Figure 21] This example shows an adaptation of a set of template reference samples for a block to be predicted using template-based intra-predictive mode derivation, according to one embodiment, when the top of the template is unavailable. [Figure 22] Another embodiment shows an example of adapting a set of template reference samples for blocks to be predicted using template-based intra-predictive mode derivation. [Figure 23] Another embodiment shows an example of adapting a set of template reference samples for blocks to be predicted using template-based intra-predictive mode derivation. [Figure 24] An example of adapting a set of template reference samples for a block to be predicted using template-based intra-prediction mode derivation is shown according to one embodiment. [Figure 25]An example of adapting a set of template reference samples for a block to be predicted using template-based intra-prediction mode derivation is shown according to one embodiment. [Figure 26] An example of adapting a set of template reference samples for a block to be predicted using template-based intra-prediction mode derivation is shown according to one embodiment. [Figure 27] An example of a method for encoding video blocks according to one embodiment is shown. [Figure 28] An example of a method for decoding a video block according to one embodiment is shown. [Figure 29] An example of a method for template-based intra-mode prediction derivation for encoding or decoding video blocks, according to one embodiment, is presented. [Figure 30] An example of a method for template-based intra-mode prediction derivation for encoding or decoding video blocks, according to another embodiment, is shown. [Modes for carrying out the invention]
[0015] This application describes various embodiments, including tools, features, embodiments, models, and approaches. Many of these embodiments are described in detail, often in a manner that sounds restrictive, at least to illustrate individual features. However, this is for clarity and not to limit the uses or scope of these embodiments. In fact, all of the different embodiments can be combined and interchangeable to provide further embodiments. Furthermore, these embodiments can also be combined with or interchangeable with embodiments described in previous applications.
[0016] The embodiments described and intended in this application can be realized in many different forms. Figures 1, 2, and 3 below provide some embodiments, but other embodiments are intended, and the considerations in Figures 1, 2, and 3 are not intended to limit the scope of implementation forms. At least one of these embodiments generally relates to video encoding and decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream. These and other embodiments can be realized as a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods, apparatus, or described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.
[0017] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” and “frame” may be used interchangeably.
[0018] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a particular order of steps or actions is required for the normal operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, actions, etc., such as “first decryption” and “second decryption.” The use of such terms does not imply a modified order of actions unless specifically required. Thus, in this example, the first decryption does not need to be performed before the second decryption, and may occur, for example, before, during, or over a period of overlap with the second decryption.
[0019] These embodiments are not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations, whether prior to exist or to be developed in the future, as well as to any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise specifically indicated or technically excluded, the embodiments described herein can be used individually or in combination.
[0020] Figure 1 shows a block diagram of an example of a system in which various embodiments and forms can be implemented. System 100 may be embodied as a device comprising various components described below and configured to perform one or more of the embodiments described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 100 may be embodied individually or in combination as a single integrated circuit, a plurality of ICs, and / or individual components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 100 are distributed across a plurality of ICs and / or individual components. In various embodiments, System 100 is communicably coupled to other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, System 100 is configured to implement one or more of the embodiments described in this application.
[0021] System 100 includes at least one processor 110, which is configured to execute instructions loaded therein to realize, for example, various embodiments described in this application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140 which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 may, in non-limiting examples, include an internal storage device, a removable storage device, and / or a network-accessible storage device.
[0022] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded or decoded video, the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in a device that performs encoding and / or decoding functions. As is known, the device may include one or both of the encoding module and the decoding module. Furthermore, the encoder / decoder module 130 may be implemented as a separate element of System 100 or may be incorporated into the processor 110 as a combination of hardware and software, as is known to those skilled in the art.
[0023] To perform the various embodiments described in this application, program code to be loaded into the processor 110 or encoder / decoder 130 may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and arithmetic logic.
[0024] In some embodiments, the internal memory of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory of the processing device (for example, the processing device can be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be memory 120 and / or storage device 140, and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG stands for Moving Picture Experts Group, also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a standard developed by JVET, i.e., Joint Video Experts Team).
[0025] Inputs to the elements of system 100 can be provided through various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) a radio frequency (RF) section for receiving RF signals wirelessly transmitted by, for example, a broadcasting station; (ii) a component (COMP) input terminal (or a set of COMP input terminals); (iii) a Universal Serial Bus (USB) input terminal; and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Another example not shown in Figure 1 is composite video.
[0026] In various embodiments, the input device of block 105 has associated input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a certain frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band so as to select a signal frequency band that may (for example) be called a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to the baseband) or to the baseband. In one embodiment of a set-top box, the RF unit and its associated input processing elements perform frequency selection by receiving, filtering, down-converting, and filtering again to a desired frequency band of RF signals transmitted via a wired (e.g., cable) medium. In various embodiments, the order of these (and other) elements is rearranged, some of these elements are removed, and / or other elements performing similar or different functions are added. Adding elements may include inserting elements between existing elements, for example, an amplifier and an analog-to-digital converter. In various embodiments, the RF unit includes an antenna.
[0027] Furthermore, the USB and / or HDMI terminals may include their respective interface processors for connecting the system 100 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be performed, for example, in a separate input processing IC or within the processor 110, as needed. Similarly, aspects of USB or HDMI interface processing may be implemented in a separate interface IC or within the processor 110, as needed. For example, demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including a processor 110 and an encoder / decoder 130, which may, as needed, process the data stream for display on an output device.
[0028] Various elements of system 100 may be provided within an integrated housing, where the various elements are interconnected and data can be transmitted between them using an internal bus known in the art, such as a suitable connection arrangement 115, including an I2C bus, wiring, and a printed circuit board.
[0029] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 190. The communication interface 150 may also include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.
[0030] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. In these embodiments, the communication channel 190 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming of the application and other over-the-top communication. In other embodiments, the streamed data is provided to system 100 using a set-top box that distributes data via an HDMI connection of input block 105. Still other embodiments provide the streamed data to system 100 using an RF connection of input block 105. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0031] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various embodiments, the display 165 includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 may be for a television, tablet, laptop, mobile phone, or other device. The display 165 may also be integrated with other components (e.g., a smartphone) or be separate (e.g., an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 185 include one or more of a standalone digital video disc (or digital multi-purpose disc) (both terms DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of System 100. For example, a disc player performs the function of playing back the output of System 100.
[0032] In various embodiments, control signals are communicated between the system 100 and the display 165, speaker 175, or other peripheral devices 185 using signaling such as AV Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be communicably coupled to the system 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to the system 100 using a communication channel 190 via a communication interface 150. The display 165 and speaker 175 may be integrated into a single unit with other components of the system 100 in an electronic device such as a television. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0033] The display 165 and speaker 175 may, alternatively, be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, for example, including an HDMI port, a USB port, or a COMP output.
[0034] The embodiments can be implemented by computer software implemented by the processor 110, by hardware, or by a combination of hardware and software. In a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 120 can be of any type appropriate for the technical environment and, in a non-limiting example, can be implemented using any appropriate data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 110 can be of any type appropriate for the technical environment and, in a non-limiting example, can include one or more of microprocessors, general-purpose computers, dedicated computers, and processors based on multi-core architectures.
[0035] Figure 2 shows the video encoder 200. While variations of this encoder 200 are considered, the encoder 200 is described below without explaining all expected variations for clarity.
[0036] In some embodiments, Figure 2 also shows encoders that are improvements over the HEVC or VVC standard, or encoders that employ HEVC or VVC-like technologies, such as the ECM encoder currently under development by JVET (Joint Video Exploration Team).
[0037] Before encoding, the video sequence may undergo pre-encoding (201), for example, applying a color conversion to the input color picture (e.g., from RGB4:4:4 to YCbCr4:2:0), or remapping the input picture components to obtain a more compression-resilient signal distribution (e.g., using histogram equalization of color components), or resizing (e.g., downscaling) the picture. Metadata may be associated with the pre-processing and attached to the bitstream.
[0038] In encoder 200, the picture is encoded by encoder elements as described below. The encoded picture is divided (202) and processed in units of, for example, coding units (CUs). Different expressions may be used in this disclosure to refer to such units or blocks resulting from the division of the picture. Such terms may be coding units or CUs, coding blocks or CBs, luminance CBs, or blocks. A coding tree unit (CTU) may refer to a group of blocks or a group of units. In some embodiments, a CTU may be considered a block or a unit in itself.
[0039] Each unit is encoded using either intra-mode or inter-mode, for example. If the unit is encoded in intra-mode, intra-prediction is performed (260). In inter-mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides whether to use intra-mode or inter-mode to encode the unit (205), and indicates the intra / inter decision, for example, by a prediction mode flag. The encoder may also blend the intra-prediction results and the inter-prediction results, or blend the results from different intra / inter-prediction methods (263). The prediction residual is calculated, for example, by subtracting the predicted blocks from the original image blocks (210).
[0040] The motion enhancement module (272) uses already available reference pictures to enhance a block's motion field without referencing the original block. The motion field of a region can be thought of as the set of motion vectors for all pixels in that region. If the motion vectors are subblock-based, the motion field can also be represented as the set of all subblock motion vectors in the region (all pixels in a subblock have the same motion vector, and the motion vectors may differ for each subblock). If a single motion vector is used for a region, the region's motion field can also be represented by a single motion vector (the same motion vector for all pixels in the region).
[0041] Next, the predicted residuals are transformed (225) and quantized (230). In addition to the quantized transformation coefficients, the motion vector and other syntactic elements are entropy coded to output a bitstream (245). The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residuals are coded directly without applying any transformation or quantization process.
[0042] The encoder decodes the encoded blocks and provides a reference for further prediction. The quantized transformation coefficients are inversely quantized (240), inversely transformed (250), and the prediction residuals are decoded. The decoded prediction residuals and the prediction blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed picture to reduce encoding artifacts, for example, by performing deblocking / SAO (Sample Adaptive Offset) filtering. The filtered image is stored in a reference picture buffer (280).
[0043] Figure 3 shows a block diagram of the video decoder 300. In the decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 generally performs a decoding path that is the reverse of the encoding path described in Figure 2. The encoder 200 also generally performs video decoding as part of encoding the video data.
[0044] Specifically, the input to the decoder includes a video bitstream, which may be generated by a video encoder 200. First, the bitstream is entropy-decoded (330) to obtain transformation coefficients, motion vectors, and other coded information. Picture partitioning information indicates how the picture is partitioned. Thus, the decoder may partition the picture according to the decoded picture partitioning information (335). The transformation coefficients are inversely quantized (340), inversely transformed (350), and the predicted residuals are decoded. The decoded predicted residuals and predicted blocks are combined (355) to reconstruct the image blocks.
[0045] The prediction block can be obtained from intra-prediction (360) or motion-compensated prediction (i.e., inter-prediction) (375) (370). The decoder may blend the intra-prediction result and the inter-prediction result, or blend the results from multiple intra / inter-prediction methods (373). Before motion compensation, the motion field may be improved by using an already available reference picture (372). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0046] The decoded picture may further undergo post-decoding (385), such as reverse color transformation (e.g., conversion from YCbCr4:2:0 to RGB4:4:4), or reverse remapping, which is the reverse of the remapping performed in the pre-encoding process (201), or resizing (e.g., upscaling) the reconstructed picture. The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0047] The embodiments described herein relate to intra-prediction. Some embodiments relate to location-dependent prediction combinations used in intra-prediction. Other embodiments relate to template-based intra-mode derivation (TIMD).
[0048] It should be understood that any one of the embodiments described herein relating to location-dependent predictive combinations can be applied to any one of the embodiments described herein relating to template-based intra-mode derivation (TIMD), and vice versa.
[0049] Any one of the embodiments described herein can be implemented, for example, in an intra-prediction module 260 of a video encoder 200, or in an intra-prediction module 360 of a video decoder 300.
[0050] To capture any edge direction presented in natural video, the number of directional intra-prediction modes in VVC is expanded from 33 to 65, as used in HEVC. New directional modes not present in HEVC are illustrated as dotted arrows in Figure 4A. These higher-density directional intra-prediction modes apply to all block sizes, as well as to both lumane and chromane intra-predictions. From HEVC to VVC, the planar and DC modes remain unchanged, with the following minor modifications. In HEVC, every intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, division is not required to generate intra-predictors using DC. In VVC, blocks can have a rectangular shape, which in the general case requires the use of division for each block. To avoid division for DC prediction, only the longer sides are used to calculate the average for non-square blocks.
[0051] In ECM, the core structure of the 67 intra-prediction modes is inherited from that of VVC. This core structure is improved in ECM as follows: the 4-tap interpolation for directional intra-prediction modes from VVC becomes 6-tap interpolation, and position-dependent intra-prediction combinations (PDPCs) are supplemented using gradient PDPCs.
[0052] In VVC and ECM, for non-square blocks, several conventional angle intra-prediction modes are replaced with wide-angle modes. The replaced modes are notified using the original method and remapped to the wide-angle mode index after analysis. The total number of core intra-prediction modes remains unchanged, i.e., 67.
[0053] For the current W×H block to be predicted, Figures 4C and 4D show a set of decoded reference samples consisting of an array of decoded reference samples above length 2W+1 and an array of decoded reference samples to the left length 2H+1. Figures 4C and 4D also show the relationship between the expansion of the decoded reference samples around the current W×H block and the range of acceptable intra-prediction angles. Table a below then presents examples of intra-prediction mode indices that are replaced by wide-angle modes in VVC and ECM, depending on the size of the current block to be predicted.
[0054] [Table 1]
[0055] Figure 4E shows an example of how angle intra-modes are replaced by wide-angle modes for non-square blocks where the width is strictly greater than the height. In this example, mode 2 is replaced by wide-angle mode 67. Mode 3 is replaced by wide-angle mode 68. For example, if the current block to be predicted is 8x4, this replacement process proceeds incrementally until mode 7 is replaced by wide-angle mode 72.
[0056] The current block to be encoded / decoded can be predicted using template-based intra-mode derivation (TIMD). TIMD derives one or two intra-prediction modes for the current block from the template of the current block. For the current block to be encoded / decoded, TIMD follows a two-step process: an intra-prediction mode index derivation step with the template of the decoded reference sample of the current block, and a step in which the current block is actually predicted.
[0057] More precisely, for a given block (1503) in Figure 15(a), the following intra-prediction mode derivation via TIMD is applied in the same manner to both the encoder and decoder sides. For each intra-prediction mode in the Most Probable Modes (MPM) list of this block, supplemented as necessary with default modes, TIMD determines the predictions of templates (1500 and 1501) for this block from the decoded reference samples of template (1502), and the SATD between this prediction and the templates for this block is calculated. The two intra-prediction modes with the minimum SATD are selected as TIMD modes. Note that in the case of TIMD, the set of directional intra-prediction modes is expanded from 65 to 129 by inserting directions between each ordinary black arrow and its neighboring dashed arrow in Figure 4A. This means that the set of possible intra-prediction modes derived via TIMD comprises 131 modes. After retaining two intra-prediction modes from the first pass of the test, which includes the MPM list supplemented in default mode, for each of these two modes, if this mode is neither PLANAR nor DC, TIMD also tests its two nearest extended directional intra-prediction modes with respect to predictive SATD. Note that above it is assumed that the block does not go outside the boundary of the current frame. If at least a portion of the block's template goes outside the boundary of the current frame, the template (1501, 1500) of block (1503), and the reference sample (1502) used to predict the template are modified as shown in Figures 15(b) and 15(c).
[0058] To predict the current block via TIMD, two predictions of the block via two TIMD modes (mode 1 and mode 2) resulting from two passes of the test are fused with weights after applying PDPC. The weights used depend on the prediction SATD(costMode1, costMode2) for the two TIMD modes.
[0059] Position-dependent intra-predictive combinations (PDPCs) are included in the derivation of TIMD modes. Therefore, any one of the embodiments described herein that applies to PDPC can also be used when applying PDPC in the derivation of TIMD modes.
[0060] In ECM, the costs of two selected modes (mode 1 and mode 2) are compared to a threshold, and in the test, a cost factor of 2 is applied as follows: costMode2<2 * costMode1 is the cost of the second-order intra-prediction mode, and costMode1 is the cost of the first-order intra-prediction mode.
[0061] If this condition is true, fusion is applied; otherwise, only mode 1 is used.
[0062] For example, the weights of the modes are determined from their SATD costs, as follows: weight1=costMode2 / (costMode1+costMode2) weight2 = 1 - weight1
[0063] In the case of TIMD, the set of directional intra-predictive modes is expanded from 65 to 129, so intra-predictive mode substitution in WAIP is applied. Table a above becomes Table b below. For example, for a given 8x4 block using TIMD, mode 2 is replaced by wide-angle mode 131, mode 3 is replaced by wide-angle mode 132, mode 4 is replaced by wide-angle mode 133, ..., mode 12 is replaced by wide-angle mode 141.
[0064] [Table 2]
[0065] Position-dependent pixel combination (PDPC) in VVC or ECM is a post-processing tool in intra-prediction. Its purpose is to eliminate discontinuities arising from initial intra-predictions in a particular prediction mode at target block boundaries adjacent to reference samples. This is achieved by using a weighted combination of the initial prediction value and one or more neighboring reference samples. This is effective not only for two non-angle modes, namely PLANAR mode and DC mode, but also for purely horizontal and purely vertical modes, as well as angle modes in the direction from the lower left corner to the upper right corner of the block and vice versa. Depending on the prediction direction, either normal PDPC or gradient PDPC is applied.
[0066] Intra prediction in VVC ("Versatile Video Coding (Draft 8)", B. Bross, J. Chen, S. Liu, and Y.-K. Wang, JVET-Q2001-vD, JVET Meeting, Jan 2020, Brussels, Belgium) and ECM ("Algorithm description of Enhanced Compression Model 6 (ECM 6)", M. Coban, F. Le Leannec, K. Naser, J. Strom, L. Zhang, JVET-AA2025, JVET Meeting, July 2022, Teleconference) include position-dependent pixel combinations (PDPC) as a post-processing tool for certain prediction modes that have the potential for intensity discontinuities to the left or above the target block. In particular, this is enabled for PLANAR mode, DC mode, pure horizontal mode, pure vertical mode, and modes associated with directions from bottom left to top right or vice versa.
[0067] PDPC is also used as a post-processing step in the derivation process of the intra-prediction mode of the TIMD coding mode described above.
[0068] The application of PDPC processing depends on the intra-prediction mode, as follows:
[0069] For PLANAR and DC modes, PDPC is applied to the first predicted values on both the upper and left sides of the target block. For purely vertical or purely horizontal modes, PDPC is applied to the first predicted values on the left or upper side of the block, respectively. For other eligible angular modes, it is applied to either the left or upper side, depending on whether the mode is vertical or horizontal. The first predicted values in the top row or left column are gracefully corrected using weighted combinations with the reference samples on the upper or left side of the block, respectively. Without PDPC, the reconstructed frame may have blocking artifacts resulting from the quantization of high-frequency coefficients. Therefore, PDPC is employed in VVC and ECM.
[0070] For angular modes eligible for PDPC, the process involves a weighted combination of the first predicted value and a secondary reference sample, as shown in Figure 5. In some modes, due to the finite length of the secondary reference array, secondary reference samples may not be available for some target pixels. In these cases, for the current block, PDPC is replaced by gradient PDPC ("Unified PDPC for Angular Intra Modes", B. Ray, GVder Auwera, M. Karczewicz, JVET-Q391, JVET Meeting, Jan 2020, Brussels, Belgium), as shown in Figure 6, where a weighted value of the gradient calculated for the secondary reference sample in the same row (for vertical mode) or column (for horizontal mode) as the target pixel is added to the initial predicted value. These two cases are determined by calculating a scale value from the predicted angle and the length of the secondary reference array. A negative scale value indicates that the length of the secondary reference array is insufficient, and therefore gradient PDPC is enabled. Gradient PDPC is similar to PDPC applied purely vertically or purely horizontally, but the gradient is calculated along the current forecast direction. The calculations involved in gradient PDPC differ from those of regular PDPC, but they both have the same form and aim for the same goal.
[0071] VVC and ECM define 67 prediction modes for intra-prediction of any target block. Of these modes, two are non-angle modes (i.e., mode 0, i.e., PLANAR mode, and mode 1, i.e., DC mode), and the remaining 65 are angle modes illustrated in Figure 4A. These modes are associated with prediction directions in the range of 45 degrees to -135 degrees clockwise. Depending on the block shape, some angle modes are replaced by an equal number of wide-angle modes defined beyond the above range. The angle modes and wide-angle modes in VVC are shown in Figure 4B.
[0072] A mode is called horizontal if it points in a direction below the diagonal (i.e., from the top left to the bottom right), and vertical otherwise. These are further called positive or negative depending on whether they belong to a purely horizontal direction (i.e., below or above a purely horizontal direction) or a purely vertical direction (i.e., to the right or left of a purely vertical direction). Thus, the downward direction, which includes a purely horizontal direction, and the rightward direction, which includes a purely vertical direction, are called positive directions. The remaining directions are called negative directions. PDPC is only enabled for positive angle modes.
[0073] The following describes an example of the PDPC process in the positive vertical direction. This is similar to the positive horizontal direction, where the reference array is swapped and the width and height of the target block are swapped.
[0074] Figure 5 shows the PDPC in the intra-prediction for the positive vertical direction for a target block of pixels shown in white in Figure 5, and the gray squares represent reference samples, which are reconstructed neighbor samples. Coordinate (0,0) addresses the top-left sample within the block. For a target pixel at location (x,y) within the block, the first prediction P(x,y) is obtained from the upper reference array at (x',-1). If the prediction direction does not have an integer slope, the reference sample is interpolated using a smoothing filter or a cubic interpolation filter. Thus, the predictor is, The prediction is estimated as P(x,y)=R(x',-1), where R(x,y) is the reconstructed neighboring sample array. Hereafter, the reconstructed neighboring sample used to obtain the first prediction is called the primary reference sample, and the reconstructed neighboring sample used in the PDPC process is called the secondary reference sample. The primary reference sample is referred to herein as the reference sample located within the primary reference array used by the intra-prediction mode to construct a prediction about the target pixel. In the example in Figure 5, the primary reference array is the upper reference array (the array of reconstructed samples above the target block). For positive horizontal intra-prediction, the primary reference array is the left reference array (the array of reconstructed samples to the left of the target block).
[0075] The secondary reference sample is referred to herein as the reference sample located within the secondary reference array. The secondary reference array is an array of reconstructed samples obtained by extending the angular prediction direction of the intra-prediction mode beyond the target block. In the example in Figure 5, the secondary reference array is the left reference array (the array of reconstructed samples to the left of the target block). For positive horizontal intra-prediction, the secondary reference array is the upper reference array (the array of reconstructed samples above the target block).
[0076] In the PDPC process, the prediction direction is extended to obtain a secondary reference sample R(-1,y') that intersects the secondary boundary reference sample array, i.e., the left reference array in the case of the vertical prediction direction. For lower complexity, when the extension does not pass through the reference sample location, the secondary reference sample is selected as the nearest neighbor. If absInvAngle represents the reciprocal of the tangent of the angle value corresponding to the prediction direction, then the y-coordinate y' of the secondary reference sample is: y' = 1 + y + (((1 + x) * It is obtained as absInvAngle+256)>>9), where >> represents a binary right shift.
[0077] The first predicted value at (x,y) is: P(x, y) = P(x, y) + (wL * is modified as (R(-1, y’) - P(x, y)) + 32) >> 6, where the weight parameter wL is wL = 32 >> (2 * x >> scale), and the parameter scale is calculated as scale = min(2, Log2(height) - (Log2(3 * absInvAngle - 2) - 8))
[0078] In this example, it is assumed that the scale parameter is a positive integer between 0 and 2. As x increases, the value of wL decreases to 0, so the scale parameter determines the number of columns in the target block to be modified in the PDPC process. Since the maximum value of wL is 32, the number of columns to receive PDPC is (3 << scale). Table 1 shows the number of columns used for PDPC and the corresponding values of wL for different scale parameter values.
[0079]
Table 3
[0080] Depending on the case, the number of columns may be larger than the target block width, so the actual number of columns modified by PDPC is given as min((3 << scale), width).
[0081] For prediction along the positive horizontal direction, the process remains the same. The secondary reference sample R(x’, -1) is obtained by calculating the coordinate x’ as x’ = 1 + x + (((1 + y) * absInvAngle + 256) >> 9)
[0082] (x,y) The first predicted value at is P(x,y)=P(x,y)+(wL * (R(x’,-1)-P(x,y))+32) is modified as >>6, where the weight parameter wL is wL=32>>(2 * y>>scale), and the parameter scale is calculated as >>scale), and the parameter scale is calculated as scale=min(2,Log2(width)-(Log2(3 * absInvAngle-2)-8)) is calculated as When the scale parameter is a positive integer between 0 and 2, as y increases, the value of wL decreases to 0. Therefore, the scale parameter determines the number of rows in the target block to be modified in the PDPC process. Since the maximum value of wL is 32, the number of rows receiving PDPC is (3<<scale).
[0083] When the prediction direction is purely vertical (or purely horizontal in the case of horizontal mode) depending on the block height (or width in the case of horizontal mode), the scale parameter calculated above may be a negative integer (i.e., less than 0). This implies that not all secondary reference samples may be available even for the three pixels in the last row (or last column in the case of horizontal mode) of the target block. In this case, gradient PDPC is enabled.
[0084] Figure 7 shows a method 700 for determining which PDPC to apply to a target block. At 701, the scale parameter is determined as described above. For the vertical prediction direction, scale=min(2,Log2(height)-(Log2(3 * absInvAngle-2)-8)), and For the horizontal prediction direction, scale=min(2,Log2(width)-(Log2(3 * absInvAngle-2)-8)).
[0085] In step 701, if the scale parameter is equal to 0 or positive, the process proceeds to step 702, where normal PDPC is applied as described above; otherwise, the process proceeds to step 703, where gradient PDPC is applied as described below. When gradient PDPC is applied in step 703, the scale parameter is: The scale is recalculated as follows: scale=(Log2(height)+Log2(width)-2)>>2
[0086] Figure 6 shows an example of gradient PDPC applied to a target block. For the target pixel at location (x,y), the reference sample R(-1,y) on the left reference array in the same row as the target pixel is used as the secondary reference sample in the positive vertical direction. The gradient value is determined by finding the predictor sample R(x",-1) of the secondary reference sample R(-1,y) in the prediction direction. The gradient is added to the first prediction P(x,y) at (x,y) using weighting. P(x,y) = Clip(P(x,y) + (wL * (R(-1,y)-R(x",-1))+32)>>6) In the formula, R(x",-1) represents the predictor sample of the second-order reference sample at (-1,y), and wL is calculated using the recalculated scale parameter wL=32>>(2 * It is calculated as x >> scale.
[0087] Predictor samples are linearly interpolated whenever x'' does not pass through the reference sample index. The values are clipped to the dynamic range of the component, as they are not guaranteed to be within a range. The recalculated scale parameter has a minimum value of 0 and a maximum value of 3 (for a maximum CU size of 128×128).
[0088] As described above, the number of columns receiving the gradient PDPC is (3 << scale), and the scale value is recalculated. Note that for all eligible pixels in any row of the target block, the secondary reference samples, and thus the determined gradient values, are the same. Therefore, they are determined outside the loop (unlike the normal PDPC) for the pixels of the row.
[0089] For the positive horizontal direction, the process is similar, where columns are replaced by rows and the secondary reference samples are in the upper reference array. The number of rows receiving the gradient PDPC is (3 << scale), and the scale value is recalculated. Similar to the case of the positive vertical direction, for all eligible pixels in any column of the target block, the secondary reference samples, and thus the determined gradient values, are the same. Therefore, they are determined outside the loop (unlike the normal PDPC) for the pixels of the column.
[0090] Some embodiments provide a method for improving intra prediction, more specifically, a method in the case of an angular prediction direction. Some embodiments provide a method for integrating the normal PDPC and the gradient PDPC applied to the angular mode in intra prediction.
[0091] In one embodiment, the two PDPC processes are combined with binary weights, and as a result, either one instead of both is used at any given moment. This leads to a simplification of the existing code but has the same result as the original code.
[0092] In another embodiment, the two PDPCs are combined with variable weights, and the weights are derived based on the prediction direction and the block size. As the prediction direction approaches a purely vertical or purely horizontal direction, the PDPC gradually changes from the normal PDPC to the gradient PDPC, with a combination of both in between.
[0093] Position-dependent intra-predictive combinations (PDPCs) are included in the derivation of TIMD modes. Therefore, any one of the embodiments described herein that applies to PDPC can also be used when applying PDPC in the derivation of TIMD modes.
[0094] Figure 8 shows an example of a method 800 for encoding video according to one embodiment. In 801, predictor blocks of video blocks to be encoded are obtained based on an intra-prediction mode using one or more primary reference samples. To this end, a first prediction is obtained for each pixel of the video block using the intra-prediction mode.
[0095] Preferably, the intra-prediction mode is an angular mode, more specifically, a mode using a positive horizontal or positive vertical direction. The first prediction is obtained using one or more primary reference samples determined from the prediction direction in the primary reference sample array.
[0096] In 802, the predictor block is modified using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the block. In some embodiments, the first prediction obtained in 801 for at least one pixel is modified using a weighted combination of at least two gradients, the at least two gradients being determined using secondary reference samples. In one embodiment, one of the at least two secondary reference samples is obtained by extending the angular intra-prediction mode toward the secondary reference array, and the other of the at least two secondary reference samples is located in the secondary reference array in the same column or row as at least one pixel, depending on whether the intra-prediction mode is vertical or horizontal. In another embodiment, if a secondary reference sample obtained by extending the angular intra-prediction mode toward the secondary reference array is not available for at least one pixel, the weighted combination uses a pixel in the predictor block in the first column or row of the predictor block and in the same row or column as at least one pixel, and a secondary reference sample obtained for that pixel in the first column or row.
[0097] In some embodiments, the weighted combination is a position-dependent pixel combination (PDPC) that uses two secondary reference samples and another primary reference sample for at least one pixel in a block.
[0098] At least two secondary reference samples are different, and the other primary reference sample is acquired as a predictor for at least one of the two secondary reference samples according to the intra-prediction mode. In 803, the video block is encoded using the modified predictor block. Encoding method 800 can be implemented in the intra-prediction module of a video encoder, such as one of the encoders 200 in Figure 2.
[0099] Figure 9 shows an example of a method 900 for decoding video according to one embodiment. In 901, predictor blocks of video blocks to be encoded are obtained based on an intra-prediction mode using one or more primary reference samples. To this end, a first prediction is obtained for each pixel of the video block using the intra-prediction mode.
[0100] Preferably, the intra-prediction mode is an angular mode, more specifically, a mode using a positive horizontal or positive vertical direction. The first prediction is obtained using one or more primary reference samples determined from the prediction direction in the reference sample array.
[0101] In 902, the predictor block is modified using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the block. In some embodiments, the first prediction obtained in 901 for at least one pixel is modified using a weighted combination of at least two gradients, the at least two gradients being determined using secondary reference samples. In one embodiment, one of the at least two secondary reference samples is obtained by extending the angular intra-prediction mode toward the secondary reference array, and the other of the at least two secondary reference samples is located in the secondary reference array in the same column or row as at least one pixel, depending on whether the intra-prediction mode is vertical or horizontal. In another embodiment, if a secondary reference sample obtained by extending the angular intra-prediction mode toward the secondary reference array is not available for at least one pixel, the weighted combination uses a pixel in the predictor block in the first column or row of the predictor block and in the same row or column as at least one pixel, and a secondary reference sample obtained for that pixel in the first column or row.
[0102] In some embodiments, the weighted combination is a position-dependent pixel combination (PDPC) that uses two secondary reference samples and another primary reference sample for at least one pixel in a block.
[0103] At least two secondary reference samples are different, and the other primary reference sample is acquired as a predictor for at least one of the two secondary reference samples, according to the intra-prediction mode.
[0104] In 903, the video block is encoded using the modified predictor block. The decoding method 900 can be implemented in an intra-prediction module of a video decoder, such as one of the decoders 300 in Figure 3.
[0105] Several embodiments for determining weighted combinations are described below.
[0106] In some embodiments, the weighted combination includes a first term (referred to herein as a normal PDPC or PDPC) relating to a position-dependent pixel combination using a first secondary reference sample determined based on an intra-prediction mode, and a second term (referred to as a gradient PDPC) relating to a gradient position-dependent pixel combination using a second secondary reference sample located in the same row or column as the target pixel, and a primary reference sample determined for the second secondary reference sample.
[0107] In a modified example, modifying a predictor block involves modifying the predicted value for at least one pixel of the predictor block using a weighted combination of the first and second terms. In this embodiment, the first weight associated with the first term and the second weight associated with the second term are non-zero. This provides a modified predicted value in which both PDPC and gradient PDPC contribute to the modification.
[0108] An embodiment of the integrated PDPC is shown in Figure 10.
[0109] For a given target pixel at (x,y), two secondary reference samples are considered: secondary reference sample 1, determined in the same way as in a normal PDPC, and secondary reference sample 2, determined in the same way as in a gradient PDPC. The gradient in secondary reference sample 2 is calculated in the same way as in a gradient PDPC, by determining the predicted sample at (x),-1. The predicted value at (x,y) is modified as follows. P(x,y) = Clip(P(x,y) + (wL * (R(-1,y')-P(x,y))+wL1 * ((R(-1,y)-R(x”,-1))-(R(-1,y')-P(x,y)))+32)>>6)
[0110] The weights wL and wL1 are determined as follows: wL=32>>(2 * (x >> scale) wL1=wL>>(6 * bPdpc) In the formula, the scale and parameter bPdpc are determined as follows: The parameter pdpcScale is first pdpcScale=min(2,Log2(height)-(Log2(3 * It is obtained as absInvAngle-2)-8)).
[0111] Next, if pdpcScale is greater than or equal to 0 (pdpcScale ≥ 0), the scale parameter is set to the pdpcscale (scale=pdpcScale) determined above, and the bPdpc flag is set to 1 (bPdpc=1).
[0112] Otherwise (if pdpcScle is less than 0), the scale parameter is: The scale is recalculated as follows: scale=(Log2(height)+Log2(width)-2)>>2 The flag bPdpc is set to 0 (bPdpc=0).
[0113] Similar to the determination of the scale parameter, the determination of the flag bPdpc is performed once per block.
[0114] As can be seen from this, in the first case (pdpcScale≧0), the maximum value of wL is 32, so wL1=(wL>>6)=0. Therefore, in this example, the integrated PDPC is P(x,y) = Clip(P(x,y) + (wL * (R(-1,y')-P(x,y))+32)>>6)) and so This is equivalent to a standard PDPC.
[0115] In the second case (pdpcScale<0), wL1=(wL>>0)=wL. Therefore, in this example, the integrated PDPC is P(x,y) = Clip(P(x,y) + (wL * (R(-1,y')-P(x,y))+wL * ((R(-1,y)-R(x”,-1))-(R(-1,y')-P(x,y)))+32)>>6) =Clip(P(x,y)+(wL * (R(-1,y)-R(x",-1))+32)>>6), This is equivalent to gradient PDPC.
[0116] The integrated PDPC can be expressed equivalently as follows: P(x,y) = Clip(P(x,y) + ((wL * G(-1,y')+wL1 * (G(-1,y)-G(-1,y'))+32)>>6)). In the formula, G(-1,y) represents the gradient calculated at (-1,y) in the prediction direction.
[0117] For positive horizontal directions, columns are swapped with respect to rows, and vice versa; the process remains the same. Two secondary reference samples are in the upper reference array, and the predictor sample (primary reference sample) is in the left reference array. The weights wL and wL1 are calculated as follows: wL=32>>(2 * y >> scale) wL1=wL>>(6 * bPdpc) In the formula, the scale and parameter bPdpc are calculated as follows: pdpcScale=min(2,Log2(width)-(Log2(3 * absInvAngle-2)-8)). If pdpcScale ≥ 0, the scale parameter is set to pdpcScale and bPdpc is set to 1. Otherwise (pdpcScale > 0), the scale parameter is set to The scale is determined as follows: scale = (Log2(height) + Log2(width) - 2) >> 2, and bPdpc is set to 0.
[0118] In this example, the integrated PDPC process is represented as the following weighted combination: P(x,y) = Clip(P(x,y) + ((wL * G(x',-1)+wL1 * (G(x,-1)-G(x',-1))+32)>>6)). In the formula, G(x,-1) represents the gradient calculated at (x,-1) in the prediction direction.
[0119] Such integration yields the same results as, for example, the ECM7.0 code. This is because, in the weighted combinations given above, either the normal PDPC or the gradient PDPC contributes to correcting the predicted values at one time.
[0120] In another embodiment, this constraint is removed so that both normal PDPC and gradient PDPC contribute to correcting the predicted values. Their partial contributions vary depending on the prediction direction and block size.
[0121] In this embodiment, the parameter pdpcScale is determined as follows for the positive vertical direction. pdpcScale=max(-2,min(2,Log2(height)-(Log2(3 * absInvAngle-2)-8))).
[0122] The weight of each PDPC contribution is set based on the parameter pdpcScale. If pdpcScale is greater than or equal to 0 (pdpcScale ≥ 0), the scale parameter is set to pdpcScale, and the weights are set as follows: scale=pdpcScale; wL=32>>(2 * (x >> scale) wL1=wL>>(2+2 * (scale)
[0123] Therefore, weighted combinations are, P(x,y) = Clip(P(x,y) + ((wL * G(-1,y')+wL1 * It is given by (G(-1,y)-G(-1,y'))+32)>>6)).
[0124] The gradient G(-1,y') at (-1,y') and the gradient G(-1,y) at (-1,y) are determined in the same manner as described above. G(-1,y') = R(-1,y') - P(x,y); G(-1,y) = R(-1,y) - R(x", -1). If pdpcScale is greater than or equal to 0, secondary reference samples 1 and 2 are available for all eligible target pixels in the predictor block, and therefore the above determination is not difficult. In this case, the values of wL and wL1 are given in Table 2 below.
[0125] [Table 4] Table 2: pdpcScale≧0. wL=32>>(2 * x >> scale) and wL1 = wL >> (2 + 2 * In the case of pdpcScale, the number of columns or rows processed in one embodiment of the integrated PDPC, and the corresponding weights for predictions along the positive vertical direction.
[0126] As can be seen, at a scale value of 2, the contribution of gradient PDPC is zero, making it equivalent to a normal PDPC. However, at scale=0 and scale1, gradient PDPC contributes to up to two columns with a higher contribution rate at scale=0 than at scale=1.
[0127] The integrated PDPC process provided herein allows one or more pixels to correct their predicted values using both normal PDPC and gradient PDPC. In other words, the integrated PDPC process allows one or more pixels to use two or more secondary reference samples to correct their predicted values when angle prediction mode is used.
[0128] If the parameter pdpcScale is negative (pdpcScale < 0), secondary reference sample 1 is not available for all eligible target pixels. However, when the parameter pdpcScale is equal to -1, it is observed that secondary reference sample 1 is always available for the first column of target pixels. Thus, the gradient in secondary reference sample 1 determined for the first column of target pixels is used in addition to the gradient determined in gradient PDPC. This is illustrated in Figure 11, which shows the target pixel P(0,y) in the first column of the predictor block and the secondary reference sample 1 R(-1,y') used for the available P(0,y).
[0129] If pdpcScale is negative (pdpcScale < 0), the scale parameter and weights are determined as follows: scale=(Log2(height)+Log2(width)-2)>>2; wL=32>>(2 * (x >> scale) wL1=wL>>(2-2 * (pdpcscale)
[0130] The weighted combination for correcting the predicted values of the pixels in the predictor block is: P(x,y) = Clip(P(x,y) + ((wL * G(-1,y)+wL1 * It is given by (G(-1,y')-G(-1,y))+32)>>6)).
[0131] The y-coordinate y' of secondary reference sample 1 is: y'=y+1+(256+absInvAngle * (pdpcScale+2))>>9 is obtained as follows: The gradient G(-1,y') is determined as G(-1,y')=R(-1,y')-P(0,y).
[0132] Unlike the previous case, here y' is not a function of x, but a fixed integer that depends on the absInvAngle value. Therefore, for a given row, the gradient G(-1,y') can be pre-calculated as gradient G(-1,y). In this case, the values of wL and wL1 are given in Table 3 below.
[0133] [Table 5] Table 3: pdpcScale<0. wL=32>>(2 * x >> scale) and wL1 = wL >> (2-2 * The number of columns or rows processed in the proposed integrated PDPC (pdpcScale), and the corresponding weights for predictions along the positive vertical direction.
[0134] The gradient at (-1,y') is calculated only once, just like the gradient at (-1,y), and is used for the target pixel at (x,y), x≧0.
[0135] As can be seen from this, when pdpcScale = -2 or less, the contribution of normal PDPC is zero, and it becomes equivalent to gradient PDPC.
[0136] In another embodiment, the original gradient can be replaced by a weighted combination of the two gradients outside the loop, using weights wL=32 and wL1=32>>4=2 when pdpcScale=-1, and wL1=32>>6=0 when pdpcScale=-2. In this embodiment, the resulting combined gradient value is used to modify the value of the predictor block with weight wL, as shown in the third column of Table 3, and the fourth and fifth columns of wL1(x) are not required.
[0137] In the embodiments described above, separate weighted combinations are used depending on whether the parameter pdpcScale is positive or negative. In another variation, the embodiments described above (pdpcScale≧0 and pdpcScale<0) are used. P(x,y) = Clip(P(x,y) + ((w1 * G(-1,y')+w2 * It can be combined into a single equation, as in G(-1,y)+32)>>6)). In the formula, weights w1 and w2 are derived as follows. pdpcScale=max(-2,min(2,Log2(height)-(Log2(3 * absInvAngle-2)-8))). If pdpcScale is positive or null (i.e., pdpcScale ≥ 0), the scale parameter is scale = pdpcScale, and the weights are as follows: wL=32>>(2 * (x >> scale) wL1=wL>>(2+2 * (pdpcscale) w1 = wL - wL1; w2 = wL1; Otherwise (when pdpcScale is negative), the scale parameter is set to scale=(Log2(height)+Log2(width)-2)>>2, and the weights are as follows: wL=32>>(2 * (x >> scale) wL1=wL>>(2-2 * (pdpcscale) w1 = wL1; w2 = wL - wL1;
[0138] This embodiment allows for having a single weighted combination for modifying the values of the predictor block. Only the weights are determined based on the values of pdpcScale and depending on whether the prediction direction is vertical or horizontal.
[0139] In this embodiment, the weight wL1 is part of wL and is derived using the parameter pdpcScale, which here depends on both the block size (height of the vertical angle and width of the horizontal angle) and the predicted angle. Alternatively, w1 can be derived in other ways, for example, based only on the predicted angle, as follows: pdpcScale=min(2,Log2(height)-(Log2(3 * absInvAngle-2)-8)); gScale1=15-Log2(3 * absInvAngle-2); gScale2=max(0,Log2(3 * absInvAngle-2)-8); If pdpcscale ≥ 0 {scale=pdpcScale; wL=32>>(2 * (x >> scale) wL1=wL>>gScale1 w1 = wL - wL1; w2=wL1;} Otherwise {scale=(Log2(height)+Log2(width)-2)>>2; wL=32>>(2 * (x >> scale) wL1 = wL >> gScale2 w1 = wL1; w2 = wL - wL1;}
[0140] The values of the parameters gScale1 and gScale2 depend either solely on absInvAngle, or equivalently, on intraPredAngle, which corresponds to the prediction direction. Table 4 below lists the gScale1 and gScale2 values corresponding to the intraPredAngle values.
[0141] [Table 6] Table 4: Values of the scale parameters gScale1 and gScale2 as functions of abs(intraPredAngle). Here, A represents intraPredAngle.
[0142] The integrated PDPC embodiment is described above using the positive vertical direction. The integrated PDPC is similar in the positive horizontal direction, where height and width are swapped and rows and columns are swapped.
[0143] In addition, an embodiment of the integrated PDPC is described herein with respect to video blocks included in the image of a video. The embodiments described herein can also be similarly applied to image blocks in an image encoder or decoder.
[0144] Furthermore, the above explanation uses the wL value, as with VVC and ECM. This can be replaced with any other decreasing function. The scale value used to derive wL1 is also an example. Instead of using it as follows, if pdpcScale≧0, then wL1=wL>>(2+2 * (scale) Other functions can be used as follows: wL1 = wL >> (2 + scale), wL1=wL>>(1+2 * scale), wL1=wL>>(1+3 * scale), These are all examples of increasing scale functions. Similarly, when pdpcScale<0, the following wL1=wL>>(2-2 * Instead of using pdpcscale, Other decreasing functions can be used as follows: wL1 = wL >> (2 - pdpcScale), wL1=wL>>(1-2 * (pdpcscale), wL1=wL>>(1-3 * (pdpcscale).
[0145] The proposed method described above uses the same interpolation method as used in PDPC or gradient PDPC from VVC and ECM. Specifically, nearest neighbor interpolation is used for the second-order reference sample 1, and linear interpolation is used for the predictor of the second-order reference sample 2.
[0146] In other embodiments, a linear or higher-order filter, such as a 4-tap filter or a 6-tap cubic filter, can be used for the secondary reference sample 1.
[0147] In another variation, depending on whether the primary reference sample uses a smoothing filter or a cubic filter, a similar filter can be used to interpolate the predictor of the secondary reference sample 2.
[0148] The integrated PDPC described herein is implemented using ECM7.0. Table 5 shows the BD rate performance. As can be seen, there is an overall BD rate gain of approximately -0.01% for the lumen component. Class F sequences yield the best results with a gain of -0.04%. It should be noted that coding efficiency is not the only advantage of the method provided herein. The integration of PDPC and gradient PDPC also simplifies the video decoding process design by specifying a single method instead of the two earlier switchable methods, PDPC and gradient PDPC.
[0149] [Table 7]
[0150] In the following embodiments, a video codec that includes PDPC in its intra-prediction is assumed, such as a VVC standard or a codec based on the ECM video compression search model.
[0151] In one embodiment, the RD performance of an integrated PDPC that enables the contributions of conventional PDPC and gradient PDPC described herein is compared to either a conventional PDPC or gradient PDPC implementation, such as those performed in VVC or ECM. The better method is selected and signaled to the decoder along with an indicator.
[0152] In a modified example, the performance of the integrated PDPC, which enables the contributions of the conventional PDPC and gradient PDPC described herein, is compared to the performance of either the conventional PDPC or gradient PDPC performed in VVC or ECM using a template, such as the one used in template-based intra-mode derivation (TIMD) of VVC. An example of such a template is shown in Figure 15(a). For positive angular modes, both PDPC (integrated PDPC or VVC PDPC) techniques are tested with the template based on their SATD between template prediction and reconstruction, and the better method for encoding or decoding the video block is selected. In this case, the selected method can be signaled with an indicator, or the decoder can infer the same thing using a similar template.
[0153] In any of the embodiments described above, instead of taking only the nearest neighbor sample, a secondary reference sample 1 can be used which is interpolated using either linear interpolation or a 4-tap filter or a 6-tap cubic filter.
[0154] In any of the embodiments described above, a predictor for the secondary reference sample 2 can be used that is interpolated using either a 4-tap filter, a 6-tap cubic filter, or a smoothing filter instead of the default linear interpolation. The cubic filter or smoothing filter is selected based on whether the primary reference sample uses a cubic filter or a smoothing filter for interpolation, respectively.
[0155] ECM version 7 uses gradient PDPC only for the lumern component, while a standard PDPC is used for both the lumern and chromar components. In another embodiment, the integrated PDPC provided herein is used for both the lumern and chromar components.
[0156] In another embodiment, one of the embodiments described above uses the activation of the integrated PDPC provided herein, which is signaled in the slice header of the picture containing the video block to be encoded or decoded, or in the PPS header or SPS header.
[0157] As described above, following an initial prediction of a given block via a given intra-prediction mode, if PDPC is permitted, blending a given predicted sample with a secondary reference sample accessed by PDPC often improves the quality of the prediction, provided that this secondary reference sample does not result from padding of available reference samples located far away. This is because the correlation between this padded secondary reference sample and the current original block sample to be predicted is likely to be small.
[0158] In a modified embodiment, for a given block, following its initial prediction via a given intra-prediction mode, if PDPC is permitted, for each predicted sample that results from padding of available reference samples located far away from the secondary reference sample accessed by PDPC, PDPC cancels for that predicted sample.
[0159] For example, Figure 16 shows two examples of this modified embodiment for a given block (shown as a white rectangle) predicted via the directional intra-prediction mode in the ECM.
[0160] In Figure 16(a), following the initial prediction of a width × height block (1600), PDPC is applied from that set (1601) of reference samples via a vertical positive intra-prediction mode, indicated by a thick arrow in the direction.
[0161]
number
[0162]
number
[0163] According to the modifications of the embodiments described herein, n bottommost reference samples are unavailable, and δ≧δ limit In this case, the PDPC for the predicted sample at position (x,y) is canceled. For example, n = height. For example, δ limit =height+1. In this case, the sample prediction provided by the primary reference sample is not modified by the unavailable secondary reference sample.
[0164] δ limitOther values for this may also be considered. For example, some of the padded reference samples (1602) that are closest to the reference sample used for padding (1603) can be used in PDPC. For example, δ limit It can be set to height+2 or height+3...
[0165] In Figure 16(b), following the initial prediction of a width × height block (1600), PDPC is applied from that set (1604) of reference samples via a horizontal positive intra-prediction mode, indicated by a thick arrow in the direction.
[0166]
number
[0167]
number
[0168] According to another variation of the embodiment described herein, p rightmost reference samples are unavailable, and γ ≥ γ limit In this case, the PDPC for the predicted sample at position (x,y) is canceled. For example, p = width For example, γ limit = width + 1. For the vertical case, γ limit Other values can be used for this.
[0169] In another embodiment, enabling or disabling PDPC can be considered on a block basis, rather than on a sample basis. For example, if the current block does not have at least one available neighboring block to its left, PDPC using the vertical angle mode is disabled. Similarly, if the current block does not have at least one available neighboring block above it, PDPC using the horizontal angle mode is disabled.
[0170] Figure 17 provides a diagram of this modified embodiment. In Figure 17(a), the light gray current width × height reference sample block is available, while the dark gray block is unavailable. Since the current block has no neighboring available blocks to its left, PDPC is disabled for any prediction of the current block via the vertical angle mode.
[0171] In Figure 17(b), the light gray current width×height reference sample block is available, but the dark gray block is unavailable. Since the current block has no neighboring available blocks above it, PDPC is disabled for any prediction of the current block via the horizontal angle mode.
[0172] In the embodiments described above, neighboring blocks may be unavailable, for example, because they do not exist, for example because the current block is on the boundary of a picture, or in another example, because neighboring blocks are unavailable to the current block, for example because neighboring blocks and the current block are encoded in separate tiles that are encoded independently of each other.
[0173] The above-described modified embodiments can be used in the same manner for all coding modes that use the PDPC tool, or a separate modification that applies PDPC can be used depending on the coding mode.
[0174] In the modified embodiment, the above-mentioned deactivation of the PDPC can be performed only in the decoder (including the decoder in the encoder). The encoder can then use a PDPC having a qualified angular mode regardless of the availability of the upper and left neighboring blocks.
[0175] As explained above, for a given block, if PDPC is allowed following an initial prediction of this block via a given intra-prediction mode, PDPC allows for the removal of some discontinuities between predicted samples and reference samples around the boundaries of the predicted block. Therefore, around the boundaries of the predicted block, the closer the predicted samples are to the reference samples, the better PDPC works.
[0176] However, for a given block using TIMD, during the TIMD derivation step, during the prediction of the template of this block from a reference sample of the template via a given intra-prediction mode, the template design is characterized by gaps between the template and its reference sample, thus reducing the effectiveness of PDPC. Such gaps are shown in Figure 15, for example, as white squares between parts 1501 and 1502 of the template of block 1503.
[0177] If this hole is removed, PDPC can become more effective. Similarly, if this hole is removed, gradient PDPC can also become more effective.
[0178] Furthermore, it should be noted that a beneficial side effect of eliminating this gap is that template samples are often more correlated with the template's reference samples, thus improving the quality of not only PDPC but also the overall template prediction.
[0179] Another aspect of this disclosure provides a method for encoding or decoding a video block to which a template-based intra-predictive mode derivation (TIMD) is applied. More specifically, in some embodiments, for a video block using TIMD, instead of defining a set of reference samples of the template during the TIMD derivation step, this set is common to the upper and left portions of the template, and each portion of the template has a different set of reference samples. This makes it possible to eliminate gaps between the template and its reference samples. The reference samples used to predict the template are closer to the template, and therefore the prediction is improved. In addition, this adaptation of the set of reference samples makes it possible to maintain the same size of the template used in TIMD in the ECM, and therefore the same prediction units can be reused. For example, the two portions of the template may have sizes that are powers of 2.
[0180] Figures 18(a.1) and 18(a.2) show the current block (1800) to be predicted during the TIMD derivation step, its template, and a reference sample of the template. Figure 18(a.1) shows templates (1801 and 1802) used in the ECM and their reference sample (1803), and Figure 18(a.2) shows embodiments provided herein of the adaptation of the set of reference samples (1810, 1811) of the TIMD templates (1801 and 1802). In Figures 18(a.1) and 18(a.2), both the upper and left portions of the template for the current block (1800) are available.
[0181] In Figure 18(a.1), for a given width × height block (1800), during the TIMD derivation step, for a given intra-prediction mode to be tested on template (1800), which is a template including the left portion of template iTw × height (1801) and the upper portion of template width × iTh (1802), some reference samples in the set of 2(width + iTw) + 2(height + iTh) + 1 reference samples of the template (1803) are used to predict both (1801) and (1802). (1812) shows the iTw × iTh gap between the template and its reference samples that exists in the prediction of the template used in ECM.
[0182] In Figure 18(a.2), for a given intra-prediction mode to be tested on template (1800), some reference samples in the set of reference samples (1810) for the left template portion (1801) are used to predict the left template portion (1801), and some reference samples in the set of reference samples (1811) for the upper template portion (1802) are used to predict the upper template portion (1802).
[0183] Unlike Figure 18(a.1), Figure 18(a.2) does not contain a gap between the (1800) template and its reference sample.
[0184] In Figures 18(a.1) and 18(a.2), to provide an example of which reference samples are used to predict the template (1800), the black dotted arrows indicate the direction of extrapolation of the template reference samples to the template for the directional intra-prediction mode at index 48. In other words, for a given sample to be predicted in the template portion, the tail of the arrow crossing this sample places the reference sample at the center of the directional interpolation filter for calculating the prediction of this template sample.
[0185] This example clearly demonstrates how changing the set of reference samples for the template (1800) from Figure 18(a.1) to Figure 18(a.2) modifies the template prediction. For example, in Figure 18(a.2), the reference sample (1813) is accessed during the prediction of (1801) from (1810). However, in Figure 18(a.1), (1813) is not involved in the prediction of (1801). Note that the black-filled dots at the edges of markers like (1813) indicate that the marker labels a single pixel instead of a set of pixels with a shared color.
[0186] Similarly, in Figure 18(a.2), the reference sample (1814) is accessed during the prediction of (1802) from (1811). However, in Figure 18(a.1), (1814) is not involved in the prediction of (1802).
[0187] As another example, Figure 19 is a copy of Figure 18, except that the intra-prediction mode at index 112 is replaced with the intra-prediction mode at index 48.
[0188] In Figure 19(a.2), the reference sample (1913) is accessed during the prediction of the left portion (1901) of the template for block (1900) from the set of reference samples (1910). However, in Figure 19(a.1), (1913) is not involved in the prediction of (1901).
[0189] In TIMD in ECM-7.0, the directional intra-prediction mode, which is twice that of VVC, covers a range of directions from "bottom left to top right" to "top right to bottom left." Therefore, please note that the index for the directional intra-prediction mode belongs to [|2,130|].
[0190] Figure 20(a) presents an embodiment in which only the upper portion (2002) of the template for the current block (2000) is available. The set of reference samples (2003) of the template during the TIMD derivation step is completed with the reference samples colored in black. Since these reference samples are not available, they are generated by padding from the reference samples (2010).
[0191] Figure 21(a) presents an embodiment in which only the left portion (2101) of the template for the current block (2100) is available. The set of reference samples (2103) of the template during the TIMD derivation step is completed with the reference samples colored in black. Since these reference samples are not available, they are generated by padding from the reference samples (2110).
[0192] As shown in Figures 20(a) and 21(a), during the TIMD derivation step, the design of the reference sample template for the current block in the ECM, and the design of the reference sample template for the current block in the modified embodiments described in these figures, correspond to the same design when only one of the two template parts is available.
[0193] Figures 18(b.1) and 18(b.2), 19(b.1) and 19(b.2), 20(b), and 21(b) show that when the TIMD derivation step returns primary and secondary TIMD modes, during the prediction of the current block (1800, 1900, 2000, 2100), the reference samples used to predict the current block share the same design as the ECM and in the embodiments described above.
[0194] Regarding (1803) in FIG. 18(a.1), (1810) and (1811) in FIG. 18(a.2), (1903) in FIG. 19(a.1), (1910) and (1911) in FIG. 19(a.2), (2003) in FIG. 20, and (2103) in FIG. 21, it should be noted that the shown relationship between the pair {size of the current block, size of its template} and the rightward extension of the set of reference samples of the template can be adapted depending on the evolution of TIMD. Similarly, the shown relationship between the pair {size of the current block, size of its template} and the downward extension of the set of reference samples of the template can also be adapted depending on the evolution of TIMD.
[0195] For example, from ECM-7.0 to ECM-8.0, the set of reference samples of the template is extended 4 times to the right and downward because TIMD tests a wider-angle intra prediction mode from ECM-8.0. This extension can be directly applied to this modified embodiment.
[0196] In a variation of the embodiment of TIMD having the adapted set of reference samples of the template described above in relation to FIGS. 18 to 21, TIMD follows the same rules that define the wider-angle intra prediction mode as TIMD used in ECM. More specifically, in ECM, for a given width×height block using TIMD, during the derivation step of TIMD, for a given intra prediction mode to be tested on the template of this block, the potential conversion of this intra prediction mode to its wider-angle version depends exclusively on width and height. Similarly, in this variation of the embodiment provided herein, for TIMD using the adapted set of reference samples of the template, for a width×height block using TIMD, during the derivation step of TIMD, the potential conversion of the intra prediction mode to its wider-angle version follows the same rules based on the width and height of the current block.
[0197] Figure 22(a.1) presents an example of TIMD used in ECM for a given width×height block (2200) during the derivation step of TIMD. In this example, the prediction of the template, which includes the left part (2201) of the template of iTw×height and the upper part (2202) of the template of width×iTh from a set of templates (2203) of 2(width+iTw)+2(height+iTh)+1 reference samples of the template, is performed via the intra prediction mode of index 12. In this example, since width = 8 and height = 4, before the template prediction, the intra prediction mode of index 12 is converted to the wide-angle mode of index 141 according to the wide-angle conversion rule.
[0198] Figure 22(a.2) illustrates an example of TIMD provided in one embodiment herein for a given width×height block 2200 using TIMD with an adapted set of reference samples of the template during the derivation step of TIMD. In this example, the prediction from that set of reference samples (2210) of the left template part (2201) of iTw×height and the prediction from that set of reference samples (2211) of the upper template part (2202) of width×iTh are performed via the intra prediction mode of index 12. Here too, since width = 8 and height = 4, before the prediction of these two template parts (2201 and 2202), the intra prediction mode of index 12 is converted to the wide-angle mode of index 141. In Figures 22(a.1) and 22(a.2), the black dotted arrows indicate the direction of the wide-angle mode of index 141.
[0199] To illustrate with another example, Figure 23(a.1) is a copy of Figure 22(a.1), and Figure 23(a.2) is a copy of Figure 22(a.2), except that the intra-prediction mode at index 13 is replaced by the intra-prediction mode at index 12. In Figures 23(a.1) and 23(a.2), the intra-prediction mode at index 13 does not undergo wide-angle transformation because width=8 and height=4.
[0200] In the variant of the above embodiment, during the TIMD derivation step, for a given block using TIMD with an adapted set of template reference samples described herein, each of the two sets of reference samples of the template portion is extended to the right and / or bottom so that the prediction of each of the two template portions is feasible. In other words, the set of reference samples is extended as needed, depending on the intra-prediction mode being tested. If the extended portion of the set of reference samples of the template portion contains unavailable pixels, padding, such as in VVC / ECM, is used to fill the extended portion.
[0201] For example, in Figures 22(a.2) and 23(a.2), the set of reference samples (2210) includes 2 widths of reference samples in the upper and upper right of the left portion of the template (2201), ensuring that predictions of the left portion of the template (2201) are always possible in the embodiments described above. Similarly, the set of reference samples (2211) includes 2 heights of reference samples in the upper left and lower left of the template (2202), ensuring that predictions of the upper portion of the template (2202) are always possible in the embodiments described above.
[0202] Note that the extension examples in Figures 22(a.2) and 23(a.2) depend on the TIMD template shape defined in the ECM. If width ≤ 8, iTw = 2; otherwise, iTw = 4. This means that if height ≤ 8, iTh = 2; otherwise, iTh = 4.
[0203] For example, if the TIMD template shape is changed, other extensions are possible to ensure that the predictions from (2210) to (2201) and from (2211) to (2202) are always feasible for the tested intra-prediction mode.
[0204] In the modified embodiment, instead of reusing wide-angle rules as used in TIMD in ECM, the modifications provided herein are based on the size of the template portion to be predicted, with the wide-angle rules being modified.
[0205] Figure 24(a.1) shows an example of the TIMD derivation steps used in ECM for a given width × height block (2400). In this example, the prediction of templates (2401, 2402) for block (2400) from a set of template reference samples (2403) is performed via the intra-prediction mode at index 12, as shown in Figure 22(a.1). Since width=8 and height=4, before template prediction, the intra-prediction mode at index 12 is converted to the wide-angle mode at index 141 according to the wide-angle conversion rules defined in ECM.
[0206] Figure 24(a.2) shows an example of the TIMD derivation step in one embodiment using an adapted set of reference samples for a given width×height block (2400). In this example, the prediction of the reference samples for the left template portion (2401) of iTw×height (2401) and the prediction of the reference samples for the upper template portion (2402) of width×iTh (2411) are performed via the intra-prediction mode at index 12.
[0207] According to this embodiment, since iTw=2 and height=4, the wide-angle transformation is not applied to the intra-prediction mode at index 12 before the prediction at (2401). However, since width=8 and iTh=2, the intra-prediction mode at index 12 is transformed to the wide-angle mode at index 141 before the prediction at (2402). Therefore, in this embodiment, the wide-angle transformation rule is applied based on the size of the template portion to be predicted.
[0208] To illustrate with another example, Figure 25(a.1) is a copy of Figure 24(a.1), and Figure 25(a.2) is a copy of Figure 24(a.2), with the intra-prediction mode at index 130 in Figure 25 replacing the intra-prediction mode at index 12 used in Figure 24. In Figure 25(a.1), since width=8 and height=4, the intra-prediction mode at index 130 does not undergo wide-angle transformation. In Figure 25(a.2), since iTw=2 and height=4, the intra-prediction mode at index 130 is transformed into the wide-angle mode at index 1 before predicting the left portion of the template of the current block. Since width=8 and iTh=2, no wide-angle transformation is applied to the intra-prediction mode at index 130 before predicting the top portion of the template of the current block.
[0209] In Figures 24(a.2) and 25(a.2), note that the extensions shown to the right and below the set of reference samples for each of the two parts of the block template are simple examples that function for prediction via any intra-prediction mode within this modified embodiment. These shown extensions can be modified without affecting the purpose of the current modified embodiment, namely the wide-angle rule based on the size of the template part to be predicted.
[0210] In ECM, the wide-angle rule does not accurately translate the directional intra-prediction mode to the opposite-direction intra-prediction mode. To correct this, in a modified embodiment, the wide-angle rule relies on the size of the portion of the template to be predicted, and this wide-angle rule always accurately translates the directional intra-prediction mode to the opposite-direction intra-prediction mode.
[0211] Figure 26 adapts Figure 25 to this modified embodiment. Unlike Figure 25(a.2), in Figure 26(a.2), since iTw=2 and height=4, the intra-prediction mode at index 130 is converted to the intra-prediction mode at index 2 before predicting the left portion of the current block template. This is because the intra-prediction mode at index 130 and the intra-prediction mode at index 2 are in exactly opposite directions. In Figure 26(a.2), for the prediction of the left portion of the current block template, the gray dotted arrow indicates the direction of the intra-prediction mode at index 130. The black dotted arrow indicates the direction of the intra-prediction mode at index 2.
[0212] Any of the above embodiments relating to TIMD having an adapted set of template reference samples can be integrated into the process of a TIMD derivation process implemented in a video codec, for example, an ECM.
[0213] For example, Figure 27 shows an example of a method 2700 for encoding video blocks using TIMD according to one of the embodiments described herein. In 2701, the set of reference samples for each part of the template is determined according to one of the embodiments described herein in relation to Figures 18 to 26. Depending on the above-described variant being used, the set of reference samples may extend to the right and / or lower left of the left or upper part of the template, as described in relation to Figures 22 to 23.
[0214] At 2702, one or more intra prediction modes are derived based on the template of the block using the intra prediction mode derivation process of TIMD with a set of reference samples for the left and upper template parts determined at 2701. Depending on the above-described deformation mode being used, the intra prediction modes evaluated in this derivation process can undergo a wide-angle transformation as necessary, as described in relation to FIGS. 22 to 26.
[0215] At 2703, the video block is encoded based on the prediction obtained from one or more intra prediction modes obtained at 2702.
[0216] FIG. 28 shows an example of a method 2800 for decoding a video block using TIMD according to any one of the embodiments described herein. At 2801, a set of reference samples for each part of the template is determined according to any one of the embodiments described herein in relation to FIGS. 18 to 26. Assume that the same embodiment used in the encoder is used on the decoder side. At 2802, one or more intra prediction modes are derived based on the template of the block using the intra prediction mode derivation process of TIMD with a set of reference samples for the left and upper template parts determined at 2801. The derivation process is similar to that performed on the encoder side. At 2803, the video block is reconstructed based on the prediction obtained from one or more intra prediction modes obtained at 2802.
[0217] FIG. 29 shows an example of a workflow of the derivation (2900) step of TIMD using an adapted set of reference samples of the template according to the embodiment shown in FIG. 18a.2 or FIG. 19a.2. The following steps are performed on the current block 1800 to be encoded or decoded.
[0218] In step 2901, the set of intra-prediction modes to be evaluated is determined. For example, the set of intra-prediction modes can be obtained from the list of Most Probable Modes (MPMs) of the current width×height block (1800). This list can be supplemented with the indices of the intra-prediction modes DC_IDX (for DC modes), VER_IDX, and HOR_IDX (for vertical and horizontal modes) if they do not yet appear in the list. In step 2902, the set of reference samples for each part of the template (left and top) is determined. The left part of the template with size iTw×height (1801) uses the set of reference samples (1810). The top part of the template with size width×iTh (1802) uses the set of reference samples (1811). These sets of reference samples are extracted from the current channel, for example, the luminance channel for TIMD used by the current luminance block.
[0219] In step 2903, a loop is performed over the set of intra-prediction modes to be evaluated on the template. For each intra-prediction mode index i in the list collected in (2901), the process proceeds to steps (2904 and 2906) for the left part of the template and to steps (2905 and 2907) for the upper part of the template.
[0220] In 2904, the prediction P for the left part of the template (1801) l,i However, it is determined from the set of reference samples (1810) via the mode index i. In 2906, the left part of the template (1801) and the predicted P l,i SATDsatd between l,i This is calculated.
[0221] In 2905, the prediction P for the upper part of the template (1802) a,iHowever, it is determined from the set of reference samples (1811) via the mode index i. In 2907, the upper part of the template (1802) and the predicted P a,i SATDsatd between a,i This is calculated.
[0222] Once all intra-prediction modes for a set have been evaluated, in (2908), one or more intra-prediction mode indices are determined based on the costs evaluated in 2906 and 2907. For example, set
[0223]
number
[0224] In addition, two blending weights w primary and w secondary However, each
[0225]
number
[0226] The prediction step for the current block 1800 (step b.2 in Figure 18) is used to obtain two predictions for the current block (1800). primary and i secondary This is done using. The prediction for the current block (1800) is obtained using the reference sample (1804) defined for the current block (1800). These two predictions are w primary and w secondary This is blended using the current block (1800) to arrive at the final prediction.
[0227] Figure 30 shows another example of the TIMD derivation (3000) step workflow using an adapted set of template reference samples for a variation of the embodiment shown in Figure 24(a.2). The workflow is not limited to this embodiment, and a similar workflow can be applied to other embodiments shown in Figures 18 to 26.
[0228] In (3001), the set of intra prediction modes to be evaluated by TIMD is obtained for the current block to be predicted, which is a block (2400) with size width × height.
[0229] For example, the set of intra-prediction modes can be obtained from the list of most probable modes (MPM) for the current width × height block (2400). This list can be supplemented with the indices of intra-prediction modes DC_IDX, VER_IDX, and HOR_IDX if they do not already appear in the list.
[0230] In (3002), a set of reference samples (2410) for the left portion (2401) of the iTw×height of the template of block (2400), and a set of reference samples (2411) for the upper portion (2402) of the width×iTh of the template of block (2400) are extracted from the current channel.
[0231] In 3003, a loop is performed over the set of intra-prediction modes to be evaluated on the template. For each intra-prediction mode index i in the list collected in (3001), the process proceeds to steps (3014, 3004, and 3006) for the left part of the template and to steps (3015, 3005, and 3007) for the upper part of the template.
[0232] In 3014, parameters for predicting the left part (2401) of the template are derived. Depending on the transform and intra prediction mode angles used, it is determined whether the intra prediction mode index i has to be converted to its associated wide angle mode index i wide-l depending on iTw and height. If so, the intra prediction mode index i is converted to its associated wide angle mode index i wide-l .
[0233] If workflow 3000 is implemented using the transforms shown in FIGS. 22-23, in 3014 the conversion is based on the width and height of the current block.
[0234] In 3004, the prediction P l,i of (2401) is determined from a set of reference samples (2410) via the mode index i, or index i wide-l if converted in 3014. In 3006, the SATD satd l,i between (2401) and P l,i is calculated.
[0235] In 3015, parameters for predicting the top of the template (2402) are derived, which may include conversion of its associated wide angle mode index i of i depending on width and iTh as necessary. If workflow 3000 is implemented using the transforms shown in FIGS. 22-23, in 3015 the conversion is based on the width and height of the current block. wide-a
[0236] In 3005, the prediction P a,i of (2402) is calculated from that set of reference samples (2411) via the mode index i, or index i wide-a if converted.
[0237] In 3007, between (2402) and P a,i The SATD between it and satd a,i is calculated.
[0238] Once all the intra prediction modes of the set have been evaluated, at 3008, one or more intra prediction mode indices are determined based on the costs evaluated at 3006 and 3007. For example, in the set
[0239]
Number
[0240] In addition, the two blending weights w primary and w secondary are, respectively
[0241]
Number
[0242] During the prediction step of the current block (2400), i primary and i secondary are used to obtain two predictions of the current block (2400). These two predictions are blended using w primary and w secondary to yield the final prediction of the current block (2400).
[0243] In FIGS. 29 and 30, the order of the steps within the TIMD derivation workflow 2900 or 3000 using the adapted set of reference samples of the template is for illustrative purposes only. Some steps can be exchanged without affecting the TIMD derivation process. For example, in FIG. 29, (2901) and (2902) can be interchanged.
[0244] In ECM, such as ECM-8.0, the TIMD derivation step is part of many template-based coding tools. For example, the TIMD derivation step is performed in Intra Block Copy (IBC), Geometric Partition Mode (GPM), and Combined Intra Inter Prediction (CIIP).
[0245] In modified embodiments, the usual TIMD derivation step in the ECM is replaced by a TIMD derivation step using an adapted set of template reference samples, as described in one of the embodiments provided herein, for one or more template-based coding tools that include a TIMD derivation step.
[0246] In another modified embodiment, the usual TIMD derivation step in the ECM is replaced by a TIMD derivation step using an adapted set of template reference samples, as described in one of the embodiments provided herein, for all template-based coding tools involved in the TIMD derivation step.
[0247] In a modified example, any one of the embodiments provided herein in relation to the adaptation of a set of template reference samples can be combined with any one of the embodiments provided herein in relation to a PDPC tool. Since the PDPC is used in the TIMD derivation and prediction process, any one of the embodiments described herein with respect to the PDPC can replace the PDPC used in the TIMD process.
[0248] For example, in the embodiments described in relation to Figures 8-11, 16, and 17, the predictor blocks acquired in 801 and 901 are predictor blocks acquired for the video block template in template-based intra-prediction mode derivation. In some variations, at least two secondary reference samples used in the PDPC process to modify the predictor blocks are located in a reference array determined for the template, the reference array includes one or more reconstructed samples located in the row immediately above the template, or in the column immediately to the left of the template, or both. In some variations, in the PDPC process, the template includes a first part located above the video block and a second part located to the left of the video block, and at least two secondary reference samples used to modify the predictor blocks acquired for either the first or second part are located in the first part and the other part of the second part.
[0249] Figure 12 shows a block diagram of a system in which an aspect of this embodiment may be implemented according to another embodiment. Figure 12 shows one embodiment of a device 1200 for encoding or decoding video according to any one of the embodiments described herein. The device comprises a processor 1210 which can be interconnected to a memory 1220 via at least one port. Both the processor 1210 and the memory 1220 may also have one or more additional interconnections to external connections.
[0250] The processor 1220 is also configured to use any one of the embodiments described herein to acquire a predictor block for a video block based on an angle intra-prediction mode using at least one primary reference sample, modify the predictor block using a position-dependent pixel combination that uses a weighted combination of values determined from at least two secondary reference samples for at least one pixel, and encode or decode the video block based at least on the modified predictor block. For example, the processor 1220 is configured using a computer program product that includes code instructions to implement any one of the embodiments described herein.
[0251] In the embodiment shown in Figure 13, in a transmission context between two remote devices A and B via a communication network NET, device A comprises a processor associated with memory RAM and ROM configured to perform a method for encoding video as described with respect to Figures 1 to 11 or Figures 15 to 30, and device B comprises a processor associated with memory RAM and ROM configured to perform a method for decoding video as described in relation to Figures 1 to 11 or Figures 15 to 30. For example, the network is a broadcast network and is adapted to broadcast / transmit coded video from device A to decoding devices including device B.
[0252] Figure 14 shows an example of the syntax of a signal transmitted via a packet-based transmission protocol. Each transmitted packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include video data from any one of the embodiments described above.
[0253] Various implementations involve decoding. As used in this application, “decoding” can encompass all or part of the processes performed on a received encoded sequence to produce, for example, a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by the decoder in the various implementations described in this application, such as entropy decoding a sequence of binary symbols to reconstruct image or video data.
[0254] As further examples, in one embodiment, “decoding” refers only to entropy decoding; in another embodiment, “decoding” refers only to differential decoding; in yet another embodiment, “decoding” refers to a combination of entropy decoding and differential decoding; and in yet another embodiment, “decoding” refers to the entire picture reconstruction process, including entropy decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to a broader decoding process will become clear from the context of the particular description and will be well understood by those skilled in the art.
[0255] Various implementations involve encoding. As with the above considerations regarding "decoding," "encoding" as used in this application may encompass all or part of the processes performed on the input video sequence to generate an encoded bitstream. In various embodiments, such processes typically include one or more processes performed by the encoder, e.g., segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by the encoder in the various implementations described in this application, e.g., determining resampling filter coefficients and resampling the decoded picture.
[0256] As further examples, in one embodiment, “encoding” refers only to entropy coding; in another embodiment, “encoding” refers only to differential coding; and in yet another embodiment, “encoding” refers to a combination of differential coding and entropy coding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or to a broader encoding process in general will become clear from the context of the particular description and will be well understood by those skilled in the art.
[0257] It should be noted that the syntactic elements used herein are descriptive terms; therefore, they do not preclude the use of other syntactic element names.
[0258] This disclosure describes various types of information that can be transmitted or stored, such as syntax. This information can be packaged or configured in various ways, including methods common in video standards, such as including the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or slice headers), or SEI messages. Other methods are also available, including methods common to system-level or application-level standards, such as including the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for the purpose of session announcement and session invitation, such as being described in an RFC and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. For example, a DASH MPD (Media Presentation Description) descriptor used in DASH and transmitted over HTTP, where the descriptor is associated with a representation or set of representations to provide additional characteristics to the content representation. c. RTP header extensions, such as those used during RTP streaming. d. An ISO-based media file format that uses boxes, which are object-oriented building blocks defined by a unique type identifier and length, as used in OMAF, for example, and also known as "atoms" in some specifications. e. An HLS (HTTP Live Streaming) manifest sent via HTTP. The manifest can, for example, be associated with a version of content or a set of versions of content, and can provide characteristics of the version or set of versions.
[0259] When a diagram is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0260] Several embodiments refer to rate-distortion optimization. In particular, during the coding process, a balance or trade-off between rate and distortion is usually considered, often given the constraints of computational complexity. Rate-distortion optimization is usually formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are different methods for solving rate-distortion optimization problems. For example, these methods may be based on extensive testing of all coding options, including all considered mode or coding parameter values, but with a complete evaluation of their coding costs, as well as the associated distortions of the reconstructed signals after coding and decoding. In particular, faster approaches can be used to reduce coding complexity by calculating approximate distortion based on the predicted or predicted residual signal rather than the reconstructed signal. These two approaches can also be combined, for example, by using approximate distortion for only some of the possible coding options and full distortion for others. Other approaches evaluate only a subset of the possible coding options. More generally, many methods employ one of various techniques to perform optimization, but the optimization is not necessarily a complete evaluation of both coding costs and associated distortions.
[0261] The implementations and embodiments described herein can be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if considered only in the context of a single form of implementation (e.g., considered only as a method), the implementations of the considered features can also be realized in other forms (e.g., apparatus or programs). Apparatus may be implemented, for example, in appropriate hardware, software, and firmware. These methods may be implemented, for example, in a processor, which refers to processing devices in general, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.
[0262] The terms "one embodiment" or "one embodiment," or "one implementation" or "one implementation," and any other variations thereof, mean that the specific features, structures, characteristics, etc., described in relation to the embodiments are included in at least one embodiment. Therefore, the appearance of phrases such as "in one embodiment" or "in a particular embodiment," or "in one implementation" or "in a particular implementation," and any other variations, found in various places throughout this application, do not necessarily all refer to the same embodiment.
[0263] Additionally, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.
[0264] Furthermore, this application may also refer to “accessing” various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0265] Additionally, this application may refer to “receiving” various types of information. Receiving is intended to be a broad term, similar to “accessing.” Receiving information may include, for example, accessing information or retrieving information (for example, from memory). Furthermore, “receiving” typically accompanies, in some way, operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0266] For example, in the cases of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" should be understood as intended to cover the selection of only the first option (A), only the second option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such phrasing is intended to cover the selection of only the first option (A), only the second option (B), only the third option (C), only the first and second options (A and B), only the first and third options (A and C), only the second and third options (B and C), or all three options (A, B, and C). This may be extended as many times as there are listed items, as will be obvious to those skilled in the art.
[0267] Furthermore, as used herein, the word “signaling” refers, in particular, to pointing something to a corresponding decoder. In this way, in one embodiment, the same parameter is used on both the encoder and decoder sides. For example, the encoder can transmit a specific parameter to the decoder (explicit signaling), and as a result, the decoder can use the same specific parameter. Conversely, if the decoder already has a specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling), allowing the decoder to easily recognize and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual function. It should be seen that signaling can be achieved in various ways. For example, one or more syntactic elements, flags, etc., can be used to signal information to the corresponding decoder in various embodiments. While this concerns the verb form of the word “signaling,” the word “signaling” can also be used as a noun in this specification.
[0268] As will be apparent to those skilled in the art, the implementations can generate various signals, for example, that are formatted to carry information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the implementations described. For example, a signal may be formatted to carry a bitstream of the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over various different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0269] Several embodiments have been described above. Features of these embodiments may be provided individually or in any combination across various claim categories and types.
Claims
1. It is a method, A predictor block for a video block is acquired based on an angle-intra prediction mode that uses at least one first primary reference sample for at least one pixel of the predictor block. Modify the predictor block using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the predictor block. A method comprising encoding the video block based at least on the modified predictor block.
2. A device comprising one or more processors, wherein the one or more processors A predictor block for a video block is acquired based on an angle-intra prediction mode that uses at least one first primary reference sample for at least one pixel of the predictor block. For at least one pixel of the predictor block, the predictor block is modified using a weighted combination of values determined from at least two secondary reference samples. An apparatus capable of encoding the video block based at least on the modified predictor block.
3. It is a method, A predictor block for a video block is acquired based on an angle-intra prediction mode that uses at least one first primary reference sample for at least one pixel of the predictor block. Modify the predictor block using a weighted combination of values determined from at least two secondary reference samples for at least one pixel of the predictor block. A method comprising decoding the video block based at least on the modified predictor block.
4. A device comprising one or more processors, wherein the one or more processors A predictor block for a video block is acquired based on an angle intra-prediction mode that uses at least one first primary reference sample for at least one pixel of the predictor block. For at least one pixel of the predictor block, the predictor block is modified using a weighted combination of values determined from at least two secondary reference samples. A device capable of decoding the video block based at least on the modified predictor block.
5. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the weighted combination is a position-dependent pixel combination.
6. The method according to any one of claims 1, 3, or 5, or the apparatus according to any one of claims 2, 4, or 5, wherein at least one of the two secondary reference samples is acquired by extending the angle intra-prediction mode toward the secondary reference array.
7. The method according to any one of claims 1, 3, or 5-6, wherein at least one of the two secondary reference samples is located in the same column or row of the secondary reference array as the at least one pixel, depending on whether the intra-prediction mode is vertical or horizontal.
8. The method according to any one of claims 1, 3, or 5-7, wherein the weighted combination uses at least one second primary reference sample in a primary reference array used to acquire the predictor block, and the at least one second primary reference sample is acquired as a predictor of one of the at least two secondary reference samples using the angular intra-prediction mode, or the apparatus according to any one of claims 2, 4, or 5-7.
9. The method according to any one of claims 1, 3, or 5-8, wherein for at least one pixel, the weight associated with the at least two secondary reference samples is different from zero, or the apparatus according to any one of claims 2, 4, or 5-8.
10. The aforementioned weighted combination is A first term relating to a position-dependent pixel combination using a first secondary reference sample determined based on the intra prediction mode, A method according to claim 1 or 3, or the apparatus according to claim 2 or 4, comprising a second term relating to a gradient position-dependent pixel combination, using a second secondary reference sample located in the same row or column as the target pixel, and a primary reference sample determined for the second secondary reference sample.
11. The method or apparatus according to claim 10, wherein modifying the predictor block includes modifying the predicted value for the at least one pixel using a weighted combination of the first term and the second term, wherein the first weight associated with the first term and the second weight associated with the second term are non-zero.
12. The method according to any one of claims 1, 3, or 5 to 11, wherein determining the weighted combination includes determining at least one difference between one of the two secondary reference samples and the predicted value of the at least one pixel.
13. The method or apparatus according to claim 8, further comprising determining the weighted combination by determining the difference between one of the at least two secondary reference samples and the second primary reference sample.
14. The method according to any one of claims 1, 3, or 5 to 13, wherein the weights used in the weighted combination depend on the angle of the prediction direction of the intra prediction mode, the block size of the video block, or both the prediction direction and the block size, or the apparatus according to any one of claims 2, 4, or 5 to 13.
15. The method according to any one of claims 1, 3, or 5-14, or the apparatus according to any one of claims 2, 4, or 5-14, wherein, in response to a determination that the scale value is negative, the scale value is determined from the height or width of the video block and the angle of the prediction direction of the intra-prediction mode, and determining the weighted combination includes using a gradient determined for a first pixel of the video block that is located in the same row as the at least one pixel and in a first column of the video block, or located in the same column as the at least one pixel and in a first row of the video block.
16. The method according to any one of claims 1, 3, or 5-14, wherein at least one of the two secondary reference samples is obtained using the same interpolation filter as the interpolation filter used to determine the first primary reference sample, or the apparatus according to any one of claims 2, 4, or 5-14.
17. The method according to any one of claims 1, 3, or 5 to 15, wherein modifying the predictor block based on the weighted combination is in response to an indicator transmitted with the video block, or the apparatus according to any one of claims 2, 4, or 5 to 145.
18. The method according to any one of claims 1, 3, or 5 to 15, wherein modifying the predictor block based on the weighted combination is in response to a cost determination determined for the video block template, or the apparatus according to any one of claims 2, 4, or 5 to 145.
19. The method according to any one of claims 1, 3, or 5-17, wherein at least one of the two secondary reference samples is determined using one of a linear interpolation filter, a four-tap filter, a six-tap cubic filter, or a smoothing filter, or the apparatus according to any one of claims 2, 4, or 5-17.
20. The method according to any one of claims 1, 3, or 5 to 18, wherein the predictor block is modified with respect to the lumens and chromens components, or the apparatus according to any one of claims 2, 4, or 5 to 18.
21. The method according to any one of claims 1, 3, or 5-20, or the apparatus according to any one of claims 2, 4, or 5-20, wherein when the distance of one of the at least two secondary reference samples to the origin of the video block exceeds a given value, one of the at least two secondary reference samples is not used to correct the predicted value of the at least one pixel in the predictor block.
22. The method according to any one of claims 1, 3, or 5-20, wherein when one of the at least two secondary reference samples is unavailable, the one of the at least two secondary reference samples is not used to correct the predicted value of the at least one pixel in the predictor block, or the apparatus according to any one of claims 2, 4, or 5-20.
23. The method according to any one of claims 1, 3, or 5-20, or the apparatus according to any one of claims 2, 4, or 5-20, wherein when one of the at least two secondary reference samples is unavailable and the distance of one of the at least two secondary reference samples to a reference sample used to pad one of the at least two secondary reference samples exceeds a given value, one of the at least two secondary reference samples is not used to correct the predicted value of the at least one pixel in the predictor block.
24. The method according to any one of claims 1, 3, or 5-20, wherein the predictor block is not modified in response to a determination that a neighboring block of the video block to which one of the at least two secondary reference samples belongs is unavailable, or the apparatus according to any one of claims 2, 4, or 5-20.
25. A computer program product comprising instructions for causing one or more processors to perform the method according to any one of claims 1, 3, or 5 to 24.
26. A non-temporary computer-readable medium for storing executable program instructions, which cause a computer executing the program instructions to carry out the method according to any one of claims 1, 3, or 5 to 24.
27. A bitstream comprising data representing video encoded using the method described in any one of claims 1, 3, or 5 to 24.
28. A non-temporary computer-readable medium for storing the bitstream described in claim 27.
29. It is a device, The apparatus according to any one of claims 4 to 24, A device comprising: (i) an antenna configured to receive or transmit a signal including data representing the video block; (ii) a band limiter configured to restrict the signal to a frequency band including the data representing the video block; or (iii) a display configured to display the video block.
30. The device according to claim 29, wherein the device includes at least one of a television, a mobile phone, a tablet, and a set-top box.