Methods and apparatus for encoding and decoding images or videos using combined intra modes

The combination of non-directional and directional intra-prediction modes in video coding enhances compression efficiency by addressing spatial and temporal redundancies, resulting in improved decoding performance and reduced bit rates.

JP2025524401APending Publication Date: 2025-07-30INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
JP2024573395
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-22
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing video coding schemes face challenges in achieving high compression efficiency due to limitations in intra-prediction modes, particularly in handling spatial and temporal redundancies within video content.

Method used

A method and apparatus for encoding and decoding video blocks using a combination of non-directional and directional intra-prediction modes, including weighted averages and template-based derivations, to enhance prediction accuracy and reduce signaling overhead.

Benefits of technology

Improves video compression efficiency by optimizing intra-prediction through combined modes, leading to reduced bit rates and enhanced decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524401000001_ABST
    Figure 2025524401000001_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding or decoding blocks of an image or video are disclosed. For at least one block of the image or video, a first intra prediction mode is obtained and a second intra prediction mode is obtained. The at least one block is encoded or decoded based on a combination of the first intra prediction mode and the second intra prediction mode. In one embodiment, the first intra prediction mode is a non-directional based intra prediction mode and the second intra prediction mode is a directional intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to Related Applications) This application claims priority to European Patent Application No. 22305953.6, filed on June 30, 2022, which is hereby incorporated by reference in its entirety.

[0002] This embodiment generally relates to video compression. This embodiment relates to a method and apparatus for encoding and decoding blocks of an image or video based on a combination of intra - prediction modes.

Background Art

[0003] To achieve high compression efficiency, video coding schemes typically employ prediction and transformation to exploit spatial redundancy and temporal redundancy within video content. Generally, intra - prediction or inter - prediction is used to utilize intra - picture or inter - picture correlation, and then the difference between the original block and the predicted block, often called the prediction error or prediction residual, is transformed, quantized, and entropy - coded. To reconstruct the video, the compressed data is decoded by reverse processes corresponding to entropy - coding, quantization, transformation, and prediction.

Summary of the Invention

[0004] According to one aspect, a method for encoding at least one block of an image or video is provided. The method includes: obtaining a first intra - prediction mode for at least one block; obtaining a second intra - prediction mode for at least one block; encoding at least one block based on a combination of the first intra - prediction mode and the second intra - prediction mode.

[0005] According to one aspect, a method for decoding at least one block of an image or video is provided. The method includes, for at least one block, obtaining a first intra prediction mode, obtaining a second intra prediction mode, and decoding the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode.

[0006] According to another aspect, an apparatus for encoding at least one block of an image or video is provided. The apparatus includes one or more processors, and the one or more processors are operable to obtain a first intra prediction mode, obtain a second intra prediction mode, and encode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode for at least one block of the image or video.

[0007] According to another aspect, an apparatus for decoding at least one block of an image or video is provided. The apparatus includes one or more processors, and the one or more processors are operable to obtain a first intra prediction mode, obtain a second intra prediction mode, and decode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode for at least one block of the image or video.

[0008] According to another aspect, a method for encoding at least one block of an image or video is provided. The method includes, for at least one block, obtaining a first intra prediction mode, obtaining the second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode, and encoding the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode.

[0009] According to another aspect, a method for decoding at least one block of an image or video is provided. The method includes, for at least one block, obtaining a first intra prediction mode; in response to a determination that the first intra prediction mode should be combined with a second intra prediction mode, obtaining the second intra prediction mode; and decoding the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode.

[0010] According to another aspect, an apparatus for encoding at least one block of an image or video is provided. The apparatus includes one or more processors, and the one or more processors are operable to obtain a first intra prediction mode for at least one block of the image or video, obtain a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode, and encode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode.

[0011] According to another aspect, an apparatus for decoding at least one block of an image or video is provided. The apparatus includes one or more processors, and the one or more processors are operable to obtain a first intra prediction mode for at least one block of the image or video, obtain a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode, and decode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode.

[0012] In some embodiments, the first intra prediction mode is obtained from a first set of intra prediction modes, the second intra prediction mode is obtained from a second set of intra prediction modes, and the first and second sets of intra prediction modes are distinct.

[0013] In some embodiments, the first set of intra prediction modes includes non - directional intra prediction modes, and the second set of intra prediction modes includes directional intra prediction modes. One of the first intra prediction mode and the second intra prediction mode is obtained from the first set of intra prediction modes, and the other of the first intra prediction mode and the second intra prediction mode is obtained from the second set of intra prediction modes.

[0014] In some embodiments, the first intra prediction mode is one of a planar prediction mode, a DC prediction mode, an intra block copy prediction mode, and a matrix - based intra prediction mode. In some embodiments, the second intra prediction mode is one of the directional intra prediction modes.

[0015] In some embodiments, the second intra prediction mode is derived from reconstructed samples adjacent to at least one block. In one variant, the second intra prediction mode is obtained from at least one of decoder - side intra mode derivation or template - based intra mode derivation.

[0016] In some embodiments, the combination is a weighted average of the prediction obtained from the first intra prediction mode and the prediction obtained from the second intra prediction mode. In some variants, the weight used in the combination is an indicator signaled in the bitstream indicating a weight from the first intra prediction mode or a set of weights used for the prediction from the first intra prediction mode, or a cost obtained when determining the second intra prediction mode, or the rank of the second prediction mode in the list of intra prediction modes, or at least one of them. In another variant, the weight used in the combination is derived from the cost obtained when determining the second intra prediction mode. In another variant, the weight varies with the position of the samples within at least one block.

[0017] In one embodiment, one or more processors are operable to encode an image or video to which a block belongs. In one embodiment, one or more processors are operable to decode an image or video to which a block belongs. Further embodiments that can be used alone or in combination are described herein.

[0018] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform a method for encoding / decoding blocks of an image or video according to any of the embodiments described herein. One or more of these embodiments also provide a non-transitory computer-readable medium and / or a computer-readable storage medium having stored thereon instructions for encoding / decoding blocks of an image or video according to the method described herein. Instructions for encoding / decoding blocks of an image or video according to the method described herein are stored.

[0019] One or more embodiments also provide a computer-readable storage medium storing a bitstream generated according to the method described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the method described above.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19A

Figure 19B

Figure 20A

Figure 20B

Figure 20C

Figure 21A

Figure 21B

Figure 21C

Figure 21D

Figure 21E

Figure 22-1

Figure 22-2

Figure 23A

Figure 23B

Figure 24

Figure 25

Figure 26

Mode for Carrying Out the Invention

[0021] In this application, various aspects are described, including tools, features, embodiments, models, techniques, etc. Many of these aspects are specifically described and often described in a way that may seem restrictive in order to show at least individual characteristics. However, this is for the purpose of clarifying the description and does not limit the application or scope of those aspects. In fact, all of the different aspects can be combined and replaced to provide further aspects. Furthermore, these aspects can also be combined and replaced with the aspects described in previous applications.

[0022] The aspects described and contemplated in this application can be implemented in many different forms. The following FIGS. 1, 2, and 3 provide some embodiments, but other embodiments are contemplated, and the consideration of FIGS. 1, 2, and 3 does not limit the width of the implementation forms. At least one of the above aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These aspects, and other aspects, can be implemented as a computer-readable storage medium that internally stores instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium that stores within itself a bitstream generated according to any of the described methods.

[0023] In this application, the terms "reconstructed" and "decoded" can be used with the same meaning, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably.

[0024] Various methods are described herein, each of which includes one or more steps or acts for achieving the described method. Unless a particular order of steps or acts is required for proper operation of the method, the order and / or use of particular steps and / or acts may be modified or combined. Note that terms such as "first", "second", etc. may be used in various embodiments to modify elements, components, steps, acts, etc., such as, for example, "first decoding" and "second decoding". The use of such terms does not imply an ordering with respect to the modified act unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding and may occur, for example, before, during, or overlapping with the second decoding.

[0025] This aspect is not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations, whether existing or future developments, and to any extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.

[0026] FIG. 1 shows a block diagram of an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device that includes various components described below and is configured to execute one or more of the aspects described in this application. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers, among other electronic devices. The elements of System 100 may be embodied, alone or in combination, in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of System 100 are distributed across multiple ICs and / or separate components. In various embodiments, System 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or dedicated input ports and / or output ports. In various embodiments, System 100 is configured to implement one or more of the aspects described in this application.

[0027] System 100 includes, for example, at least one processor 110 configured to execute instructions loaded internally to implement various aspects described in this application. Processor 110 may include built-in memory, an input / output interface, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage device 140 may include, by way of non-limiting example, an internal storage device, a removable storage device, and / or a network-accessible storage device.

[0028] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, the device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.

[0029] To execute the various aspects described in this application, the program code loaded on the processor 110 or the encoder / decoder 130 may be stored in the storage device 140 and then loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstream, matrix, variable, and intermediate or final results from the processing of equations, expressions, operations, and operation logic.

[0030] In some embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide a working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as a working memory for video coding operations and decoding operations such as MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, H.265 and MPEG-H Part 2 are also known), or VVC (Versatile Video Coding, a new standard under development by the Joint Video Experts Team (JVET)).

[0031] Inputs to the elements of system 100 may be provided through various input devices as shown in block 105. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals transmitted over the airwaves by, for example, a broadcaster, (ii) a Component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other embodiments, not shown in FIG. 1, include composite video.

[0032] In various embodiments, the input device of block 105 has each associated input processing element known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in certain embodiments, band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments may include one or more elements that perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner that performs these various functions, for example, including down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, down-converting, and filtering again to a desired frequency band an RF signal transmitted over a wired (e.g., cable) medium. In various embodiments, the order of the above (and other) elements may be rearranged, some of these elements may be omitted, and / or other elements performing similar or different functions may be added. Adding elements may include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0033] In addition, the USB and / or HDMI terminals may each include an interface processor for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 110. Demodulation, error correction, and multiplexing of the separated streams are provided to various processing elements, including, for example, the processor 110, and an encoder / decoder 130 that operates in combination with memory and storage elements that process the data streams as needed to present them on the output devices.

[0034] The various elements of the system 100 may be provided within an integrated housing, in which the various elements are interconnected using an internal bus known in the art, including a suitable connection configuration 115, such as an I2C bus, wiring, and a printed circuit board, and may transmit data to each other.

[0035] The system 100 includes a communication interface 150 that enables it to communicate with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired medium and / or a wireless medium.

[0036] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of such embodiments are received via communication channel 190 and communication interface 150 that are adapted for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to an external network including the Internet to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that distributes data via the HDMI connection of input block 105 is used to provide the data streamed to system 100. In yet other embodiments, the RF connection of input block 105 is used to provide the data streamed to system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0037] System 100 may provide an output signal to various output devices, including display 165, speaker 175, and other peripheral devices 185. The display 165 in various embodiments may include, for example, one or more of a touch screen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be for a television, a tablet, a laptop, a mobile phone, or other device. The display 165 can also be integrated with other components (such as in the case of a smartphone) or separate (such as an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 185 include one or more of a stand-alone digital video disc (or digital versatile disc) (both terms referred to as DVR), a disc player, a stereo system, and / or an illumination system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of system 100. For example, a disc player performs the function of playing back the output of system 100.

[0038] In various embodiments, the control signal is communicated between the system 100 and the display 165, speaker 175, or other peripheral device 185 using signaling such as AV.Link, CEC, or other communication protocols that enable control between devices with or without user intervention. The output devices may be communicatively coupled to the system 100 via dedicated connections through their respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to the system 100 using the communication channel 190 via the communication interface 150. The display 165 and speaker 175 may be integrated into a single unit with other components of the system 100 in an electronic device, such as a television, for example. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0039] Alternatively, the display 165 and speaker 175 may be separated from one or more of the other components, for example, if the RF section of the input 105 is part of a separate set-top box. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0040] The embodiments can be implemented by computer software implemented by the processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120 can be of any type suitable for the technical environment, and can be implemented using any suitable data storage technology, such as, as a non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 110 can be of any type suitable for the technical environment, and can include, as a non-limiting example, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a multi-core architecture-based processor.

[0041] Figure 2 shows the encoder 200. Although variants of this encoder 200 are contemplated, hereinafter, for clarity, the encoder 200 will be described without describing all the variants that are expected.

[0042] In some embodiments, Figure 2 also shows an encoder that has improved the HEVC standard or the VVC standard, or an encoder that employs a technology similar to HEVC or VVC, such as an encoder ECM being developed by the JVET (Joint Video Exploration Team).

[0043] Before being encoded, the video sequence can go through pre-encoding processing (201), such as applying a color conversion to the input color picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or performing remapping of the input picture components (e.g., using histogram equalization of color components) or resizing of the picture (e.g., downscaling) to obtain a signal distribution that is more resistant to compression. Metadata can be associated with the pre-processing and attached to the bitstream.

[0044] In encoder 200, a picture is encoded by encoder elements as follows. The picture to be encoded is partitioned (202) into units such as, for example, coding units (CUs) or blocks, and processed. In the present disclosure, different expressions may be used to refer to such units or blocks resulting from the partitioning of a picture. Such expressions may be coding unit or CU, coding block or CB, luminance CB, or block... A coding tree unit (CTU) may refer to a group of blocks or a group of units. In some embodiments, a CTU may be regarded as a block or a unit by itself.

[0045] Each unit is encoded using, for example, either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction (260) is performed. In the inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) which of the intra mode or the inter mode to use to encode the unit, and indicates the intra or inter decision, for example, by a prediction mode flag. The encoder may also mix (263) the intra prediction result and the inter prediction result, or may mix the results from different intra / inter prediction methods. The prediction residual is calculated, for example, by subtracting (210) the block predicted from the original image block.

[0046] The motion improvement module (272) uses available reference pictures to improve the motion field of a block without referring to the original block. The motion field for a region can be regarded as a set of motion vectors for all pixels having that region. When the motion vectors are sub-block based, the motion field can also be represented as a set of all sub-block motion vectors within the region (all pixels within a sub-block have the same motion vector, and the motion vectors can be different for each sub-block). When a single motion vector is used for a region, the motion field for the region can also be represented by a single motion vector (the same motion vector for all pixels within the region).

[0047] Next, the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is directly coded without applying the transformation process or the quantization process.

[0048] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (240), inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (255) to reconstruct an image block. A loop filter (265) is applied to the reconstructed picture, for example, to perform deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (280).

[0049] Figure 3 shows a block diagram of video decoder 300. As described below, in decoder 300, the bitstream is decoded by decoder elements. Video decoder 300 generally implements a decoding path that is the reverse of the encoding path described in FIG. 2. Encoder 200 also generally performs video decoding as part of video data encoding.

[0050] In particular, the input to the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (335). The transform coefficients are inverse quantized (340), inverse transformed (350), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (355) to reconstruct the image block.

[0051] A predicted block can be obtained from intra prediction (360) or motion compensation prediction (i.e., inter prediction) (375) (370). The decoder may mix the intra prediction result and the inter prediction result (373), or may mix the results from multiple intra / inter prediction methods. Before motion compensation, the motion field can be improved by using available reference pictures (372). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (380).

[0052] The decoded picture can further undergo post - decoding processing (385), for example, inverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4), or inverse remapping that performs the inverse of the remapping process performed in pre - encoding processing (201), or size change (e.g., upscaling) of the reconstructed picture. The post - decoding processing can use metadata derived in the pre - encoding processing and signaled in the bitstream.

[0053] Some of the embodiments described herein relate to intra - prediction used in image or video coding. As an example, for a given block to be predicted, the encoder selects the best intra - prediction mode with respect to rate - distortion and signals the selected intra - prediction mode index to the decoder. In this way, for this block, the decoder can perform the same prediction. Signaling the index of the selected intra - prediction mode may add extra overhead and reduce the coding gain from the intra part of the coded video. An example of coding the index of the intra - prediction mode selected to predict a given block is to create a list of the Most Probable Modes (MPM), and thus, if the index of the selected intra - prediction mode belongs to that list, reduce the signaling overhead. This is a classical method for signaling the intra - prediction mode index, known as MPM - list - based signaling. This method is adopted, for example, in VVC and HEVC. This method is extended in ECM (Enhanced Compression Model) where two instead of one MPM lists are used. Hereinafter, for brevity, the MPM - list - based signaling used for signaling the mode index is abbreviated to mode signaling.

[0054] To further limit signaling overhead, the ECM also features two tools that derive from the decoded pixels surrounding a given block the intra prediction mode that is most likely to be the best intra prediction mode for predicting the given block with respect to rate distortion. For each of these two tools, the reduction in signaling overhead results from the fact that the signaling of the tool alone enables the decoder to obtain the index of the most likely "best" intra prediction mode.

[0055] The first tool is known as Decoder-side Intra Mode Derivation (DIMD). The second tool is called Template-based Intra Mode Derivation (TIMD). More specifically, in DIMD, templates of the decoded pixels above and to the left of the current block to be predicted are analyzed to infer the directionality of the template, from which two directional intra prediction modes are selected. The prediction signal is generated by mixing these two intra prediction modes with the planar mode. In TIMD, several intra prediction modes are tested against a template of the decoded pixels above and to the left of the current block. The cost (Sum of Absolute Transform Differences (SATD)) is determined for each of the tested intra prediction modes between the decoded samples of the template and the predicted samples of the template using the tested intra prediction mode. The two intra prediction modes that result in the two lowest costs are retained. The prediction signal is generated either by applying the intra prediction mode with the lowest SATD or by mixing the two intra prediction modes that provide the two lowest costs.

[0056] In VVC and ECM, there are other intra prediction tools such as Intra Block Copy (IBC) and Matrix-based Intra Prediction (MIP). In IBC, a reference block from the reconstructed part of the current frame is copied and used as the prediction of the current block. In MIP, to calculate the prediction of the current block, a matrix-vector multiplication is performed between the reference samples and a matrix selected from a set of available matrices.

[0057] In HEVC and VVC, there are also inter prediction tools that predict the current block from the reconstructed frame. Some of these tools use bi-prediction and perform a weighted average from two inter predictions to predict the current block. Some tools such as Combined Intra Inter Prediction (CIIP) also perform a weighted average between intra prediction and inter prediction.

[0058] Some embodiments of the present disclosure relate to a method for encoding or decoding blocks of an image or video, in which a weighted average between two intra prediction modes is used to predict a block, enabling an increase in the performance of intra compression in ECM and VVC.

[0059] Examples of intra prediction modes are described below with respect to FIGS. 4 to 18.

[0060] As used herein, the term "intra mode" or "intra prediction mode" may refer to any one of the intra prediction tools described herein, such as MIP, IBC, CIIP, or any one of the core 67 intra prediction modes known from HEVC, VVC, or ECM, or any other mode that generates a prediction of a unit from the reconstructed samples of the picture to which the unit belongs.

[0061] Core 67 Intra Prediction Mode: To capture any edge directions presented in the raw video, the number of directional intra prediction modes in VVC is extended from 33 used in HEVC to 65. Figure 4 shows the directional intra prediction modes in VVC, where the directional modes from HEVC are indicated by solid black arrows, and the new directional modes in VVC not in HEVC are shown as dashed black arrows. These denser directional intra prediction modes are applied to all block sizes and to both luma intra prediction and chroma intra prediction. In VVC, in addition to 65 directional intra prediction modes, a planar mode and a DC mode are provided, making the number of core intra prediction modes 67.

[0062] In the change from HEVC to VVC, the planar mode and the DC mode remain unchanged except for the following minor modifications. In HEVC, each of the intra-coded blocks has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate the intra predictor using DC. In VVC, a block can have a rectangular shape that generally requires the use of per-block division operations. To avoid the division operation for DC prediction, only the long side is used to calculate the average of non-square blocks.

[0063] In ECM, the core structure of the 67 intra prediction modes is inherited from the core structure in VVC. This core structure is improved in ECM, where the 4-tap interpolation for the directional intra prediction modes from VVC becomes 6-tap interpolation in ECM. The Position Dependent Intra Prediction Combination (PDPC) is supplemented by gradient-based PDPC in ECM.

[0064] Intra Prediction Mode Signaling in ECM Intra Prediction Mode Signaling in Luma In the ECM (currently ECM-4.0), when the intra prediction mode selected to predict the current luminance coding block (Coding Block, CB) is not any of DIMD, matrix-based intra prediction (MIP) mode, and TIMD, that is, when it is one of the 67 intra prediction modes described above, the MPM list of this CB is used and its index is signaled.

[0065] In the ECM (currently ECM-4.0), the general-purpose MPM list is decomposed into a list of 6 primary MPMs and a list of 16 secondary MPMs as shown in FIGS. 5A and 5B. The general-purpose MPM list is constructed by sequentially adding candidate intra prediction mode indexes from the candidate with the highest probability to the candidate with the lowest probability of being the selected intra prediction mode for predicting the current luminance CB. FIG. 5A shows the sequential addition of candidate intra prediction mode indexes when the current luminance CB to be predicted belongs to an intra slice and the current luminance CB has a width W and a height H, from left to right. The candidate intra prediction modes are inserted into the primary list or the secondary list in a specific order, and some intra prediction modes are inserted into one of the lists according to the conditions shown in FIGS. 5A and 5B. In the case of the primary list, the first MPM is the planar intra prediction mode, and then the subsequent intra prediction modes to be inserted are the intra prediction modes used to predict the adjacent CBs of the current luminance CB. In the case of the secondary list, two intra prediction modes obtained from DIMD are inserted, and then, until the size limit of the secondary list is reached, other intra prediction modes are inserted into the secondary list by considering the adjacent intra prediction modes of one or more intra prediction modes inserted into the primary list. If the size limit is not reached, some default modes may be added.

[0066] Note that there is no redundancy in the general list of MPMs, that is, the list cannot contain two identical intra prediction mode indices. For readability, FIGS. 5A and 5B show the case where each candidate intra prediction mode index is different from each other. However, in the general case, the slots of indices 0 to i-1 included in the general list of MPMs are already filled. If the current candidate intra prediction mode index already exists in the current general list of MPMs, this candidate is skipped, and if the next candidate intra prediction mode does not exist in the general list of MPMs, it is inserted in the slot of index i. Otherwise, the current intra prediction mode index is inserted in the slot of index i, and if the next candidate intra prediction mode does not exist in the general list of MPMs, it is inserted in the slot of index i+1.

[0067] The signaling of the intra prediction mode selected to predict the current luminance CB in the ECM (currently ECM-4.0) is shown in FIGS. 6A and 6B. Note that FIGS. 6A and 6B depict the signaling of the intra prediction mode selected to predict the current luminance CB on the encoder side. However, the same applies to the decoder side. In FIGS. 6A and 6B, MRL indicates Multiple Reference Line. When the TIMD flag is equal to 1, the MRL index belongs to {0,1,3}. An MRL index of 0 means that the MRL is not used to predict the current luminance CB. An MRL index of 1 means that the second row of the decoded reference samples above the current luminance CB and the second column of the decoded reference samples to the left of the current luminance CB are used for prediction. An MRL index of 3 means that the fourth row of the decoded reference samples above the current luminance CB and the fourth column of the decoded reference samples to the left of the current luminance CB are used for prediction. When the TIMD flag is equal to 0, the MRL index belongs to {0,1,3,5,7,12}. ISP represents Intra Sub-Partition. The ISP mode index belongs to {0,1,2}. An ISP mode index of 0 means that the ISP is not used for the current luminance CB. An ISP mode index of 1 indicates that the current luminance CB is horizontally divided into Transform Blocks (TBs). An ISP mode index of 2 indicates that the current luminance CB is vertically divided into luminance TBs. In this figure, the intra prediction modes BDPCM (Intra Block-DPCM), TMP (Template Matching based intra prediction), IBC (Intra Block Coding), and palette are omitted because these tools are exclusively turned on for a specific video sequence.

[0068] Intra Prediction Mode Signaling in Chrominance In ECM (currently ECM-4.0), the signaling of the intra prediction mode selected to predict the current pair of chrominance CB, i.e., the collocated Cb and Cr CB of the current luma CB, is shown in FIG. 7. In FIG. 7, when the Direct Mode (DM) flag is equal to 1, the four possibilities for the current intra prediction mode index are the index of the planar mode, the index of the horizontal mode, the index of the vertical mode, and the index of DC. To avoid any redundancy, when DM is one of the four above-mentioned modes, in this set of four modes, the index of the redundant mode is replaced by the index of the vertical diagonal mode. In ECM (currently ECM-4.0), the Cross-Component Linear Model (CCLM) collects six different intra prediction modes represented as LM, MMLM, MDLM_L, MDLM_T, MMLM_L, and MMLM_T, but it should be noted that in VVC, the CCLM collects only three intra prediction modes.

[0069] Wide Angle Intra Prediction (WAIP) In VVC and ECM, for non-square blocks, some of the conventional angular intra prediction modes are replaced by the wide angle mode. The replaced modes are signaled using the original method and remapped to the index of the wide angle mode after syntax analysis. The total number of core intra prediction modes remains unchanged, i.e., 67.

[0070] For the current WxH block to be predicted, FIG. 8 shows a set of decoding reference samples consisting of an array of upper decoding reference samples of length 2W+1 and an array of left decoding reference samples of length 2H+1. FIG. 8 also shows the relationship between the range of decoded reference samples around the current WxH block and the range of allowable intra prediction angles. Table 1 presents the indices of intra prediction modes that are replaced by the wide-angle mode in VVC and ECM, depending on the size of the current block to be predicted.

[0071]

Table 1

[0072] FIG. 9 shows an example of how the angular intra mode is replaced by the wide-angle mode for a non-square block whose width is strictly larger than its height. In this example, mode 2 is replaced by wide-angle mode 67. Mode 3 is replaced by wide-angle mode 68. For example, when the current block to be predicted is 8×4, this replacement process proceeds incrementally until mode 7 is replaced by wide-angle mode 72.

[0073] Template-Based Intra Mode Derivation (TIMD) For a given luminance CB (1003) in FIG. 10(a), the following mode derivation via TIMD is applied in the same way on both the encoder side and the decoder side. For each intra prediction mode in the MPM list of this luminance CB, if necessary, supplemented with the default mode, the TIMD mode determines the prediction of the templates (1000 and 1001) of the luminance CB (1003) from the decoded reference samples of the template (1002), and the SATD between the predicted reference samples of the template and the decoded reference samples of the template of the luminance CB is determined. In the first pass, two intra prediction modes with the minimum SATD are selected as the TIMD mode. This means that the set of possible intra prediction modes derived via TIMD gathers 131 modes. After retaining two intra prediction modes in the first pass involving the MPM list supplemented with the default mode, for each of these two selected intra prediction modes, if this mode is neither planar nor DC, TIMD also checks, with respect to the SATD cost, the two closest extended directional intra prediction modes for the selected intra prediction mode. In this second pass, the set of directional intra prediction modes is extended from 65 to 129 by inserting the direction between each black arrow and its adjacent dotted black arrow in FIG. 4 to provide extended directional intra prediction modes.

[0074] Note that in the above description, it is assumed that the templates of the luminance CB do not extend beyond the boundaries of the current frame. If at least a part of the template of the luminance CB extends beyond the boundaries of the current frame, FIGS. 10(b) and (c) show how the template is adapted.

[0075] In FIG. 10(a), the current W×H luminance CB (1003) is surrounded by its fully available template consisting of its left w t ×H part (1000) and its upper W×h t part (1001). During the TIMD derivation step, the tested intra prediction modes are those of the template with 1 + 2wt +2W + 2h t Predict the template of the current luminance CB from the set (1002) of decoded reference samples of +2H. In the current version of ECM (ECM - 4.0), when W ≤ 8, w t is equal to 2, otherwise w t is equal to 4. When H ≤ 8, h t is equal to 2, otherwise h t is equal to 4. In Figure 10(b), the current W×H luminance CB (1003) is surrounded by its template, and only the W×h t portion (1001) above it is available. During the TIMD derivation step, the tested intra prediction mode is 1 + 2W + 2h t + 2H of the set (1002) of decoded reference samples to predict the template of the current luminance CB. In Figure 10(c), the current W×H luminance CB (1003) is surrounded by its template, and only the w t ×H portion (1000) on its left is available. During the TIMD derivation step, the tested intra prediction mode is 1 + 2w t + 2W + 2H of the set (1002) of decoded reference samples to predict the template of the current luminance CB.

[0076] To predict the current luminance CB via TIMD, the two intra predictions obtained from the two TIMD modes selected in the first or second pass for the luminance CB are weighted and fused after applying PDPC. The weights used depend on the prediction SATD of the two TIMD modes.

[0077] In the case of TIMD, since the set of directional intra prediction modes is extended from 65 to 129, the intra prediction mode replacement in WAIP is applied. Table 1 is applied to Table 2. For example, in the case of a given 8×4 luminance CB using TIMD, mode 2 is replaced by wide-angle mode 131, mode 3 is replaced by wide-angle mode 132, mode 4 is replaced by wide-angle mode 133, …, mode 12 is replaced by wide-angle mode 141.

[0078]

Table 2

[0079] Matrix-based Intra Prediction (MIP) The matrix-based intra prediction (MIP) method is an intra prediction technique newly added to VVC. To predict samples of a rectangular block with width W and height H, MIP takes as input one column of H reconstructed adjacent boundary samples on the left side of the current block and one line of W reconstructed adjacent boundary samples above the block. If the reconstructed samples are not available, they are generated as done in conventional intra prediction. The generation of the prediction signal is based on the following three steps: optional averaging of the randomly selected reconstructed adjacent boundary samples as shown in FIG. 11, matrix-vector multiplication between the MIP weight matrix and the averaged adjacent boundary samples, and optional linear interpolation of the result from the previous multiplication.

[0080] In ECM, up to ECM-4.0, MIP has not been modified regarding its implementation in VVC.

[0081] Decoder side Intra Mode Derivation (DIMD) For a given luminance CB to be predicted, DIMD derives two intra prediction modes from a template of reconstructed adjacent samples surrounding this luminance CB, and these two intra predictors are combined with the planar mode predictor using weights derived from the determined gradients within the template. The division operation in weight derivation is performed using the same lookup table (LUT)-based integerization method used by the Cross Component Linear Model (CCLM). For example, the division operation in direction calculation Orient=G y / G x is calculated by the following LUT-based method. x=Floor(Log2(Gx)) normDiff=((Gx<<4)>>x)&15 x+=(3+(normDiff!=0)?1:0) Orient=(Gy * (DivSigTable[normDiff]|8)+(1<<(x-1)))>>x where DivSigTable

[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}.

[0082] The two derived intra prediction modes are included in the primary list of the MPM. Therefore, the DIMD process is executed before creating the MPM list for a given luminance CB to be predicted. For a given luminance CB, the primary derived intra prediction mode via DIMD is stored and used for constructing the MPM list of adjacent luminance CBs.

[0083] Geometric Partition Mode (GPM) In VVC and ECM, the geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signaled using a CU-level flag as a type of merge mode, and other merge modes include the regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. A total of 64 partitions, excluding 8×64 and 64×8, have w×h = 2 m ×2 n For each possible CU size of m,n ∈{3...6}, it is supported by the geometric partitioning mode.

[0084] When the GPM mode is used, the CU is divided into two parts by geometrically arranged lines. Some examples are shown in Figure 12. The position of the dividing line is mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition within the CU is inter-predicted using its own motion, and only one-sided prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Similar to conventional dual prediction, a one-sided prediction motion constraint is applied to ensure that only two motion compensation predictions are required for each CU.

[0085] When the geometric partitioning mode is used for the current CU, a geometric partitioning index indicating the partitioning mode (angle and offset) of the geometric partition, and two merge indices (one for each partition) are further signaled. The maximum number of GPM candidate sizes is explicitly signaled in the SPS, which specifies the syntax binarization for the GPM merge index. After predicting each part of the geometric partition, the sample values along the geometric partition edge are adjusted using a mixing process with adaptive weights. This is the prediction signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction modes.

[0086] Mixing along the geometric partition edge After predicting each part of the geometric partition using its own motion, mixing is applied to the two prediction signals to derive samples around the geometric partition edge. The mixing weights for each position of the CU are derived based on the distance between the individual position and the partition edge.

[0087] The distance of the position (x, y) with respect to the partition edge is derived as follows.

[0088]

Equation

[0089] The weights for each part of the geometric partition are derived as follows.

[0090]

Equation

[0091] partIdx depends on the angle index i. An example of the weight w0 is shown in FIG. 13.

[0092] GPM Intra In the exploration experiment in ECM-3.1 for ECM-4.0, the GPM Intra prediction mode is provided, and the Intra prediction mode is added to GPM to combine inter prediction with intra prediction. Four tests were conducted.

[0093] In the first variant (Test a), GPM with inter prediction and intra prediction is provided. In GPM with inter prediction and intra prediction, the final prediction sample is generated by weighting the inter prediction sample and the intra prediction sample for each GPM separation region. The inter prediction sample is derived in the same manner as the GPM within the current ECM, and the intra prediction sample is derived by the intra prediction mode (IPM) candidate list and the index signaled from the encoder. The IPM candidate list size is predefined as 3. The available IPM candidates are the parallel angle mode (parallel mode) with respect to the GPM block boundary, the vertical angle mode (vertical mode) with respect to the GPM block boundary, and the plane mode, as shown in FIGS. 14(a) to (c), respectively. Further, the GPM with intra and intra prediction as shown in FIG. 14(d) is restricted in the GPM using intra, reduces the signaling overhead for IPM, and avoids an increase in the size of the intra prediction circuit on the hardware decoder. In addition, to further improve the coding performance, the direct motion vector and IPM memory on the GPM mixing area are introduced.

[0094] In the second variant (Test b), two modifications, namely DIMD and adjacent mode-based IPM derivation, and the combination of GPM-intra and GPM-MMVD, are introduced into the first variant to achieve higher coding performance.

[0095] The IPM candidate list size is the same as that of the first variant form, and the parallel mode is registered first. Therefore, if the same IPM candidates do not exist in the list, up to two IPM candidates derived from ECM-3.1 and / or the decoder-side intra-mode derivation (DIMD) method in the adjacent blocks can be registered. Regarding adjacent mode derivation, there are up to five positions for the available adjacent blocks, which are restricted by the angle of the GPM block boundary as shown in Table 3 below, and this has already been used for GPM with template matching (GPM-TM) in ECM-3.1.

[0096] Unlike the first variant form, GPM with intra prediction (GPM-intra) can be used together with GPM with merge with motion vector difference (GPM-MMVD), which has already been implemented in ECM-3.1.

[0097]

Table 3

[0098] In the third variant form (Test c), template-based intra-mode derivation (TIMD) in ECM-3.1 can be additionally used for the IPM candidates of GPM-intra in order to further improve the coding performance. The IPM candidate list size is also the same as that of the first variant form. The parallel mode is registered first, and then TIMD, DIMD, and the IPM candidates of the adjacent blocks can be registered in this order.

[0099] In the fourth variant form (Test d), in addition to the third variant form, GPM-intra with GPM with template matching (GPM-TM) can be used to increase the application rate of the GPM-intra block.

[0100] Intra Block Copy (IBC) Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. The IBC tool significantly improves the coding efficiency of screen content materials. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is executed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has already been reconstructed within the current picture. The luma block vector of an IBC-coded CU is of integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pel motion vector precision and 4-pel motion vector precision. An IBC-encoded CU is treated as a third prediction mode other than the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luma samples.

[0101] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs an RD check on blocks that have either a width or a height not greater than 16 luma samples. In the non-merge mode, the block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a block-matching-based local search is performed.

[0102] In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowable block sizes. The hash key calculation for all positions within the current picture is based on 4×4 sub-blocks. For a current block of a larger size, the hash key is determined to match the hash key of the reference block when the hash keys of all 4×4 sub-blocks all match the hash keys at the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost for each matching reference is calculated, and the one with the minimum cost is selected.

[0103] In block matching search, the search range is set to cover both the previous CTU and the current CTU.

[0104] At the CU level, the IBC mode is signaled using a flag, which can be signaled as the IBC AMVP mode or the IBC skip / merge mode as follows. IBC skip / merge mode: The merge candidate index is used to indicate which of the block vectors in the list from adjacent candidate IBC-coded blocks is used to predict the current block. The merge list consists of spatial candidates, HMVP candidates, and pairwise candidates.

[0105] IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. For the block vector prediction method, one is from the left neighbor (if IBC-coded), one is from the upper neighbor, and two candidates are used as predictors. If neither neighbor is available, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index.

[0106] IBC reference region To reduce memory consumption and decoder complexity, IBC in VVC allows only the reconstructed parts of a pre-defined area that includes the area of the current CTU and an area of the left CTU. Figure 15 shows an example of the reference area in IBC mode, where each block represents a 64×64 luma sample unit. The current block to be predicted is shown in a striped pattern, the gray blocks correspond to the reconstructed blocks, and the blocks with an X mark are the blocks within the reconstructed area that are not available for IBC. As shown in Figure 15, depending on the position of the current coding CU position within the current CTU, the following applies. When the current block falls within the upper left 64×64 block of the current CTU, in addition to the already reconstructed samples within the current CTU, the reference samples within the lower right 64×64 block of the left CTU can also be referred to using the CPR mode. The current block can also refer to the reference samples within the lower left 64×64 block of the left CTU and the reference samples within the upper right 64×64 block of the left CTU using the CPR mode.

[0107] When the current block falls within the upper right 64×64 block of the current CTU, in addition to the already reconstructed samples within the current CTU, if the luma position (0, 64) for the current CTU has not yet been reconstructed, the current block can also refer to the reference samples within the lower left 64×64 block and the lower right 64×64 block of the left CTU using the CPR mode. Otherwise, the current block can also refer to the reference samples within the lower right 64×64 block of the left CTU.

[0108] If the current block falls within the bottom left 64×64 block of the current CTU, in addition to the already reconstructed samples within the current CTU, if the luma position (64, 0) for the current CTU has not yet been reconstructed, the current block can also refer to the reference samples within the top right 64×64 block and the bottom right 64×64 block of the left CTU using the CPR mode. Otherwise, the current block can refer to the reference samples within the bottom right 64×64 block of the left CTU using the CPR mode.

[0109] If the current block falls within the bottom right 64×64 block of the current CTU, the current block can only refer to the already reconstructed samples within the current CTU using the CPR mode.

[0110] This limitation enables the IBC mode to be implemented using local on-chip memory for hardware implementation.

[0111] IBC Adaptation for Camera-Captured Content IBC is an effective tool for screen content coding. It also shows an improvement in coding efficiency for some camera-captured content at the expense of a significant increase in coding time. The IBC adaptation method based on EE2-3.2 software is described below. This adaptation shows that IBC can obtain good coding gains with a controllable increase in coding time.

[0112] The decoder is exactly the same as EE2-3.2 when the CTU size is 128×128. This means that, as shown in Figure 16, the reference area for IBC is extended to the two CTU rows above the current CTU. Specifically, for the CTU (m,n) to be coded, the reference area includes the CTUs with indices (m - 2,n - 2)...(W,n - 2), (0,n - 1)...(W,n - 1), (0,n)...(m,n), where W represents the maximum horizontal index within the current picture.

[0113] However, when the CTU size is 256×256, the two additional rows of the upper CTU may require extra memory. To prevent IBC from using extra memory, when the CTU size is 256×256, the reference area is shown in FIG. 17. Specifically, assuming that the current CTU index is (m,n), the reference area includes the CTUs with indices (0,n)...(m,n) and (m - 1,n - 1)...(W,n - 1) as shown by the light gray blocks in FIG. 17, and the dark gray block is the current CTU (m,n).

[0114] In addition to the changes to the EE2 - 3.2 decoder when the CTU size is 256×256, the encoder of EE2 - 3.2 is modified to limit the sample - by - sample block vector search (or what is called local search) range to [-12,12] horizontally and [-12,12] vertically centered around the first block vector predictor for each IBC block.

[0115] Combined Intra - Inter Prediction (CIIP) In VVC and ECM, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., CU width × CU height is 64 or more) and both the CU width and CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. CIIP prediction combines the inter - prediction signal with the intra - prediction signal as its name indicates. The inter - prediction signal P inter in the CIIP mode is derived using the same inter - prediction process applied to the normal merge mode, and the intra - prediction signal P intraIt is derived following the normal intra prediction process using the planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, and the weight values are calculated according to the coding modes of the upper and left adjacent blocks (shown in FIG. 18) as follows. If the upper neighbor is available and is intra - coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. If the left neighbor is available and is intra - coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. When (isIntraLeft + isIntraTop) is equal to 2, set wt to 3. Otherwise, when (isIntraLeft + isIntraTop) is equal to 1, set wt to 2. Otherwise, set wt to 1.

[0116] The CIIP prediction is formed as follows. P CIIP = ((4 - wt) * [[ID=2A]]P inter + wt * P intra + 2) >> 2

[0117] Embodiments of a method and apparatus for encoding or decoding an image or video block are described herein in connection with FIGS. 19 - 26. The block is predicted based on a combination of a first predictor block and a second predictor block, and the first and second predictor blocks are obtained from a first intra prediction mode and a second intra prediction mode, respectively. Hereinafter, in any one of the embodiments described herein, the mode of predicting a block based on such a combination is referred to as luma intra fusion.

[0118] Note: There seems to be a formatting issue in the original text where "P" is used in a formula - like context without clear definition. I've left it as "P" in the translation. If it's a variable with a specific meaning in the original patent context, it should be translated according to that meaning. Also, in the formula translation, I've made a minor adjustment by adding "2A" for the repeated "P" in the formula for better readability of the formula structure.Any one of the embodiments described in this specification can be implemented in an intra prediction module of an image or video encoder / decoder, such as the intra prediction module 260 of the encoder 200 and the intra prediction module 360 of the decoder 300.

[0119] In one embodiment, the first intra prediction mode is obtained from a first set of intra prediction modes, the second intra prediction mode is obtained from a second set of intra prediction modes, and the first and second intra prediction modes are distinct. In one variant, the first and second sets of intra prediction modes are also distinct.

[0120] In one embodiment, the first set of intra prediction modes includes non-direction-based intra prediction modes, and the second set of intra prediction modes includes directional intra prediction modes (IPM). One of the first intra prediction mode and the second intra prediction mode is obtained from the first set of intra prediction modes, and the other of the first intra prediction mode and the second intra prediction mode is obtained from the second set of intra prediction modes.

[0121] For example, the first set of intra prediction modes includes one or more of the following intra prediction modes, namely, the planar mode, the DC mode, the MIP mode, and the IBC mode. In one variant, when included in the first set of intra prediction modes, the MIP mode is an adaptive MIP mode that uses a matrix specifically trained so as not to capture the directional characteristics of the block. In another variant, the first set of intra prediction modes can also include any other intra prediction mode that does not capture the directional characteristics of the block or captures only the lower frequencies of the block signal.

[0122] For example, the second set of intra prediction modes includes one or more of the following intra prediction modes, i.e., one or more of the 67 directional intra prediction modes of VVC, or an extended directional intra prediction mode of ECM or TIMD mode, one or more of the wide-angle intra prediction modes, a directional intra prediction mode parallel to the partition edge of the geometric partition mode, a directional intra prediction mode perpendicular to the partition edge of the geometric partition mode, an adaptive MIP mode using a matrix trained to collect only the directional features of the block, or an intra prediction mode provided by the DIMD or TIMD process, etc. The second set of intra prediction modes can also include any other intra prediction mode that captures one or more directional features of the block or captures the high frequency of the block signal.

[0123] Unless otherwise specified, it is understood that any one of the above and below embodiments can be combined with any other one or more of the above and below embodiments.

[0124] FIG. 19A shows an example of a method 1900 for encoding blocks of an image or video according to one embodiment. At 1901, a first predictor block is obtained from a first intra prediction mode. In one variant, the first intra prediction mode can be determined by testing a first set of intra prediction modes of the intra prediction mode, determining a rate-distortion cost for each of the first set of intra prediction modes, and selecting the best intra prediction mode with respect to rate-distortion. At 1902, a second predictor block is obtained from a second intra prediction mode. In one variant, the second intra prediction mode can be determined by testing a second set of intra prediction modes of the intra prediction mode in the same way as the first intra prediction mode. In another variant, the second intra prediction mode can be derived from reconstructed samples adjacent to the block. Other variants are further described below with respect to determining the second prediction mode. At 1903, a luma intra fusion prediction is obtained, that is, a prediction for the block is obtained by combining the first block predictor and the second block predictor. Embodiments for combining the first block predictor and the second block predictor are further described below. At 1904, the block is encoded using the prediction from the luma intra fusion.

[0125] FIG. 19B shows an example of a method 1910 for decoding an image or video block according to an embodiment. At 1911, a first predictor block is obtained from a first intra prediction mode. As an example, the first intra prediction mode can be determined by decoding syntax elements from a bitstream indicating the first intra prediction mode. At 1912, a second predictor block is obtained from a second intra prediction mode. In one variant, the second intra prediction mode can be determined by decoding one or more syntax elements providing the second intra prediction mode. In another variant, the second intra prediction mode can be derived on the decoder side in the same way as on the encoder side. Some variants for determining the second intra prediction mode are further described below. At 1913, a luminance intra fusion prediction for the block is obtained by combining the first block predictor and the second block predictor. Embodiments for combining the first block predictor and the second block predictor are further described below. At 1914, the block is decoded / reconstructed using the prediction from the luminance intra fusion.

[0126] Figure 20A shows an example of method 2000 for encoding blocks of an image or video according to another embodiment. In this embodiment, the determination of the prediction using luma-intra fusion responds to a determination of whether two intra prediction modes should be combined. In 2001, a first predictor block is obtained from a first intra prediction mode. In one variant, the first intra prediction mode can be determined in the same way as the embodiment described with respect to Figure 19A. In 2002, a second predictor block is obtained from a second intra prediction mode. In one variant, the second intra prediction mode can be determined in the same way as the embodiment described with respect to Figure 19A, or using any other variant further described below. In 2003, it is determined whether the first intra prediction mode should be combined or fused with the second intra prediction mode. In other words, it is determined whether the first predictor block should be combined with the second predictor block.

[0127] For example, based on cost, it is determined that the first intra prediction mode should be combined with the second intra prediction mode. In this variant, the cost is obtained when determining the second intra prediction mode. For example, the cost is determined when each intra prediction mode within a second set of intra prediction modes is tested, or when the second intra prediction mode is determined using the TIMD process. Other variants can be used to determine whether the first and second intra prediction modes should be combined.

[0128] In 2003, if it is determined that the first intra prediction mode should be combined with the second intra prediction, then in 2004, the prediction for the block is obtained by combining the first block predictor and the second block predictor. Embodiments for combining the first block predictor and the second block predictor are further described below.

[0129] In 2003, if it is determined that the first intra prediction mode should not be combined with the second intra prediction, in 2005, the prediction for the block is obtained without combining the first block predictor and the second block predictor. The prediction for the block can be obtained from the first block predictor, or from the second intra prediction mode, or any other prediction such as another intra prediction mode, or an inter prediction mode.

[0130] In 2006, the block is encoded using the prediction obtained in 2004 or 2005.

[0131] In another variant of FIG. 20A, the determination of whether two intra prediction modes should be combined is made before determining the second intra prediction mode. For example, this determination can be made based on the size of the block, or based on the first intra prediction mode, or other variants described below can be used.

[0132] In this other variant, the second intra prediction mode is determined only if it is determined that two intra prediction modes should be combined.

[0133] FIG. 20B shows an example of a method 2010 for decoding a block of an image or video according to another embodiment. In 2011, a first predictor block is obtained from the first intra prediction mode. As an example, the first intra prediction mode can be determined by decoding one or more syntax elements from a bitstream indicating the first intra prediction mode.

[0134] In 2012, it is determined whether the first intra prediction mode should be combined or fused with the second intra prediction mode. This decision can be made based on the block size, or based on the first intra prediction mode, or based on one or more syntax elements decoded from the bitstream, or based on other variations described below.

[0135] In 2012, if it is determined that the first intra prediction mode should be combined with the second intra prediction mode, then in 2013, the second intra prediction mode is determined and the second predictor block is obtained from the second intra prediction mode. In one variation, the second intra prediction mode can be determined in the same manner as described for the embodiment with respect to FIG. 19B, or using any other variation further described below. Then, in 2014, a prediction for the block is obtained by combining the first block predictor and the second block predictor. Embodiments for combining the first block predictor and the second block predictor are further described below.

[0136] In 2012, if it is determined that the first intra prediction mode should not be combined with the second intra prediction, then in 2015, a prediction for the block is obtained without combining the first block predictor and the second block predictor. The prediction for the block can be obtained from the first block predictor, or from the second intra prediction mode, or from any other prediction such as another intra prediction mode or an inter prediction mode. In any case, the prediction mode of the block is the same as the prediction mode used when encoding the block.

[0137] In 2016, the block is decoded / reconstructed using the prediction obtained in 2014 or 2015.

[0138] Figure 20C shows an example of method 2020 for decoding blocks of an image or video according to a further embodiment. In 2021, in a manner similar to that of Figure 20B, a first predictor block is obtained from a first intra prediction mode. In this embodiment, the determination in 2023 regarding whether two intra prediction modes should be combined is made in 2022 after obtaining a second intra prediction mode. Depending on the variant form, the second intra prediction mode can be determined by decoding one or more syntax elements that provide the second intra prediction mode, or can be derived on the decoder side in the same way as on the encoder side. Some variant forms for determining the second intra prediction mode are further described below. Next, a second block predictor is obtained from the determined second intra prediction mode.

[0139] In 2023, it is determined whether the first intra prediction mode should be combined or fused with the second intra prediction mode. This determination can be made in the same way as in Figure 20A or based on other variant forms described below.

[0140] In 2023, if it is determined that the first intra prediction mode should be combined with the second intra prediction mode, in 2024, the prediction for the block is obtained by combining the first block predictor and the second block predictor. Embodiments for combining the first block predictor and the second block predictor are further described below. In 2023, if it is determined that the first intra prediction mode should not be combined with the second intra prediction, in 2025, the prediction for the block is obtained without combining the first block predictor and the second block predictor. The prediction for the block can be obtained from the first block predictor, or from the second intra prediction mode, or any other prediction such as another intra prediction mode or an inter prediction mode. In any case, the prediction mode of the block is the same as the prediction mode used when encoding the block.

[0141] In 2026, the block is decoded / reconstructed using the prediction obtained in 2024 or 2025.

[0142] Some embodiments for determining a second intra prediction mode are described below.

[0143] In some embodiments, the second intra prediction mode is a directional intra prediction mode.

[0144] In some embodiments, the directional intra prediction modes to be combined are derived from a decoder-based process, i.e., decoder-side intra mode derivation (DIMD) or template-based intra mode derivation (TIMD), in order to avoid having to signal an IPM index for the second intra prediction mode. Below, it is assumed that the directional intra prediction modes to be combined are derived from a decoder-based process.

[0145] In a variant where the first prediction mode is determined to be the planar mode, the following variant can be used to derive the second intra prediction mode.

[0146] Classically, the DIMD process combines two IPMs with the planar mode. Thus, in one embodiment where DIMD is used to derive the second intra prediction mode (directional mode), the combination of two intra prediction modes (hereinafter also referred to as lum intra fusion) may not be applied to the lum CB using the planar mode.

[0147] In one embodiment where TIMD is used to derive the second intra prediction mode (additional directional intra prediction mode), the planar mode may be removed from the TIMD search.

[0148] In one embodiment, the second intra prediction mode can be determined as the second MPM (MPM[1]) in the list. Thus, in this embodiment, since the plane is MPM[0] and MPM[1] is often a direction-based mode, the first and second MPMs (MPM[0] and MPM[1]) from the MPM list are combined. In one variant, such a combination of the first MPM and the second MPM in the MPM list is performed only under specific conditions. For example, this can be done when the left neighbor and the upper neighbor are close, i.e., when the left and the upper use the same IPM, or when the indices of the left and the upper have an absolute difference of 1.

[0149] In one variant where the first prediction mode is determined to be the DC mode, the following variant can be used to derive the second intra prediction mode.

[0150] In one embodiment where TIMD is used to derive the second intra prediction mode (additional directional intra prediction mode), DC may be removed from the TIMD search.

[0151] In one variant where the first prediction mode is determined to be the MIP mode, the following variant can be used to derive the second intra prediction mode or to adapt the MIP mode. As described above, the MIP mode uses a matrix from a set of trained matrices to generate a prediction for the block. To efficiently apply a direction-based IPM (second intra prediction mode) on top of the MIP mode (used as the first intra prediction mode), the matrix used in the MIP mode can be retrained so as not to capture the direction of the block. This means that the retrained matrix will predict only the non-directional information in the block, leaving the directional information that should be predicted from the loop intra fusion mode that combines the first and second intra prediction modes.

[0152] In another embodiment, only a subset of the matrices (e.g., half) is retrained not to predict the direction, and the other matrices are retrained in the same way as described above for the classical MIP mode.

[0153] In those embodiments, when the MIP mode is used as the first intra prediction mode of a combination of two intra prediction modes (luma intra fusion), only the matrices trained not to predict the directionality information can be selected in the MIP mode.

[0154] In a variant, the combination of two intra prediction modes (luma intra fusion) is always applied to the retrained matrices, and in this way, no additional signaling is required. In fact, "no additional signaling is required" because (1) the decoder-based process derives the index of the directional intra prediction mode (the second intra prediction mode), so the identification of the directional intra prediction mode involved in luma intra fusion incurs no signaling cost, and (2) luma intra fusion is always applied to the MIP mode (the first intra prediction mode), so luma intra fusion does not need to be signaled.

[0155] In a variant where the first prediction mode is determined to be the IBC mode, the following variant can be used to derive the second intra prediction mode.

[0156] In one embodiment, when using luma intra fusion for luma CB with IBC as the first intra prediction mode, the second intra prediction mode uses the intra mode used by the original luma CB indicated by the IBC motion vector.

[0157] In one embodiment, the application of luma intra fusion is limited to the case where the block indicated by the IBC motion vector used is a normal intra-coded block using the directional IPM.

[0158] In one embodiment, loop intra fusion is applied even when the indicated block was not coded using a normal intra, but the propagated intra information is in a directional mode.

[0159] Depending on the embodiment used to combine the first and second intra prediction modes, i.e., loop intra fusion, one or more syntax elements are signaled in the bitstream together with the coded data representing the block. The one or more syntax elements can signal one or more of the following items: the first intra prediction mode, the second intra prediction mode, an indicator indicating whether the first intra prediction mode should be combined with the second intra prediction mode, an indicator indicating whether the second intra prediction mode is obtained from decoder-side intra mode derivation or from template-based intra mode derivation, an indicator indicating a weight from a set of weights to be used for the first predictor block when combining the first and second predictor blocks.

[0160] Several embodiments for signaling one or more syntax elements are described.

[0161] FIG. 21A shows an example of a method 2100 for signaling or decoding from a bitstream one or more syntax elements that enable determination of one or more intra prediction modes used to predict blocks of an image or video according to one embodiment. For example, method 2100 can be combined with the embodiments described with respect to FIGS. 19A and 19B. Here, it is assumed that the block is coded using a prediction that combines a first intra prediction mode and a second prediction mode as described in any one of the embodiments described herein. At 2101, one or more syntax elements indicating the first intra prediction mode are encoded within the bitstream and decoded from the bitstream, respectively. At 2102, one or more syntax elements indicating the second intra prediction mode are encoded within the bitstream and decoded from the bitstream, respectively.

[0162] FIG. 21B shows an example of method 2110 for signaling or decoding from a bitstream one or more syntax elements that enable determination of one or more intra prediction modes used to predict blocks of an image or video according to another embodiment. For example, method 2110 can be combined with the embodiments described with respect to FIGS. 20A-20B-20C. Here, assume that the block is coded using a prediction that combines a first intra prediction mode and a second prediction mode, as described in any one of the embodiments described herein. At 2111, one or more syntax elements indicating the first intra prediction mode are encoded within the bitstream and decoded from the bitstream, respectively. At 2112, one or more syntax elements indicating whether the first intra prediction mode should be combined with the second intra prediction mode are encoded within the bitstream and decoded from the bitstream, respectively. In other words, here it is signaled whether the lum intra fusion mode is used for the block. In one variant, the second intra prediction mode is not signaled within the bitstream. Thus, if the first intra prediction mode should be combined with the second intra prediction mode, the second intra prediction mode is derived at the decoder in the same manner as at the encoder. In another variant, one or more syntax elements can be signaled to indicate the second intra prediction mode.

[0163] FIG. 21C shows an example of method 2120 for signaling or decoding from a bitstream one or more syntax elements that enable determining one or more intra prediction modes used to predict blocks of an image or video according to another embodiment. For example, method 2120 can be combined with the embodiments described with respect to FIGS. 19A-19B or FIGS. 20A-20B-20C. Here, it is assumed that the block is coded using a prediction that combines a first intra prediction mode and a second prediction mode, as described in any one of the embodiments described herein. In 2121, one or more syntax elements indicating the first intra prediction mode are encoded within the bitstream and decoded from the bitstream, respectively. In 2122, it is determined whether the use of the loop-intra fusion mode is signaled in the bitstream. In other words, it is determined whether it is signaled in the bitstream whether the first intra prediction mode should be combined with the second intra prediction. As described in some of the above embodiments, the use of the combination of the first and second intra prediction modes can be disabled or enabled based on the size of the block or based on the first intra prediction mode. In another variant, the use of the combination of the first and second intra prediction modes can be disabled or enabled based on, for example, the cost evaluated when determining the second intra prediction mode using TIMD search. This determination of whether the use of the combination is signaled within the bitstream is performed in the same way in the encoder and decoder. If it is determined that the use of the combination of the first and second intra prediction modes is disabled, the same determination is performed for both the encoder and the decoder, so there is no need to signal whether the combination is used.

[0164] When it is determined that a combination of the first and second intra prediction modes (use of lumina intra fusion) should be signaled, the use of the combination of the first and second intra prediction modes is signaled to indicate whether the block is effectively predicted by the combined prediction of the first and second intra prediction modes or by another prediction. Then, at 2123, one or more syntax elements indicating whether the first intra prediction mode should be combined with the second intra prediction mode are encoded in the bitstream and decoded from the bitstream respectively. In other words, here it is signaled whether the lumina intra fusion mode is used for the block. The same variations for signaling or deriving the second intra prediction described above are also possible.

[0165] FIG. 21D shows an example of a method 2130 for signaling or decoding from a bitstream one or more syntax elements that enable determining one or more intra prediction modes used to predict a block of an image or video according to another embodiment. For example, method 2130 can be combined with the embodiments described with respect to FIGS. 19A-19B or FIGS. 20A-20B-20C. Here, it is assumed that the block is coded using a prediction that combines a first intra prediction mode and a second prediction mode as described in any one of the embodiments described herein. At 2131, one or more syntax elements indicating the first intra prediction mode are encoded in the bitstream and decoded from the bitstream respectively. At 2132, one or more syntax elements indicating the mode for deriving the second intra prediction mode are encoded in the bitstream and decoded from the bitstream respectively. For example, the one or more syntax elements indicate whether DIMD or TIMD is used to derive the second intra prediction mode. This embodiment can be combined with the embodiments described with respect to FIGS. 21A-21B-21C.

[0166] Figure 21E shows an example of method 2140 for signaling or decoding from a bitstream one or more syntax elements that enable determination of one or more intra prediction modes used to predict blocks of an image or video according to another embodiment. For example, method 2140 can be combined with the embodiments described with respect to FIGS. 19A - 19B or FIGS. 20A - 20B - 20C. Here, it is assumed that the block is coded using a prediction that combines a first intra prediction mode and a second prediction mode as described in any one of the embodiments described herein. In 2141, one or more syntax elements indicating the first intra prediction mode are coded within the bitstream and decoded from the bitstream, respectively. In 2142, one or more syntax elements that enable derivation of weights for combining the first and second intra prediction modes, as further described below, are coded within the bitstream and decoded from the bitstream, respectively. This embodiment can be combined with the embodiments described with respect to FIGS. 21A - 21B - 21C - 21D.

[0167] In the above-described embodiments, the first intra prediction mode is signaled to the decoder using one or more syntax elements. However, in some embodiments, the first intra prediction mode is not explicitly signaled to the decoder. For example, one or more syntax elements can be used to signal the use of luma intra fusion for blocks in which the first intra prediction mode and the second intra prediction mode are derived at the decoder. For example, the combination of the first and second intra prediction modes is known to the decoder.

[0168] In another embodiment, only the first intra prediction mode is signaled to the decoder, and the use of the combination of the first and second intra prediction modes is always activated. The second intra prediction mode is determined based on the first intra prediction mode, or is derived from the DIMD or TIMD process.

[0169] Further embodiments regarding signaling, which can be combined with the embodiments described above, are described below.

[0170] In some embodiments, the SPS flag is used to indicate whether the use of loop intra fusion may be used on a slice. In embodiments where TIMD (respectively, DIMD) is used to derive the second intra prediction mode of loop intra fusion, when TIMD (respectively, DIMD) is not enabled, this flag is not transmitted and is presumed to be 0.

[0171] In some embodiments, the use of loop intra fusion is limited to a specific block size. Specifically, loop intra fusion may not be permitted for blocks that are too small (e.g., if their width × their height is less than 32, or if their width or height is below a certain value, e.g., 8) in order to reduce latency issues for smaller blocks. Additionally, in some embodiments, loop intra fusion may not be permitted for blocks that are considered too large (e.g., if their width × their height is greater than 1024, or if their width or height is above a certain value, e.g., 32) for better performance.

[0172] In another embodiment, called "per-block signaling embodiment", when permitted within a slice and over a block, if the intra prediction mode selected to predict the block can be one of the first intra prediction modes of loop intra fusion, i.e., one of the above non-directional intra modes such as normal intra plane, MIP, IBC, etc., the loop intra fusion process is always signaled over this block.

[0173] For example, if the intra prediction mode selected to predict the block is the MIP mode, the "per-block signaling embodiment" is applied when the loop intra fusion process is always signaled over this block. FIG. 22 shows an example of the updated signaling of the intra prediction mode selected to predict the current loop CB in ECM-4.0. Given the legend of FIG. 22, note that the loop intra fusion flag (MIP) is coded using the CABAC context model. However, FIG. 22 is an example. The loop intra fusion flag (MIP) may be bypass-coded. Also, in FIG. 22, note that the loop intra fusion flag (MIP) is placed after the truncated binary coding of the MIP matrix index. Further, the loop intra fusion flag (MIP) may be placed between the MIP transpose flag and the truncated binary coding of the MIP matrix index. The loop intra fusion flag (MIP) may also be placed between the MIP flag and the MIP transpose flag. In FIG. 22, for example, TIMD is systematically used as a decoder-based process to derive the index of the second intra prediction mode involved in loop intra fusion, e.g., the directional intra prediction mode.

[0174] As another example, when the intra mode selected to predict a block is either the MIP mode or the planar mode, and the lum intra fusion process is always signaled on this block, the "per-block signaling embodiment" is applied. FIGS. 23A and 23B show an example of the updated signaling of the intra prediction mode selected to predict the current lum CB in ECM-4.0. Given the legend of FIGS. 23A and 23B, note that the lum intra fusion flag (PLANAR) is coded using the CABAC context model. However, the lum intra fusion flag (PLANAR) may be bypass-coded. In FIGS. 23A and 23B, for example, TIMD is systematically used as a decoder-based process to derive the index of a second intra prediction mode involved in lum intra fusion, such as the directional intra prediction mode.

[0175] FIGS. 23A and 23B show the signaling of the intra prediction mode selected to predict the current lum CB in ECM-4.0 in the case of the "per-block signaling embodiment" when the intra mode selected to predict a block is either the MIP mode or the planar mode and the lum intra fusion process is always signaled on this block. Here, note that when the ISP is selected to predict the current block or lum CB, i.e., when the ISP flag is 1, the lum intra fusion flag (PLANAR) is not signaled and is presumed to be 0. Also, note that when the MRL is selected to predict the current block, the planar mode cannot be the intra prediction mode selected to predict the current block, and thus, the lum intra fusion flag (PLANAR) is presumed to be 0.

[0176] In some embodiments, the loop-intra fusion process is always applied without additional signaling. In some embodiments, the process is always applied and not signaled for specific cases. For example, in embodiments where the MIP matrix has a retrained subset specifically for loop-intra fusion, the loop-intra fusion process is always applied to blocks that use the matrix from the retrained subset.

[0177] In some embodiments, both TIMD and DIMD can be used to derive the directional intra prediction mode. In those embodiments, a flag can be used to indicate which derivation process should be used.

[0178] In some embodiments, there is no decoder-side derivation process, and the directional mode used for loop-intra fusion is signaled to the decoder.

[0179] In some embodiments, when the block to be coded is eligible for loop-intra fusion, for example, a cost evaluation such as the TIMD process is always applied. If the SATD cost obtained by this evaluation is below a specific threshold, and only in that case, loop-intra fusion is performed. In some embodiments, this threshold is used to determine whether the loop-intra fusion process should be sent. For example, if the SATD cost exceeds the threshold, loop-intra fusion is never applied. Otherwise, whether to use it or not is signaled. Or, if the SATD cost is below the threshold, it is always applied, and otherwise, whether to apply it or not is signaled.

[0180] Several embodiments for determining the weights used when combining the first and second intra prediction modes in the loop-intra fusion mode are described below.

[0181] Before applying luma intra fusion, a first predictor block obtained from the prediction by the first intra prediction mode of the luma intra fusion mode (i.e., the prediction obtained from a normal plane, MIP, IBC, or other mode eligible for luma intra fusion) is called predA, and a second predictor block obtained from the prediction by a second intra prediction mode, for example, the prediction by a directional mode selected for luma intra fusion (i.e., the prediction from an intra mode selected by, for example, the TIMD or DIMD process) is called predB. The final prediction block obtained from luma intra fusion is called predF by the following equation. predF[x][y]=α[x][y]×predA[x][y]+β[x][y]×predB[x][y]

[0182] Thereby, α + β = 1.

[0183] In some embodiments, α and β have fixed values for the entire block. For example, in some embodiments, α = β = 0.5 for all samples, or in some other embodiments, α = 0.75 and β = 0.25.

[0184] In some embodiments, the weights used depend on the first intra prediction mode used for luma intra fusion. For example, in some embodiments, when luma intra fusion is used on MIP, the weights are α = β = 0.5, otherwise α = 0.75 and β = 0.25.

[0185] In some embodiments, after signaling that a set of weights exists and that loop intra fusion is used on a block, the selected weights are sent to the decoder. For example, the set of weights for α can be {1 / 4, 1 / 2, 3 / 4} (where β = 1 - α). In some embodiments, the weights vary according to the first intra prediction mode used for loop intra fusion. For example, the weights for α when using loop intra fusion on MIP can be {1 / 4, 1 / 2, 3 / 4}, but when used on IBC, they can be {3 / 8, 1 / 2, 5 / 8}.

[0186] In some embodiments, a process similar to the process used to determine whether loop intra fusion should be sent can be applied to determine the weights to be used. For example, if the cost evaluation of the additional mode (e.g., SATD given by the TIMD derivation process) is below a certain threshold, the weights can be, for example, α = 0.75 and β = 0.25, and if not, α = β = 0.5. Similarly, in some embodiments, if the SATD cost exceeds a certain threshold, the weights can be, for example, α = 0.25 and β = 0.75, and if not, α = β = 0.5.

[0187] In some embodiments, the weights for loop intra fusion may be derived for each block. For example, in some embodiments that use TIMD for loop intra fusion, the SATD cost costA of the first intra prediction mode can be calculated for the TIMD template. Using costB, which is the cost of the best direction base mode found by TIMD, the final weights can be derived as α = costA / (costA + costB) and β = 1 - α.

[0188] In some embodiments, the weights can vary within a block. For example, for each sample at position (x,y), where (0,0) is the upper left corner and (width, height) is the lower right corner, the weights can be given by the following equation.

[0189] [Number]

[0190] In some embodiments, the weights may depend on the rank r in the MPM list of the second intra prediction mode (direction-based IPM). For example, in some embodiments, when the direction-based mode selected by TIMD is within the first N modes of the MPM list, α is

[0191] [Number] calculated as, where N < M and β = 1 - α. In some embodiments, N = 8 and M = 16.

[0192] The embodiments described herein provide a new image or video compression tool that combines the predictions of two intra prediction modes. In some embodiments, the tool is used by combining a non-directional intra tool, such as a planar mode, a matrix-based intra prediction (MIP), or an intra block copy (IBC), with a directional intra prediction mode (IPM). To reduce the signaling cost, a tool such as a template-based intra mode derivation (TIMD) or a decoder-side intra mode derivation (DIMD) can be used to derive the direction-based mode at the decoder.

[0193] FIG. 24 shows a block diagram of a system in which aspects of the present embodiment can be implemented, according to another embodiment. FIG. 24 shows an embodiment of an apparatus 2400 for encoding or decoding blocks of an image or video, where the blocks are predicted using a luma fusion mode as described according to any one of the embodiments described herein. The apparatus includes a processor 2410 and can be interconnected to a memory 2420 through at least one port. Both the processor 2410 and the memory 2420 can also have one or more additional interconnections to external connections.

[0194] The processor 2420 is also configured to obtain a first intra prediction mode, obtain a second intra prediction mode, and encode or decode at least one block based on a combination of the first intra prediction mode and the second intra prediction mode, using any one of the embodiments described herein. For example, the processor 2421 is configured using a computer program product that includes code instructions implementing any one of the embodiments described herein.

[0195] In one embodiment shown in FIG. 25, in a transmission context between two remote devices A and B through a communication network NET, device A includes a processor associated with memories RAM and ROM configured to implement a method for encoding blocks of an image or video as described with respect to FIGS. 1, 2, 19A, 20A, 21A - 21E, and device B includes a processor associated with memories RAM and ROM configured to implement a method for decoding blocks of an image or video as described with respect to FIGS. 1, 3, 19B, 20B, 20C, 21A - 21E. Depending on the embodiment, devices A and B are also configured to predict blocks using a luma fusion mode as described with respect to FIGS. 4 - 23B.

[0196] According to one embodiment, the network is a broadcast network adapted to broadcast / transmit an encoded image or video from device A to a decoding device including device B.

[0197] FIG. 26 shows an example of the syntax of a signal transmitted through a packet-based transmission protocol. Each transmission packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include image or video data coded according to any one of the embodiments described above. In one variant, the signal includes data representing any one of the following items. - A first intra prediction mode, - A second intra prediction mode, - An indicator indicating whether the second intra prediction mode is obtained from decoder-side intra mode derivation or from template-based intra mode derivation is signaled in the bitstream. - An indicator indicating whether the first intra prediction mode should be combined with the second intra prediction mode is signaled in the bitstream for at least one block. - An indicator indicating the use of the loop filter mode for a block, - An indicator indicating the matrix to be used by the first intra prediction mode, where the first intra prediction mode is the MIP mode used in the loop filter mode. - An indicator indicating a weight from a set of weights used for the prediction obtained from the first intra prediction mode in the loop filter mode.

[0198] Various implementation modes involve decoding. When used in this application, "decoding" can include, for example, all or part of the process performed on the received encoded sequence to produce a final output suitable for a display. In various embodiments, such processing typically includes one or more of the processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such a process may also or alternatively include processes performed by decoders of various implementation forms described in this application, such as decoding resampling filter coefficients and resampling the decoded picture.

[0199] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstructed picture process including entropy decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a broader decoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.

[0200] Various implementation modes involve encoding. Similar to the above considerations regarding "decoding", "encoding" as used in this application can include, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more of the processes generally performed by an encoder, such as segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such a process may also or alternatively include processes performed by encoders of various implementation forms described in this application, such as determining resampling filter coefficients and resampling the decoded picture.

[0201] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a more general encoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.

[0202] Note that the syntactic elements used in this specification are for illustrative purposes only. Therefore, they do not exclude the use of other syntactic element names.

[0203] In this disclosure, various information such as, for example, syntax that can be transmitted or stored has been described. This information can be packaged or arranged in various ways, including, for example, general ways in video standards such as putting the information into an SPS, PPS, NAL unit, header (e.g., NAL unit header, or slice header), or SEI message. Other ways are also available, including, for example, general ways in system-level or application-level standards such as putting the information into one or more of the following. a. SDP (Session Description Protocol), for example, a format for describing a multimedia communication session for the purposes of session announcement and session invitation, as described in the RFC and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. For example, a DASH MPD (Media Presentation Description) descriptor as described in DASH and transmitted via HTTP, where the descriptor is associated with a representation or set of representations to provide additional characteristics to the content presentation. c. For example, an RTP header extension as used during RTP streaming. d. For example, an ISO Base Media File Format that uses, as used in OMAF and in some specifications, boxes which are object - oriented building blocks defined by a unique type identifier and length, also known as "atoms" in some cases. e. An HLS (HTTP Live Streaming) manifest transmitted via HTTP. The manifest can be associated with a version or set of versions of the content, for example, to provide characteristics of the version or set of versions.

[0204] When a figure is presented as a flow chart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flow chart of the corresponding method / process.

[0205] Some embodiments refer to rate distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is often considered, often due to computational complexity constraints. Rate distortion optimization is typically formulated to minimize a rate distortion function that is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, these techniques can be based on an extensive test of all encoding options including all considered modes or encoding parameter values, but involve their encoding cost and a complete evaluation of the associated distortion of the reconstructed signal after encoding and decoding. To reduce the encoding complexity, it is also possible to use faster techniques, especially with the calculation of approximate distortion based on predicted or prediction residual signals rather than the reconstructed signal. A mixture of these two approaches can also be used, such as by using approximate distortion for only some of the considered encoding options and complete distortion for other encoding options. In other approaches, only a subset of the considered encoding options is evaluated. More generally, many approaches employ any of various techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.

[0206] The implementations and aspects described herein can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if considered only in the context of a single form of implementation (e.g., only considered as a method), the implementation of the considered features can also be implemented in other forms (e.g., an apparatus or a program). For example, the apparatus can be implemented in appropriate hardware, software, and firmware. This method can be implemented, for example, by a processor that generally refers to a processing device including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes, for example, communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information among end users.

[0207] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and other variations that appear in various places throughout this application do not necessarily all refer to the same embodiment.

[0208] In addition, this application may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0209] Furthermore, this application may refer to "accessing" various information. Accessing information can include, for example, one or more of receiving information, obtaining information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0210] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information can include, for example, one or more of accessing the information or obtaining the information (e.g., from a memory). Further, "receiving" typically involves, in some manner, during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information.

[0211] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that any use of the following " / ", "and / or", and "at least one of" is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art of this and related arts, this can be extended for as many listed items as there are.

[0212] Also, as used herein, the term "signaling" specifically means indicating something to the corresponding decoder. Thus, in certain embodiments, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, an encoder can send specific parameters to a decoder (explicit signaling) so that the decoder can use the same specific parameters. In contrast, if the decoder already has other parameters along with that specific parameter, signaling that simply enables the decoder to know and select that specific parameter (implicit signaling) can be used without sending. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It will be understood that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the term "signal", although the term "signal" may also be used as a noun herein.

[0213] As will be apparent to those skilled in the art, an implementation can result in various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for implementing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. As is known, signals can be transmitted over various different wired or wireless links. A signal can be stored in a processor-readable medium.

[0214] Numerous embodiments have been described above. The features of these embodiments can be provided either alone or in any combination across various claim categories and types.

Claims

**Claim 1** A method, comprising, for at least one block of an image or video: decoding an indicator indicating a first intra prediction mode; acquiring a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode; decoding the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode. **Claim 2** An apparatus comprising one or more processors, the one or more processors being operable to, for at least one block of an image or video: decode an indicator indicating a first intra prediction mode; acquire a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode; decode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode. **Claim 3** A method, comprising, for at least one block of an image or video: encoding an indicator indicating a first intra prediction mode; acquiring a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode; encoding the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode. **Claim 4** An apparatus comprising one or more processors, the one or more processors being operable to, for at least one block of an image or video: encode an indicator indicating a first intra prediction mode; acquire a second intra prediction mode in response to a determination that the first intra prediction mode should be combined with the second intra prediction mode; encode the at least one block based on a combination of the first intra prediction mode and the second intra prediction mode. **Claim 5** The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the first intra prediction mode is obtained from a first set of intra prediction modes, the second intra prediction mode is obtained from a second set of intra prediction modes, and the first and second sets of intra prediction modes are distinct.

6. The method according to any one of claims 1, 3, or 5, or the apparatus according to any one of claims 2, 4, or 5, wherein one of the first intra prediction mode or the second intra prediction mode is obtained from a first set of intra prediction modes including a non-directional intra prediction mode, and the other of the first intra prediction mode and the second intra prediction mode is obtained from a second set of intra prediction modes including a directional intra prediction mode.

7. The method according to any one of claims 1, 3, or 5 to 6, or the apparatus according to any one of claims 2, 4, or 5 to 6, wherein the first intra prediction mode is one of a planar prediction mode, a DC prediction mode, an intra block copy prediction mode, and a matrix-based intra prediction mode.

8. The method according to any one of claims 1, 3, or 5 to 7, or the apparatus according to any one of claims 2, 4, or 5 to 7, wherein the second intra prediction mode is one of the directional intra prediction modes.

9. The method according to any one of claims 1, 3, or 5 to 8, or the apparatus according to any one of claims 2, 4, or 5 to 8, wherein at least one of the first intra prediction mode or the second intra prediction mode is signaled in a bitstream.

10. The method according to any one of claims 1, 3, or 5 to 9, or the apparatus according to any one of claims 2, 4, or 5 to 9, wherein obtaining the second intra prediction mode includes deriving the second intra prediction mode from reconstructed samples adjacent to the at least one block.

11. The method according to any one of claims 1, 3, or 5 to 10, or the apparatus according to any one of claims 2, 4, or 5 to 10, wherein the second intra prediction mode is obtained from at least one of decoder-side intra mode derivation or template-based intra mode derivation.

12. The method or apparatus according to claim 11, wherein an indicator indicating whether the second intra prediction mode is obtained from decoder-side intra mode derivation or obtained from template-based intra mode derivation is signaled in the bitstream.

13. The method according to any one of claims 3 or 5 to 12, or the apparatus according to any one of claims 4 to 12, wherein the determination that the first intra prediction mode should be combined with the second intra prediction mode is based on an indicator signaled for the at least one block in the bitstream.

14. The method according to any one of claims 3 or 5 to 13, or the apparatus according to any one of claims 4 to 13, wherein for the at least one block, the determination that the first intra prediction mode should be combined with the second intra prediction mode is based on a cost obtained when determining the second intra prediction mode.

15. The method or apparatus according to claim 13, wherein signaling of the indicator indicating whether the first intra prediction mode should be combined with the second intra prediction mode is based on a cost obtained when determining the second intra prediction mode.

16. The method or apparatus according to any one of claims 3 or 5 to 15, or the apparatus according to any one of claims 4 to 15, wherein for the at least one block, the determination that the first intra prediction mode should be combined with the second intra prediction mode is based on the size of the at least one block.

17. The first intra prediction mode is a matrix-based intra prediction mode that uses one matrix from a set of matrices, and at least one of the matrices from the set is trained to predict non-directional information within a block to be predicted, the method according to any one of claims 3 or 5 to 16, or the apparatus according to any one of claims 4 to 16.

18. For the at least one block, the determination that the first intra prediction mode should be combined with the second intra prediction mode is based on an indicator signaled in a bitstream indicating a matrix to be used by the first intra prediction mode, the method or apparatus according to claim 17.

19. The combination is a weighted average of the first intra prediction mode and the second intra prediction mode, the method according to any one of claims 1, 3, or 5 to 18, or the apparatus according to any one of claims 2, 4, or 5 to 18.

20. The weight used in the combination is based on at least one of an indicator signaled in a bitstream indicating a weight from a set of weights used for the first intra prediction mode, the first intra prediction mode, a cost obtained when determining the second intra prediction mode, and a rank of the second prediction mode in a list of intra prediction modes, the method or apparatus according to claim 19.

21. The weight used in the combination is derived from a cost obtained when determining the second intra prediction mode, the method or apparatus according to claim 19.

22. The weight used in the combination varies with the position of samples within the at least one block, the method or apparatus according to any one of claims 19 to 21.

23. A computer program product comprising instructions for causing one or more processors to perform the method according to any one of claims 1, 3, or 5 to 22.

24. A non-transitory computer-readable medium storing executable program instructions for causing a computer executing the instructions to perform the method according to any one of claims 1, 3, or 5 to 22.

25. A bitstream comprising data representing at least one block of an encoded image or video using the method according to any one of Claims 1, 3, or 5 to 22.

26. A non-transitory computer-readable medium storing the bitstream according to Claim 25.

27. A device, - the apparatus according to any one of Claims 2 or 4 to 22, and - at least one of: (i) an antenna configured to receive a signal comprising data representing an image or video, (ii) a band limiter configured to limit the signal to a frequency band comprising the data representing the image or video, or (iii) a display configured to display the image or video.

28. The device according to Claim 27, wherein the device comprises at least one of a television, a mobile phone, a tablet, and a set-top box.

Citation Information

Cited By

  • Intra Prediction Fusion Method, Video Encoding Method and Apparatus, Video Decoding Method and Apparatus, and Video Coding System

    JP2025521765A

  • Intra predictive fusion method, video encoding method and apparatus, video decoding method and apparatus, and video coding system

    JP7892808B2