Histogram normalization of blocks for decoder side intra mode derivation
By normalizing and merging histograms of oriented gradients for neighboring blocks, the method addresses inefficiencies in intra prediction, improving video compression efficiency and decoding accuracy.
Patent Information
- Application Number
- PCT/EP2024/086529
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video compression technologies face challenges in efficiently leveraging spatial and temporal redundancy for intra prediction, leading to suboptimal compression efficiency and decoding accuracy.
The method involves obtaining and normalizing histograms of oriented gradients for neighboring blocks to determine intra prediction modes, using spatial and HOG distance normalization for non-adjacent blocks, and merging these gradients to enhance decoder-side intra mode derivation.
Improves video compression efficiency and decoding accuracy by accurately predicting intra modes, reducing computational complexity, and enhancing the quality of reconstructed images.
Smart Images

Figure EP2024086529_17072025_PF_FP_ABST
Abstract
Description
[0001]HISTOGRAM NORMALIZATION OF BLOCKS FOR DECODER SIDE INTRA MODE DERIVATION This application claims the priority to European Application No. 24305025.9, filed on 8 January 2024, which is incorporated herein by reference in its entirety. TECHNICAL FIELD The present embodiments generally relate to video compression. The present embodiments relate to a method and an apparatus for encoding or decoding an image or a video. More particularly, the present embodiments relate to improving DIMD (decoder side intra mode derivation) prediction mode. BACKGROUND To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter picture correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. In inter prediction, motion vectors used in motion compensation are often predicted from motion vector predictor. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction. SUMMARY According to an aspect, a method for encoding an image or a video is provided. For at least one block of an image, at least one first histogram of oriented gradients is obtained based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block. The at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block. A prediction is determined for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, and the at least one block is encoded using the determined prediction. According to another aspect, an apparatus for encoding an image or a video is provided. The apparatus comprises one or more processors operable to, for at least one block of an image, obtain at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, determine a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, and encode the at least one block using the determined prediction. According to an aspect, a method for decoding an image or a video is provided. For at least one block of an image, at least one first histogram of oriented gradients is obtained based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block. The at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block. A prediction is determined for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, and the at least one block is decoded using the determined prediction. According to another aspect, an apparatus for decoding an image or a video is provided. The apparatus comprises one or more processors operable to, for at least one block of an image, obtain at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, determine a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, and decode the at least one block using the determined prediction. Further embodiments that can be used alone or in combination are described herein. One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any one of the methods for encoding or decoding an image or a video according to any of the embodiments described herein. One or more of the present embodiments also provide a non- transitory computer readable medium and / or a computer readable storage medium having stored thereon instructions for encoding or decoding an image or a video according to the methods described herein. One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving the bitstream generated according to the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1A illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to an embodiment. FIG. 1B illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment. FIG. 1C illustrates a block diagram of a system within which aspects of the present embodiments may be implemented according to another embodiment. FIG.2 illustrates a block diagram of an embodiment of a video encoder within which aspects of the present embodiments may be implemented. FIG.3 illustrates a block diagram of an embodiment of a video decoder within which aspects of the present embodiments may be implemented. FIG.4 illustrates an example of extraction of the gradients from the context of a WxH block to be predicted. FIG.5 illustrates an example of identification of a range of a target intra prediction mode index from absolute values of GVER and GHOR and the signs of GVER and GHOR in ECM-6.0. FIG.6 illustrates an example of computation of an angle Ɵ between the reference axis and thedirection being perpendicular to the gradient G of components GVER and GHOR when |^^^^^^^^| >|^^^^^^^^|, in the case that |^^^^^^^^|<0 and |^^^^^^^^|<0. FIG.7 illustrates an example of computation of an angle Ɵ between the reference axis and thedirection being perpendicular to the gradient G of components GVER and GHOR when |^^^^^^^^| ≥|^^^^^^^^|, in the case that |^^^^^^^^|<0 and |^^^^^^^^|<0. FIG. 8 illustrates an example of computation of the index of the target intra prediction modeindex in the conditions of FIG.6, i.e. |^^^^^^^^| > |^^^^^^^^|, in the case that |^^^^^^^^|<0 and |^^^^^^^^|<0.FIG. 9 illustrates an example of computation of the index of the target intra prediction modeindex in the conditions of FIG.7, i.e. |^^^^^^^^| ≥ |^^^^^^^^|, in the case that |^^^^^^^^|<0 and |^^^^^^^^|<0.FIG.10 illustrates an example of DIMD regions used to infer the location dependency of DIMD modes. FIG. 11 illustrates an example of determining an MHOG of a current WxH block from two HOGs of two blocks being neighbors of the current block selecting DIMD and a merge DIMD mode for intra prediction. FIG.12 illustrates an example of a workflow for normalizing each HOG involved in a merge of HOGs before applying the merge. FIG.13 illustrates an example of non-adjacent candidates for DIMD merge. FIG.14 illustrates an example of a method for encoding an image or a video according to an embodiment. FIG.15 illustrates an example of a method for decoding an image or a video according to an embodiment. FIG.16 illustrates an example of a method for obtaining a merge hOG for a current block to encode or decode, according to an embodiment. FIG.17 illustrates an example of non-adjacent blocks used for obtaining a merge HOG for a block according to an embodiment. FIG.18 illustrates an example of a method for determining a Merge HOG (MHOG) of a current WxH block from two HOGs of two blocks being non-adjacent of the current block, according to an embodiment that uses spatial distance normalization. FIG.19 illustrates an example of a method for determining a Merge HOG (MHOG) of a current WxH block from two HOGs of two blocks being non-adjacent of trhe current block, according to an embodiment that uses HOG distance (SAD) normalization. FIG. 20 shows two remote devices communicating over a communication network in accordance with an example of the present principles. FIG.21 shows the syntax of a signal in accordance with an example of the present principles. DETAILED DESCRIPTION This application describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well. The aspects described and contemplated in this application can be implemented in many different forms. FIGs. 1A, 1B, 1C, 2 and 3 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs.1A, 1B, 1C, 2 and 3 does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described. In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably. Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding. The present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. FIG. 1A-1C illustrates block diagrams of examples of systems in which various aspects and embodiments can be implemented. Any one of the systems 100A, 100B or 100B may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. In various embodiments, the system 100A, 100B or 100C is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100A, 100B or 100C is configured to implement one or more of the aspects described in this application. FIG. 1A illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. The system 100A includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100A includes at least one memory 120, e.g., a volatile memory device, and / or a non-volatile memory device, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The memory 120 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples. The processor 110 may be interconnected to the memory 120 by an interconnection bus 115. Program code to be loaded onto processor 110 to perform the various aspects described in this application is subsequently loaded onto memory 120 for execution by processor 110. In some embodiments, memory inside of the processor 110 is used to store program code instructions and to provide working memory for processing that is needed during encoding or decoding. The input to the elements of system 100A may be provided through various input devices (not represented). Both Processor 110 and memory 120 can also have one or more additional interconnections to external connections. FIG.1B illustrates a block diagram of an example of a system 100B in which various aspects and embodiments can be implemented. The system 100B includes the processor 110 and memory 120 as described in relation with FIG.1A. The input to the elements of system 100B may be provided through various input devices as indicated in block 105 which is described further below with FIG. 1C. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG.1B, include composite video. The various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards. The system 100B includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium. The system 100B may provide an output signal to various output devices, including a display, speakers, and other peripheral devices. The output devices may be communicatively coupled to system 100B via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100B using the communications channel 190 via the communications interface 150. FIG.1C illustrates a block diagram of an example of a system 100C in which various aspects and embodiments can be implemented according to another embodiment. Elements of system 100C, singly or in combination, may be embodied in a single integrated circuit, multiple Ics, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100C are distributed across multiple Ics and / or discrete components. The system 100C includes the processor 110 and memory 120 as described in relation with FIG.1A or 1B. System 100C includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples. System 100C includes an encoder / decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100C or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art. Program code to be loaded onto processor 110 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input data (image, video, volumetric content), the decoded data (image, video, volumetric content) or portions of the decoded data, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic. In some embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for data encoding and decoding operations, such as for MPEG-2, HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding also known as H.266, standard developed by JVET, the Joint Video Experts Team). The input to the elements of system 100C may be provided through various input devices as indicated in block 105, also mentioned in FIG.1B. Such input devices of system 100B or 100C include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG.1B or 1C, include composite video. In various embodiments, the input devices of block 105 in system 100B or 100C have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to- digital converter. In various embodiments, the RF portion includes an antenna. Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100B or 100C to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device. Various elements of the systems 100A, 100B or 100C may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using the suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards. Similarly as for the system 100B of FIG. 1B, the system 100C includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium. Data is streamed to the system 100B or 100C, in various embodiments, using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100B or 100C using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100B or 100C using the RF connection of the input block 105. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network. The system 100C may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other devices. The display 165 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide a function based on the output of the system 100C. For example, a disk player performs the function of playing the output of the system 100C. In various embodiments, control signals are communicated between the system 100C and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100C via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100C using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100C in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip. The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs. In any of the systems 100A, 100B or 100C, the embodiments can be carried out by computer program product comprising code instructions that implements any one of embodiments described herein. The computer program product may be computer software implemented by the processor 110 or by hardware, or by a combination of hardware and software. As a non- limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 120 of any one of the systems 100A, 100B or 100C can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 110 of any one of the systems 100A, 100B or 100C can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples. FIG. 2 illustrates an example of a block-based hybrid video encoder 200. Variations of this encoder 200 are contemplated, but the encoder 200 is described below for purposes of clarity without describing all expected variations. In some embodiments, FIG. 2 also illustrate an encoder in which improvements are made to the HEVC standard or a VVC standard (Versatile Video Coding, Standard ITU-T H.266, ISO / IEC 23090-3, 2020) or an encoder employing technologies similar to HEVC or VVC, such as an encoder ECM (Enhanced Compression Model) under development by JVET (Joint Video Exploration Team). Before being encoded, the video sequence may go through pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of color components), or re-sizing the picture (ex: down-scaling). Metadata can be associated with the pre-processing and attached to the bitstream. In the encoder 200, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units of, for example, CUs (Coding units) or blocks. In the disclosure, different expressions may be used to refer to such a unit or block resulting from a partitioning of the picture. Such wording may be coding unit or CU, coding block or CB, luminance CB, or block. A CTU (Coding Tree Unit) refers to a group of blocks or group of units or group of coding units (CUs). In some embodiments, a CTU may be considered as a block, or a unit as itself. Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction (260). In an inter mode, motion estimation (275) and compensation (270) are performed. The intra mode and / or the inter mode may comprise several distinct sub-modes. For example, the intra mode may comprise directional intra predictions, template-based intra prediction, decoder side intra mode derivation prediction, intra block copy prediction or other modes spatially predicting the samples values of the unit. The inter mode may comprise skip mode, merge mode according to which motion information is derived from a list of motion candidates and no motion vector prediction residual is encoded, an inter mode according to which motion information is derived from a list of motion candidates and further refined either by encoding motion vector prediction residual or by template-matching performed both at the encoder and the decoder, further inter modes are also possible. The encoder decides (205) which one of the intra mode or inter mode to use for encoding the unit. When different intra modes and / or inter modes are possible, the encoder decides (205) which of the intra modes or inter modes to use. The encoder indicates the intra / inter decision by, for example, one or more syntax element signaling the prediction mode. The encoder may also blend (205) intra prediction result and inter prediction result, or blend results from different intra / inter prediction methods. Prediction residuals are calculated, for example, by subtracting (210) the predicted block from the original image block. The motion refinement module (272) uses already available reference picture in order to refine the motion field of a block without reference to the original block. A motion field for a region can be considered as a collection of motion vectors for all pixels with the region. If the motion vectors are sub-block-based, the motion field can also be represented as the collection of all sub-block motion vectors in the region (all pixels within a sub-block have the same motion vector, and the motion vectors may vary from sub-block to sub-block). If a single motion vector is used for the region, the motion field for the region can also be represented by the single motion vector (same motion vectors for all pixels in the region). The prediction residuals are then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes. The encoder decodes (reconstructs) an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (240) and inverse transformed (250) to decode prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (265) are applied to the reconstructed picture to perform, for example, one or more of a deblocking filtering, an SAO (Sample Adaptive Offset) filtering or an ALF (Adaptive Loop Filter) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (280). Such filtered image is also referred to as a reference image in the following. FIG.3 illustrates a block diagram of a video decoder 300. In the decoder 300, a bitstream is decoded by the decoder elements as described below. Video decoder 300 generally performs a decoding pass reciprocal to the encoding pass as described in FIG.2. The encoder 200 also generally performs video decoding as part of encoding video data. In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (335) the picture according to the decoded picture partitioning information. The transform coefficients are de- quantized (340) and inverse transformed (350) to decode the prediction residuals. Combining (355) the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). In a similar manner as in the encoder, intra prediction and / or inter prediction may comprise several distinct sub-modes. The decoder obtains (370) the predictor block based on one or more syntax elements signaling the prediction mode among the available intra modes and inter modes. The decoder may blend (370) the intra prediction result and inter prediction result, or blend results from multiple intra / inter prediction methods. Before motion compensation, the motion field may be refined (372) by using already available reference pictures. In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380). Note that, for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side is identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture. The decoded picture can further go through post-decoding processing (385), for example, an inverse color transform (e.g. conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (201), or re-sizing the reconstructed pictures (ex: up-scaling). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream. Some of the embodiments described herein relates to intra prediction of a block of an image or a video to encode or to decode and more particularly to improving a Decoder Side Intra Mode Derivation (DIMD). Any one of the embodiments described herein can be implemented for instance in an intra prediction module of a video encoder and an intra prediction module of a video decoder. For instance, the embodiments described herein can be implemented in the intra prediction module 260 of the video encoder 200 in FIG.2 or the intra prediction module 360 of the video decoder 300 in FIG.3. Decoder Side Intra Mode Derivation (DIMD) is an intra prediction mode for a block that relies on the assumption that the reconstructed pixels surrounding the block to be predicted carries information to infer the texture directionality in this block, i.e. the intra prediction modes that most likely generate the predictions with the highest qualities. DIMD is a prediction mode for a block whose determination is performed in a same manner on the encoder and on the decoder. In the following, the explanations apply the same way on both the encoder and decoder sides. Inference in DIMD The inference of the indices of the intra prediction modes that most likely generate the predictions of highest qualities according to DIMD is decomposed into three steps. First, gradients are extracted from a context of reconstructed pixels around a given block to be predicted. Then, these gradients are used to fill a Histogram of Oriented Gradients (HOG). Finally, the indices of the intra prediction modes that most likely, according to DIMD derivation process, give the predictions with highest qualities are derived from this HOG, and a blending may be performed. Extraction of gradients from the context For a given block to be predicted, a L-shape context of ℎ rows of reconstructed pixels above this block and ^^ columns of reconstructed pixels on the left side of this block is considered, as illustrated on FIG.4. On FIG.4, the block to be predicted is displayed in white, the context is displayed in gray. The context contains ℎ rows of reconstructed pixels located above the block and ^^ columns of reconstructed pixels located on the left side of the block. The gradient filter is framed in black in FIG.4. The L-shape context is also usually referred as a template of the given block. At each reconstructed pixel of interest in this context, a local vertical gradient and a local horizontal gradient are computed. In Mohsen Abdoli, Thomas Guionnet, Mickael Raulet, Gosala Kulupana, Saverio Blasi. Decoder-side intra mode derivation for next-generation video coding. In ICME, 2020 (referred as [1] in the following), and in Mohsen Abdoli, Elie Mora, Thomas Guionnet, Mickael Raulet. NonCE-3: decoder-side intra mode derivation with prediction fusion. Contribution JVET-N0342 at the 4th JVET meeting in Geneva, from 19 to 29 March 2019 (referred as [2] in the following), and in ECM-6.0, the local vertical and horizontalgradients are computed via 3 × 3 vertical and horizontal Sobel filters respectively. Moreover,in [1], [2], and ECM-6.0, a reconstructed pixel of interest in this context refers to a reconstructed pixel at which the gradient filter does not go out of the context bounds. Therefore, in [1], [2], and ECM-6.0, the complete extraction of gradients can be summarized by the “valid”convolution of the 3 × 3 vertical and horizontal Sobel filters with the context. Note that, inECM-6.0, ℎ = 3 and ^^ = 3.Filling the Histogram of Oriented Gradients (HOG) In the HOG, each bin is associated to the index of a different directional intra prediction mode. At initialization, all the HOG bins are equal to 0. For each reconstructed pixel of interest at which the local vertical gradient ^^^^^^^^and the local horizontal gradient ^^^^^^^^are computed, a direction is derived from ^^^^^^^^and ^^^^^^^^, and the bin associated to the index of the directional intra prediction mode whose direction is the closest to the derived direction is incremented. This index is called the “target intra prediction mode index”. More precisely, for a given reconstructed pixel of interest, the derivation of the direction from ^^^^^^^^and ^^^^^^^^is based on the following observation. During the prediction of a block via a directional intra prediction mode, the largest gradient in absolute value usually follows the perpendicular to the mode direction. Therefore, the direction derived from ^^^^^^^^and ^^^^^^^^must be perpendicular to the gradient of components ^^^^^^^^and ^^^^^^^^. For instance, in ECM- 6.0 using the 65 VVC directional intra prediction modes, considering vertical and horizontal gradient filters for which the direction of positive vertical gradient goes from top to bottom and the direction of positive horizontal gradient goes from right to left, the mapping from the absolute values of ^^^^^^^^and ^^^^^^^^and the signs of ^^^^^^^^and ^^^^^^^^to the range of the target intra prediction mode index is displayed in FIG.5. In the framework of ECM using the VVC directional intra prediction modes, from FIG.5: in (1) the target intra prediction mode index belongs to ^2,17^,in (2) the target intra prediction mode index belongs to ^19, 33^,in (3) the target intra prediction mode index belongs to^34,49^,in (4) the target intra prediction mode index belongs to ^51, 66^,if ^^^^^^^^==0, the target intra prediction mode is vertical, i.e. its index is 50, if ^^^^^^^^==0, the target intra prediction mode is horizontal, i.e. its index is 18.Now, if |^^^^^^^^| > |^^^^^^^^|, the reference axis is the horizontal axis. Otherwise, the referenceaxis is the vertical axis. The angle ^^ between the reference axis and the direction beingperpendicular to the gradient ^^ of components ^^^^^^^^ and ^^^^^^^^ is given by tan(^^) =|^^^^^^^^|⁄ |^^^^^^^^| if |^^^^^^^^| > |^^^^^^^^|, tan(^^) = |^^^^^^^^|⁄ |^^^^^^^^| otherwise, as illustrated on FIG.6 and FIG.7. For the current reconstructed pixel of interest at which the local vertical gradient ^^^^^^^^and the local horizontal gradient ^^^^^^^^are computed, for the range of intra prediction mode indices found as in FIG.5, it is now possible to find the index of the intra prediction mode whose angle with respect to the reference axis is the closest to ^^. The bin associated to the index of the foundtarget intra prediction mode is then incremented by |^^^^^^^^| + |^^^^^^^^|. This means that, bydenoting ^^ the bin associated to the index of the found target intra prediction mode, ^^^^^^[^^] =^^^^^^[^^] + |^^^^^^^^| + |^^^^^^^^| . Note that, for the current reconstructed pixel of interest, if^^^^^^^^ = ^^^^^^^^ = 0, no bin in the HOG is incremented.Angle discretization For a given reconstructed pixel at which the local vertical gradient ^^^^^^^^and the local horizontal gradient ^^^^^^^^are computed, for the found range of the target intra prediction mode index (see FIG.5), the angle ^^ mentioned above is not directly compared to the angle of each intra prediction mode with respect to the reference axis in this range. Indeed, the absolute angle of each intra prediction mode with respect to its reference axis is stored in a scaled integer form.Therefore, ^̇^ = floor(tan(^^) × (1 ≪ 16)) is compared to the scaled integer form ^^^^ of theangle of the directional intra prediction mode of index ^^ from the reference axis, ^^ ∈ [|0, 16|].floor denotes the floor operation, operation that for an input x as a result of floor(x) returns the greatest integer less than or equal to x. Then, the absolute shift ^^∗from the index of thereference axis to the index of the target intra prediction mode is ^^∗ = m^i^n | ^^^^ − ^̇^|. The targetintra prediction mode index is finally equal to the index of the reference axis shifted by ^^∗. In the conditions of FIG. 6, FIG. 8 illustrates the computation of the index of the target intra prediction mode using the above-mentioned discretization of ^^. In the conditions of FIG. 7, FIG. 9 illustrates the computation of the index of the target intra prediction mode using the above-mentioned discretization of ^^. Inference of the intra prediction mode(s) Once the filling of the HOG is completed, the index of the directional intra prediction mode that most likely generates the prediction with the highest quality is selected as the one associated to the histogram bin of the largest magnitude. In [1] and ECM-6.0, the two bins with the largest magnitudes are identified to find indices of the directional intra prediction modes that most likely yield the two predictions with the highest qualities according to DIMD, and these two modes are linearly combined, optionally with PLANAR. In JVET-AB0148, the number of selected bins has been extended to 5, that is 5 predictions provided using the 5 selected indices of directional intra prediction modes are combined. Signaling of DIMD In ECM-6.0, for a given luma Coding Block (CB) to be predicted, DIMD is signaled via a DIMD flag, placed first in the decision tree of the signaling of the intra prediction mode selected to predict this luma CB, i.e. before the Template-Matching Prediction (TMP) flag and the Matrix-based Intra Prediction (MIP) flag. Location-dependent DIMD modes by regions In Saverio Blasi, Jani Lainema. AHG12 – Location-dependent Decoder-side Intra Mode Derivation. Contribution JVET-AB0116 at the 28th JVET meeting in Mainz, from 20 to 28 October 2022 (referred as [3] in the following), non-uniform, sample-based weights are proposed to blend the predictions obtained from the selected indices of the HOG. The usage of sample-based blending, and the specific weights to use for a given prediction, are inferred during the DIMD derivation process. When deriving a DIMD mode (that is a directional intra prediction mode), it is determined whether the derivation of such mode was mostly influenced by the template region above or on the left of the current block. If a DIMD mode was mostly derived from samples above the current block, then when blending the corresponding prediction, higher weights should be used for samples closer to the above portion of the block. To determine whether specific samples in the template contribute to inferring specific DIMD modes, three separate regions are considered within the DIMD template as illustrated on FIG. 10. The gradient computation is performed separately for samples in each region, resulting inthree histograms, ^^^^^^^^^^^^ , ^^^^^^^^^^ and ^^^^^^^^^^^^^^^^^^^^ respectively. For a directional mode ^^ ,^^^^^^^^^^^^[^^] represents the cumulative magnitude of all samples in the region ABOVE atdirection ^^. It should be noticed that the template area is extended by one sample on the top- left and one sample on the bottom-right, with respect to conventional DIMD. The full histogram of gradients for the whole template can then be computed as the sum of the three separate histograms. As in conventional DIMD, the two directional modes with largest and second-largest cumulative magnitude in the histogram are selected as main and secondary DIMD modes, ^^^^^^^^^^^^^^^^0and ^^^^^^^^^^^^^^^^1, respectively. Additionally, the histograms ^^^^^^^^^^^^and ^^^^^^^^^^can be used to determine whether ^^^^^^^^^^^^^^^^0and / or ^^^^^^^^^^^^^^^^1depend on a specific template region ABOVE or LEFT. In particular, the location-dependency of ^^^^^^^^^^^^^^^^^^, denoted as ^^^^^^^^^^^^^^, can be defined as:If : (^^^^^^^^^^^^[^^^^^^^^^^^^^^^^^^] > 2^^^^^^^^^^[^^^^^^^^^^^^^^^^^^]), then:^^^^^^^^^^^^^^ = 1, that is ^^^^^^^^^^^^^^^^^^ depends on region ABOVE.Else if : (^^^^^^^^^^[^^^^^^^^^^^^^^^^^^] > 2^^^^^^^^^^^^[^^^^^^^^^^^^^^^^^^]), then:^^^^^^^^^^^^^^ = 2, that is ^^^^^^^^^^^^^^^^^^ depends on region LEFT.Else: ^^^^^^^^^^^^^^ = 0, that is ^^^^^^^^^^^^^^^^^^ is not location-dependent.Blending is then performed to fuse the main and secondary DIMD predictions, ^^^^^^^^^^^^^^^^0and ^^^^^^^^^^^^^^^^1 , with the Planar prediction ^^^^^^^^^^^^^^^^^^^^ . In case no DIMD mode isdetermined to be location-dependent (meaning ^^^^^^^^^^^^0 == ^^^^^^^^^^^^1 == 0) then uniformblending is applied. Uniform weights ^^^^^^^^^^0, ^^^^^^^^^^1and ^^^^^^^^^^^^^^ are derived based on the relative magnitudes of the modes in the histogram, and the final DIMD prediction is computed as: Else, if at least one of the DIMD modes is inferred to be location-dependent, then sample-basedblending is used. A different weight is used to blend the predictions at each location (^^, ^^).If ^^^^^^^^^^^^^^ ≠ 0 the sample-based weights ^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^) for prediction ^^^^^^^^^^^^^^^^^^are computed so that the average weight used within the block is approximately equal to the uniform weight ^^^^^^^^^^^^and so that higher weights are used in the portion of the block closer to the region ABOVE or LEFT, depending on ^^^^^^^^^^^^^^. A range ∆^^is pre-defined,corresponding to the largest deviation of ^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^) from ^^^^^^^^^^^^. Higher valuesof ∆^^result in a higher variation of the weights within the block. In particular for a block ofsize ^^ × ^^, if ^^^^^^^^^^^^^^ = 1, then:^^^^^^^^^^^^^^^^^^^^^^ (^^, ^^)^^ ^^ = ^^^^^^^^^^^^ + ∆^^ − 2∆^^(^^−1)Else if ^^^^^^^^^^^^^^ = 2, then: If both ^^^^^^^^^^^^^^ ≠ 0, ^^ = 0,1, then the weights ^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^) are computed for bothpredictions as in one of the two above equations, depending on the value of ^^^^^^^^^^^^^^.Conversely, if ^^^^^^^^^^^^^^ = 0 and ^^^^^^^^^^^^(1−^^) ≠ 0, then the weights for ^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^)are determined as: Finally, the weights for the Planar prediction ^^^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^) are then determined as:^^^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^) = 64 - ∑1 ^^=0 {^^^^^^^^^^^^^^^^^^^^^^^^(^^, ^^)}The final location-dependent DIMD prediction is then determined as: DIMD merge In Saverio Blasi, Ivan Zupancic, Jani Lainema. AHG12 – Decoder Side Intra Mode Derivation Merge. Contribution JVET-AE0071 at the 31st JVET meeting in Geneva, from 11 to 19 July 2023 (referred as [4] in the following), a new DIMD merge mode is introduced in ECM-9.0. In this mode, for the current block, a Merged Histogram of Gradients (MHOG) is determined by combining the HOGs of neighboring blocks that are encoded using the conventional DIMD or the DIMD in merge mode for intra prediction. The conventional DIMD is the DIMD intra prediction wherein the HOG is determined for the current block using reconstructed samples of the template of the current block while the DIMD in merge mode is the DIMD intra prediction wherein the HOG (MHOG) is determined for the current block using HOG previously determined for neighboring blocks. In the DIMD in merge mode, the indices of the intra prediction modes used to predict the current block and the weights for the potential DIMD blending are derived from this MHOG. More precisely regarding the combination of HOGs, for the current block, if a single neighboring block selects conventional DIMD or the DIMD in merge mode, then the HOG of this neighboring block is equal to the MHOG of the current block. If at least two neighboring blocks select conventional DIMD or DIMD in merge mode, the two HOGs of these two neighboring blocks are averaged to form the MHOG of the current block. For a given block, the DIMD in merge mode is allowed only if the current block has at least one neighbor selecting conventional DIMD or DIMD in merge mode. If allowed, the use of DIMD in merge mode is signalled with a CABAC coded CU level flag.FIG.11 presents an example for determining the MHOG of a current ^^ × ^^ block. The current^^ × ^^ block (1100) has, among others, three neighboring blocks (1101), (1102), and (1103).Blocks (1102) and (1103) select DIMD for intra prediction whereas (1101) selects neither conventional DIMD nor DIMD in merge mode for intra prediction. During the encoding / decoding of the current frame to which the blocks belong, the HOG, denoted HOG1, of block (1103) was filled from the template (1104) of reconstructed samples of the block (1103). From the template (1105) of reconstructed samples of the block (1102), the HOG, denoted HOG2, of block (1102) was filled. The MHOG of block (1100) is then determined by averaging the two HOGs: HOG1and HOG2. DIMD merge with histogram normalization In [4], in the DIMD in merge mode, for the current block having at least two neighboring blocks selecting conventional DIMD or DIMD in merge mode for intra prediction, the merge of HOGs does not consider the dynamic of the HOGs to be merged. For instance, in FIG.11, as block (1105) is larger than block (1104), HOG2contains more gradient magnitude incrementations than HOG1. Therefore, during the merge of HOG1and HOG2, the contribution of HOG2to MHOG overshoots that of HOG1. This makes the principle of HOG merging less relevant. This can be solved by incorporating a histogram normalization before merging. The normalization of each HOG involved in the merge of HOGs before combining them to form the MHOG of the current block may be summarized in FIG.12.At 1200, the determination of the MHOG of the current ^^ × ^^ block may be triggered. At1201, the neighboring blocks that selects conventional DIMD or DIMD in merge mode for intra prediction may be put into a group of candidate blocks blockSet. At 1202, as optional step, the set of candidate neighboring blocks blockSet may be reduced via filtering according to a given criteria, yielding a reduced set of candidate neighboring blocks blockSetFiltered. As an example, the criteria for eliminating a given neighboring block from the candidates set may be based on a distance between the current block and the given neighboring block. As another example, the criteria for eliminating a given neighboring block may be based on a difference between the size of the current block and the size of the given neighboring block. At 1203, normalization information normInfo may be collected from thecurrent block and the candidates set blockSetFiltered . At 1204, using normInfo andblockSetFiltered, the HOG of each neighboring block in blockSetFiltered may be normalized,yielding {HOGnorm1, HOGnorm2, … }.At 1205, the normalized HOGs {HOGnorm1, HOGnorm2, … } may be merged into the MHOGof the current block. Normalization information collected at 1203 may be one or more of the following normalization factors. In a variant, the HOG may be normalized by the number of reference samples inside the template of the associated neighboring block. In this variant, the normInfo at (1203) in FIG.12 may be the size of the template of each neighboring block involved in the merge. In another variant, the HOG may be normalized by the number of pixels of the associated neighboring block. In this variant, the normInfo at (1203) in FIG. 12 may be the size of each neighboring block involved in the merge. In another variant, the HOG may be normalized by any value derived from a template or block size. As a first example, for each HOG of neighboring block involved in the merge, each of its bin may be divided by 2 times the number of pixels of the associated neighboring block. As a second example, for each HOG of neighboring block involved in the merge, each of its bin may be divided by the number of reference samples inside the template of the associated neighboring block plus 2. In another variant, the HOG of a neighboring block of the current block involved in the merge yielding MHOG may be normalized with respect to the bin magnitudes. For instance, the normalizationof this HOG may be as follows: HOGnorm =HOG ∑^^^^=1 HOG[^^], ^^ denoting the number of bins of thisHOG. For instance, in ECM-9.0, ^^ = 65, associated to the 65 directional intra predictionmodes in luma. As another example, the normalization of this HOG may be as follows:HOGnorm DIMD merge with non-adjacent blocks In [4] and as described above, when neighbor blocks are coded with one of a DIMD mode, the DIMD histograms are combined to form a new DIMD Merge histogram for the current block. DIMD Merge modes and weights are computed based on this merged histogram. In Junyan Huo, Jiawei Fan, Zhenyao Zhang, Yanzhuo Ma, Fuzheng Yang, Ming Li . AHG12 – Non-adjacent spatial candidates for DIMD merge. Contribution AF0106 at the 32nd Meeting, Hannover, from 13 to 20 October 2023 (referred as [5] in the following), a new DIMD merge mode with non-adjacent blocks is presented. In this contribution, non-adjacent spatial candidates are used for DIMD merge candidates list construction to determine the intra prediction for the current block. While DIMD in merge mode employs the DIMD information from neighboring (adjacent blocks) to predict the current block [4], as described above, the new DIMD merge with non- adjacent spatial block uses candidates from non-adjacent spatial blocks as illustrated in FIG. 13. FIG.13 illustrates some examples of non-adjacent spatial blocks (1301, 1303, 1305, 1304) that can be used as candidates for the DIMD in merge mode fr the current block (1302). As detailed in [5], the distances between non-adjacent candidates blocks (for e.g (1303), (1304), (1305)) and current block (1302) are defined based on the width and height of current coding block. In [5], the DIMD in merge mode includes a list of non-adjacent blocks. Normalization og HOG as describd above can be applied also to the HOG obtained from the non-adjacent blocks, however this does not consider the spatial distance between the current block and the non- adjacent blocks or the HOG distance between the HOG candidates from non-adjacent blocks and the HOG of the current block. This makes the principle of HOG merging less relevant in case of non-adjacent blocks. Some embodiments described herein provide a method wherein an histogram normalization is used for at least non-adjacent blocks that take into account spatial distance and / or HOG distance. In some embodiments, before merging, HOGs of non-adjacent blocks are normalized by their spatial distance to the current block and / or by their HOG distance with the current block. It is to be noted that, the descriptions hereafter apply the same way at the encoder side and the decoder side. In a variant, for each HOG of non-adjacent blocks involved in the merge into the MHOG of the current block, this HOG is normalized by the HOG distance (SAD) with the current block. In another variant, for each HOG of non-adjacent blocks involved in the merge into the MHOG of the current block, this HOG is normalized by the spatial distance with the current block. In a further variant, the HOG of each non-adjacent block of the current block involved in the merge giving rise to MHOG can be filtered by a filter before this merge. In another variant, after merging the HOGs of non-adjacent blocks of the current block involved in the merge, yielding MHOG, the MHOG can be filtered. As described above, each bin of an HOG (histogram of oriented gradients) is associated to an index of a directional intra prediction mode, such that one or more intra prediction modes can be derived from an HOG to predict a current block. In the following, the wording HOG, histogram of oriented gradients are used interchangeably. Also, the wording histogram of intra prediction directions can also be used. FIG.14 illustrates an example of a method 1400 for encoding an image or a video according to an embodiment. For example, the method 1400 is the one performed by the encoder described in relation with FIG. 2. It is considered here that at least one current block of the image or video to encode is to be encoded using a DIMD in merge mode. At 1401, a merge HOG (MHOG) is obtained for the current block based on one or more HOG previously obtained for one or more neighboring blocks. The neighboring blocks can be adjacent to the current block or non-adjacent to the current blocks. For example, the neighboring blocks can be obtained from a set of candidate blocks that are encoded using a DIMD prediction mode: either a conventional DIMD or DIMD in merge mode. For each considered neighboring block, a HOG has been previously obtained either from a determination from reconstructed samples of the template of the neighboring block or from a merging of HOG of neighboring blocks of the neighboring block. At 1401, when obtaining the MHOG for the current block, the HOG obtained for each one of the neighboring block of the current block using for obtaining the MHOG is normalized based on a distance between the current block and the neighboring block. In a variant, the distance is a spatial distance between the current block and the neighboring block. In another variant, the distance is an HOG distance between the current block and the neighboring block. In a variant, at 1401, the MHOG is obtained for the current block by merging one or more HOG previously obtained for one or more neighboring blocks and an HOG obtaiend for the current block using reconstructed samples of the template of the current block. In a variant, the MHOG is obtained for the current block by merging a part of the HOGs mentioned above, for intstance by merging only indices that have the highest magnitudes, or whose magnitudes is above a given value. Further details and variants for determining the MHOG for the current block are provided further below. At 1402, a prediction is determined for the current block using the MHOG obtained for the current block. One or more intra prediction modes are derived from the MHOG obtained for the current block. The prediction for the current block is obtained by merging all the predictions provided by the derived one or more intra prediction modes. For example, the prediction in a similar manner as described above using the one or more intra prediction directions that have the highest magnitudes in the obtained MHOG. At 1403, the current block is encoded using the prediction for example based determining a residue between the current block and the prediction, and encoding the residue into a bitstream. FIG.15 illustrates an example of a method 1500 for decoding an image or a video according to an embodiment. For example, the method 1500 is the one performed by the decoder described in relation with FIG. 3. It is considered here that at least one current block of the image or video to decode is to be decoded using a DIMD in merge mode. At 1501, a merge HOG (MHOG) is obtained for the current block in a similar manner as at the encoder, for example as in 1401 of FIG.14. At 1502, a prediction is determined for the current block using the MHOG obtained for the current block in a similar manner as at the encoder, for example as in 1402 of FIG. 14. At 1503, the current block is reconstructed using the prediction, for example by decoding a residue from a bitstream and adding the decoded residue to the prediction. An example of a method 1600 for obtaining the MHOG of a current block according to an embodiment is illustrated in FIG.16 wherein each HOG involved in the merge of HOGs before combining them to form the MHOG of the current block is normalized.At 1601, the determination of the MHOG of the current ^^ × ^^ block is triggered, for examplewhen it is decided that the current block is to be encoded or decoded using a DIMD in merge mode. At 1602, neighboring blocks of the current block that select conventional DIMD or the DIMD in merge mode for intra prediction are put into a set of candidate blocks: blockSet. These neighboring blocks are preferably non-adjacent blocks of the current block. In some variants, the neighboring blocks can also comprise adjacent blocks to the current block. In some variants, the set of candidate blocks can be pruned according to a given criteria and a part of the set blockSet may be selected. As an example, the given criteria for eliminating a given block from the set can be based on a distance between the current block and the given block. For example, if a spatial distance between the given block and the current block is above a given value. As another example, the given criteria for eliminating a given block from the set can be based on a difference between the size of the current block and the size of this given block. For example, if the difference is above a given value or if the current block and the given block have not a same shape. At 1603, as an optional step, the set of blocks blockSet can be reduced via filtering according to the given criteria mentioned above, yielding a reduced set of blocks blockSetFiltered. At 1604, distance information ^^^^^^^^^^^^^^^^ is collected from the current block and the blocks inthe set blockSetFiltered . As an example, ^^^^^^^^^^^^^^^^ can be the SAD (Sum of AbsoluteDistance) distance between a block in the set blockSetFiltered and the current block. Asanother example, ^^^^^^^^^^^^^^^^ can be the spatial distance between a block in the setblockSetFiltered and the current block. At 1605, using ^^^^^^^^^^^^^^^^ and the blocks in the set blockSetFiltered, the HOG of each block inthe set blockSetFiltered is normalized, yielding {HOGnorm1, HOGnorm2, … }.At 1606, the normalized HOGs {HOGnorm1, HOGnorm2, … } are merged into the MHOG ofthe current block. Normalization with respect to the spatial distance In a variant, for each HOG of a block involved in the merge into the MHOG of the current block, this HOG is normalized by the spatial distance d with the current block. For example, the block is not adjacent to the current block. The number of blocks involved in the obtention of the MHOG may vary regarding the strategy for selecting and eliminating blocks in the set of candidate blocks. For example, non-adjacent candidates can be searched around the current block until a given spatial distance. For example, as illustrated on FIG. 17, a range for searching non-adjacent candidates (1701) around the current block (1702) may have a size depending on the shape of the current block (1702). Examples of non-adjacent blocks are shown in FIG.17 in grey squares. Some of them are listed on the left of FIG. 17: for instance, (1703) is the block number 30, (1704) is the block number 27, and (1705) is the block number 27. For example, the current block (1702) and the non-adjacent blocks (e.g: (1703), (1704) and (1705)) may be a CU (Coding Unit). In this variant, the distInfo determined at (1604) in FIG. 16 can be the spatial distance to the current block of each non-adjacent block involved in the merge. In the example illustrated on FIG.17, d19 corresponds to the spatial distance between the neighboring block (1705) and the current block (1702), d30 corresponds to the spatial distance between the neighboring block (1703) and the current block (1702), d27 corresponds to the spatial distance between the neighboring block (1704) and the current block (1702). FIG.18 provides an example of a method for determining a Merge HOG (MHOG) of a current WxH block (1702) from two HOGs of two blocks (1703, 1705) being non-adjacent to the current block where the two non-adjacent blocks (1703, 1705) select DIMD for intra prediction. Some blocks, for e.g (1704), may select neither conventional DIMD nor DIMD in merge mode for intra prediction. Another criteria selection of non-adjacent block can be the spatial distance between this non-adjacent block and the current block. For e.g the non-adjacent block (1704) can be considered too far (d27) to the current block. In this case, this non-adjacent block (1704) is not selected. (1704) is an example of a non-selected block for obtaining a Merge HOG for the current block. During the encoding / decoding of the current frame to which the current block (1702) belongs, at 1840, the HOG, denoted HOG19, of the block (1705) is obtained. For example, HOG19is filled from the template of reconstructed pixels (1706). At 1830, the HOG, denoted HOG30, of the block (1703) is obtained. For example, HOG30is filled from the template of reconstructed pixels (1708). At 1850, HOG30is normalized by dividing each of its bins by the spatial distance between the block (1703) to the current block (1702). For instance, here, the spatial distance is d30 and maybe equal variant, (^^0, ^^0) and (^^30, ^^30) may bethe center of their respective blocks. At 1860, HOG19is normalized by dividing each of its bins by the spatial distance between the block (1705) to the current block (1702). For instance, here, the spatial distance is d19 and maybe equal to ^^19 = √(^^0 − ^^19)2 + (^^0 − ^^19)2. In a variant, (^^0, ^^0) and (^^19, ^^19) may be thecenter of their respective blocks. At 1870, the MHOG of the block (1702) is obtained by merging the normalized HOG of the blocks (1703, 1705): HOGnorm19and HOGnorm30. For instance, HOGnorm19and HOGnorm30are averaged, yielding the MHOG of the current block (1702). Normalization with respect to HOG distance In another variant, for each HOG of a block involved in the merge into the MHOG of the current block, this HOG is normalized by the HOG distance with the current block. For example, the block is a neighboring block not adjacent to the current block. The number of blocks involved in the obtention of the MHOG may vary regarding the strategy for selecting and eliminating blocks in the set of candidate blocks. For example, non-adjacent candidates are searched in a given range around the current block. In another example, the non-adjacent candidates (1701) around the current block (1702) can be selected depending on the HOG distance to the current block (1702). For example, if the distance between the HOG of a non-adjacent block and the HOG of the current block is below a given value, the non-adjacent block is selected and added to a list of candidate neighboring blocks, otherwise it is not selected. Examples of non-adjacent blocks are shown in FIG.17 in grey squares. Some of them are listed on the left of FIG.17: for instance, (1703) is the block number 30, (1704) is the block number 27, and (1705) is the block number 27. For example, the current block (1702) and the non- adjacent blocks (e.g: (1703), (1704) and (1705)) may be a CU (Coding Unit). In this variant, the distInfo determined at (1604) in FIG.16 is the Sum of Absolute Distance (SAD) between the HOG of each non-adjacent block involved in the merge and the HOG of the current block (1702). In a variant, another HOG distance can be the Euclidean distance. As illustrated on FIG.17, (1706) corresponds to the template of reconstructed pixels of the block (1705) (block number 19), (1708) corresponds to the template of reconstructed pixels of the block (1703) (block number 30) and (1707) corresponds to the template of reconstructed pixels of the block (1704) (block number 27). FIG.19 provides an example of a method for determining a Merge HOG (MHOG) of a current WxH block (1702) from two HOGs of two blocks (1703, 1705) being non-adjacent to the current block where the two non-adjacent blocks (1703, 1705) select DIMD for intra prediction according to this variant. Some blocks, for e.g (1704), may select neither conventional DIMD nor DIMD in merge mode for intra prediction. Another criteria selection of non-adjacent block can be as in the variant above the spatial distance between the non-adjacent block and the current block or the criteria selection can be the SAD distance determined between the template of the non-adjacent block and the template of the current block or the SAD distance determined between the HOG of the non-adjacent block (determined using reconstructed samples of the template of the non-adjacent block) and the HOG of the current block (determined using the reconstructed samples of the template of the current block). If the SAD distance is too high comparing to other SAD from non-adjacent block or comparing to an absolute value Th, the non-adjacent block is not selected. On FIG. 17, the block (1704) is an example of non-selected block. A variant can be to use a combination of SAD distance and spatial distance to select the non- adjacent blocks involved in the merging of the HOG for the current block. In a variant, the number of non-adjacent blocks can be limited to a given number, for e.g, the N non-adjacent blocks below Th are selected. During the encoding / decoding of the current frame to which the current block belongs, at 1940, the HOG, denoted HOG19, of the block (1703) is obtained. For example, HOG19is filled from the template of reconstructed pixels (1704) of the block (1703). At 1930, the HOG, denoted HOG30, of the block (1702) is obtained. HOG30is for example filled from the template of reconstructed pixels (1705) of the block 30 (1702). At 1920, the HOG, denoted HOG0, of the current block (1702) is filled from the template of reconstructed pixels (1902) of the current block (1702). At 1950, the HOG30of the block (1703) is normalized by dividing each of its bins by the HOG distance, here, ^^^^^^30determined between HOG30associated to block (1703) and HOG0associated to the current block (1702). For instance, here, the HOG distance ^^^^^^30isdetermined at 1990 as ^^^^^^30 = ∑^^ ^^=1 ^^^^^^(HOG30[^^] − HOG0[^^]) . At 1960, the HOG19of the block (1705) is normalized by dividing each of its bins by the ^^^^^^19determined between the HOG19associated to the block (1705) and the HOG0associated to the current block (1702). For instance, here, the HOG distance ^^^^^^19is determined at 1980 as^^^^^^19 = ∑^^ ^^=1 ^^^^^^(HOG19[^^] − HOG0[^^]).In a variant, the HOG distance can be the Euclidean distance between the HOG of the non- adjacent block and the HOG of the current block. For instance, in this variant Euclidean distance may be used instead of SAD. At 1970, the MHOG of the block (1702) is determined by merging the normalized HOG HOGnorm19and HOGnorm30obtained respectively at 1960 and 1950. For instance, HOGnorm19and HOGnorm30can be averaged, yielding the merge MHOG of the block (1702). In a variant, only the SAD distance between the most significant indices of the HOGs can be considered in the merging. For example, the SAD may be computed only on the indices i that corresponds to the 5 higher amplitudes of the HOG19and HOG30. In another variant, at 1970, the MHOG of the block (1702) is determined by merging the normalized HOG HOGnorm19and HOGnorm30obtained respectively at 1960 and 1950 and the HOG0 determined for the current block at 1920. Combination of HOG distance and spatial distance to normalize HOG. In a variant, the spatial distance and SAD can be combined to normalize HOGs. For e.g, concerning the normalization of HOG19may be normalized as below: with^^^^^^19=∑^^ ^^=1 ^^^^^^(HOG19[^^] − HOG0[^^])− ^^ 2 2 19) + (^^0 − ^^19)^^^^is a scale factor to weight the spatial distance. ^^^^^^^^is a scale factor to weight the SAD distance. Combining a first HOG normalization before spatial distance and / or SAD distance normalization In a variant, the HOGs may be normalized by the method described above in relation with FIG. 12 before the spatial distance and / or SAD distance normalization is done. Smoothing HOG Smoothing of HOGs involved in the merge before merging. In a variant, the HOG of each non-adjacent block of the current block involved in the merge giving rise to a MHOG for the current block can be filtered by a filter before this merge. Any filter that reduces the pixel noise in the template of reconstructed samples being reflected in the HOG bins can be used. For instance, a Gaussian filter can be used: gaussian ^^^^is the mean, and may be equal to 0. ^^^2^is the variance of the kernel of the gaussian filter. Then, the process of filtering may be expressed as follows: HOGnorm . gaussian(^^ − ^^) A LUT of the Gaussian filter could be used to approximate gaussian(y). In another example, a Laplace filter can be used: laplace ^^^^ is the mean, and may be equal to 0. ^^ > 0 is a scale parameter. Then, the process of filteringcan be expressed as follows. HOGnorm . laplace(^^ − ^^) A LUT of the Laplace filter could be used to approximate laplace(y). Smoothing of MHOG after merging In a variant, after merging the HOGs of non-adjacent blocks of the current block involved in the merge, yielding the MHOG for the current block, the MHOG can be filtered. Any filter can be used. For instance, in the case where a Gaussian filter is used, the process of filtering can be expressed as follows: .gaussian(^^ − ^^) Combination of HOG normalization and HOG smoothing In a variant, any of the embodiments related to the normalization of HOGs described above in relation with FIG. 14-19 can be straightforwardly combined with any of the embodiments related to the smoothing of HOGs described above. In an embodiment, illustrated in FIG.20, in a transmission context between two remote devices A and B over a communication network NET, the device A comprises a processor in relation with memory RAM and ROM which are configured to implement a method for encoding an or a video according to any one of the embodiments described herein and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement a method for decoding an image or a video according to any one of the embodiments described herein. In accordance with an example, the network is a broadcast network, adapted to broadcast / transmit a coded video from device A to decoding devices including the device B. FIG. 21 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may comprise data representative of at least one part of an image encoded according to any one of the embodiments described above. The payload can also comprise any signaling required for DIMD prediction mode. Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, entropy decoding a sequence of binary symbols to reconstruct image or video data. As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding, and in another embodiment “decoding” refers to the whole reconstructing picture process including entropy decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art. Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application, for example, determining re-sampling filter coefficients, re-sampling a decoded picture. As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art. Note that the syntax elements as used herein, are descriptive terms. As such, they do not preclude the use of other syntax element names. This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example. This information can be packaged or arranged in a variety of manners, including for example manners common in video standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, picture header or a slice header), or an SEI message. Other manners are also available, including for example manners common for system level or application level standards such as putting the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission. b. DASH MPD (Media Presentation Description) Descriptors, for example as used in DASH and transmitted over HTTP, a Descriptor is associated to a Representation or collection of Representations to provide additional characteristic to the content Representation. c. RTP header extensions, for example as used during RTP streaming. d. ISO Base Media File Format, for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications. e. HLS (HTTP live Streaming) manifest transmitted over HTTP. A manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions. When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process. Some embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion. The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users. Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment. Additionally, this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory. Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information. Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information. It is to be appreciated that the use of any of the following “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed. Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun. As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium. A number of embodiments has been described above. Features of these embodiments can be provided alone or in any combination, across various claim categories and types.
Claims
CLAIMS 1. A method, comprising: Obtaining, for at least one block of an image, at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, Determining a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, Encoding the at least one block using the determined prediction.
2. An apparatus comprising one or more processors operable to, obtain, for at least one block of an image, at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, determine a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, encode the at least one block using the determined prediction.
3. A method, comprising: obtaining, for at least one block of an image, at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, determining a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, decoding the at least one block using the determined prediction.
4. An apparatus comprising one or more processors operable to: obtain, for at least one block of an image, at least one first histogram of oriented gradients based on at least one second histogram of oriented gradients obtained for a neighboring block of the at least one block, wherein the at least one second histogram of oriented gradients is normalized based on a distance determined between the at least one block and the neighboring block, determine a prediction for the at least one block using one or more intra prediction modes determined using the at least one first histogram of oriented gradients, decode the at least one block using the determined prediction.
5. The method of claim 1 or 3 or the apparatus of claim 2 or 4, wherein the neighboring block is a block of the image that is not adjacent to the at least one block.
6. The method of any one of claim 1, 3 or 5 or the apparatus of any one of claims 2, 4 or 5, wherein the distance determined between the at least one block and the neighboring block is a spatial distance in the image between the at least one block and the neighboring block.
7. The method of any one of claim 1, 3 or 5 or the apparatus of any one of claims 2, 4 or 5, wherein the distance determined between the at least one block and the neighboring block is a distance determined between a third histogram of oriented gradients determined for the at least one block and the at least one second histogram of oriented gradients obtained for the neighboring block.
8. The method of any one of claim 1, 3 or 5-7 or the apparatus of any one of claims 2, or 4-7, wherein the neighboring block is encoded using a prediction determined using one or more intra prediction modes determined using the at least one second histogram of oriented gradients.
9. The method of any one of claim 1, 3 or 5-8 or the apparatus of any one of claims 2, or 4-8, wherein the at least one first histogram of oriented gradients is obtained by merging the at least one second histogram of oriented gradients with at least one other histogram of oriented gradients obtained for at least one other block of the image.
10. The method or the apparatus of claim 9, wherein the at least one other histogram of oriented gradients is determined for the at least one block.
11. The method or the apparatus of claim 9 or 10, wherein the neighboring block is part of a group of neighboring blocks of the at least one block and the at least one first histogram of oriented gradients is obtained by merging the at least one second histogram of oriented gradients with one or more histograms of oriented gradients obtained respectively for one or more neighboring blocks of the group.
12. The method of any one of claim 1, 3 or 5-11 or the apparatus of any one of claims 2, or 411, wherein obtaining the at least one second histogram of oriented gradients for the neighboring block comprises normalizing the at least one second histogram of oriented gradients based on a number of reference samples in a template of the neighboring block or based on a number of pixels in the neighboring block or based on magnitudes of the histogram of oriented gradients.
13. The method of any one of claim 1, 3 or 5-12 or the apparatus of any one of claims 2, or 4-12, wherein obtaining the at least one first histogram of oriented gradients comprises filtering the at least one second histogram of oriented gradients.
14. The method of any one of claim 1, 3 or 5-13 or the apparatus of any one of claims 2, or 4-13, wherein the obtained at least one first histogram of oriented gradients is filtered.
15. The method of any one of claim 1, 3 or 5-14 or the apparatus of any one of claims 2, or 4-14, wherein each bin of the at least one first histogram of oriented gradients is associated to an index of a directional intra prediction mode.
16. A computer program product including instructions for causing one or more processors to carry out the method of any of claims 1, 3 or 5-15.
17. A non-transitory computer readable medium storing executable program instructions to cause a computer executing the program instructions to perform the method of any of claims 1, 3 or 5-15.
18. A device comprising: an apparatus according to claim 4; and at least one of (i) an antenna configured to receive or transmit a signal, the signal including data representative of the image, (ii) a band limiter configured to limit the signal to a band of frequencies that includes the data representative of the image, or (iii) a display configured to display the image.
19. A device according to claim 18, wherein the device comprises at least one of a television, a cell phone, a tablet, a set-top box.
Citation Information
Patent Citations
Weighted prediction mode for scalable video coding
US20140072041A1
Usage of templates for decoder-side intra mode derivation
US20220224915A1
Systems and methods for histogram-based weighted prediction in video encoding
US20220303525A1
Method, device, and medium for video processing
WO2023051532A1
EP24305025A