Encoding and decoding method using template-based tools and corresponding device

JP2025531731A5Pending Publication Date: 2026-09-08INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025512937
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-07
Filing Date
2023-08-31
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

Existing image and video coding schemes face challenges in achieving high compression efficiency due to limitations in exploiting spatial and temporal redundancy within video content.

Method used

A template-based tool is used to identify unavailable pixels in a current block of a picture, allowing for improved encoding and decoding by determining information to be used for encoding or decoding the block, leveraging a decoding device with processors and memory to perform these methods.

Benefits of technology

Enhances encoding and decoding processes by optimizing the use of template-based tools, leading to improved compression efficiency and effective reconstruction of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A decoding method is disclosed. Information is obtained to identify which pixels (e.g., decoded pixels) are unavailable in a template of a current block of a picture. A template-based tool is further applied using the obtained information to determine information to be used to decode the current block. Finally, the current block is decoded using the determined information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of European Patent Application Publication No. 22306323.1, filed September 7, 2022, which is incorporated herein by reference in its entirety.

[0002] FIELD OF THE INVENTION At least one of the present embodiments generally relates to a method and apparatus for encoding and decoding picture blocks using the output of a template-based tool. [Background technology]

[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancy within video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-picture correlation, and then the difference between the original block and the predicted block, often called the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention

[0004] In one embodiment, a decoding method is disclosed, the decoding method comprising: Obtaining information to identify which pixels (e.g., decoded pixels) are unavailable within a template of a current block of a picture; applying a template-based tool using the obtained information to determine information to be used to decrypt the current block; and and decoding the current block using the determined information.

[0005] A decoding device is disclosed, the decoding device comprising one or more processors and at least one memory coupled to the one or more processors, the one or more processors configured to perform the method disclosed above.

[0006] In another embodiment, a method of encoding is disclosed, the method comprising: Obtaining information to identify which pixels (e.g., decoded pixels) are unavailable in a template of a current block of a picture; applying a template-based tool using the obtained information to determine information to be used to encode the current block; encoding the current block using the determined information.

[0007] An encoding device is disclosed, the encoding device comprising one or more processors and at least one memory coupled to the one or more processors, the one or more processors configured to perform the method disclosed above.

[0008] Further embodiments that can be used alone or in combination are described herein.

[0009] One or more embodiments also provide a computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method for predicting chroma samples or encoding / decoding image or video data according to any of the embodiments described herein. One or more of the present embodiments also provide a non-transitory computer-readable medium and / or computer-readable storage medium having stored thereon instructions for predicting chroma samples or encoding / decoding image or video data according to the methods described herein.

[0010] One or more embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide methods and apparatus for transmitting or receiving a bitstream generated according to the methods described above. [Brief explanation of the drawings]

[0011] [Figure 1] 1 illustrates a block diagram of a system in which aspects of the present embodiments may be implemented. [Figure 2] 1 illustrates a block diagram of one embodiment of a video encoder. [Figure 3] 1 illustrates a block diagram of one embodiment of a video decoder. [Figure 4A] Illustrates the HOG bin increment from the horizontal gradient GHOR and vertical gradient GVER calculated at a given decoded reference sample in the center row or center column of the current W×H luminance CB template being encoded / decoded. [Figure 4B] Illustrates the HOG bin increment from the horizontal gradient GHOR and vertical gradient GVER calculated at a given decoded reference sample in the center row or center column of the current W×H luminance CB template being encoded / decoded. [Figure 4C] Illustrates the HOG bin increment from the horizontal gradient GHOR and vertical gradient GVER calculated at a given decoded reference sample in the center row or center column of the current W×H luminance CB template being encoded / decoded. [Figure 5] 1 illustrates the input to the spatial 5-tap components of the filter when the CCCM predicts the current chrominance CB to be coded / decoded from a potentially downsampled version of the reconstructed luminance CB that is collocated with the current chrominance CB. [Figure 6A]Illustrates the current WxH chrominance CB (201) to be coded / decoded, its chrominance reference area (203), the downsampled reconstructed luminance CB (200) that is collocated with this chrominance CB, and the luminance reference area (202) for chroma format 4:2:0. [Figure 6B] Illustrates the current WxH chrominance CB (201) to be coded / decoded, its chrominance reference area (203), the downsampled reconstructed luminance CB (200) that is collocated with this chrominance CB, and the luminance reference area (202) for chroma format 4:2:0. [Figure 6C] Illustrates the current WxH chrominance CB (201) to be coded / decoded, its chrominance reference area (203), the downsampled reconstructed luminance CB (200) that is collocated with this chrominance CB, and the luminance reference area (202) for chroma format 4:2:0. [Figure 7] For SAD, the template (303) illustrates the search for a candidate reconstructed W×H luminance block (302) belonging to the search range of the TMP that is closest to the template (301) of the current W×H luminance CB (300) being coded / decoded. [Figure 8A] 1 illustrates the identification of reference samples (in terms of motion) of a template T for the current block to be coded / decoded in the case of a bidirectional merge candidate. [Figure 8B] 1 illustrates the identification of reference samples (for motion) for each sub-template in the template T of the current block being coded / decoded in the case of sub-block based merging candidates. [Figure 9A] index

[0012]

number

[0013]

number

[0014] This application describes various aspects, including tools, features, embodiments, models, approaches, and the like. Many of these aspects are described in specific, and often definitive, terms to at least illustrate their individual characteristics. However, this is for purposes of clarity of description and does not limit the applicability or scope of the aspects. In fact, all of the different aspects can be combined and substituted to provide further aspects. Furthermore, the aspects can be combined and substituted with aspects described in previous applications.

[0015] Aspects described and contemplated in this application can be implemented in many different forms. While Figures 1, 2, and 3 below provide some embodiments, other embodiments are contemplated, and discussion of Figures 1, 2, and 3 is not intended to limit the scope of implementations. At least one aspect generally relates to encoding and decoding video, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.

[0016] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "encoded" and "coded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, although not necessarily, the term "reconstructed" is used on the encoder side and the term "decoded" is used on the decoder side.

[0017] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. The order and / or use of specific steps and / or actions may be varied or combined, unless a specific order of steps or actions is required for proper operation of the method. Additionally, terms such as “first,” “second,” and the like may be used in various embodiments to modify elements, components, steps, operations, etc., e.g., “first decode” and “second decode.” The use of such terms does not imply any ordering of the modified operations unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but may occur before, during, or during an overlapping period with the second decode.

[0018] The present aspects are not limited to VVC or HEVC, but may, for example, be applied to other standards and recommendations, whether existing or developed in the future, and extensions of any such standards and recommendations (including VVC and HEVC).Unless otherwise indicated or technically precluded, the aspects described in this application may be used alone or in combination.

[0019] FIG. 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, alone or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, through a communication bus or dedicated input and / or output ports. In various embodiments, system 100 is configured to perform one or more of the aspects described herein.

[0020] System 100 includes at least one processor 110 configured to execute instructions loaded therein, for example, to implement various aspects described herein. Processor 110 may include internal memory, input / output interfaces, and various other circuits as known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes storage device 140, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage device 140 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0021] System 100 includes, for example, encoder / decoder module 130, which may include its own processor and memory, configured to process data to provide encoded or decoded video. Encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software, as is known to those skilled in the art.

[0022] Program code loaded onto processor 110 or encoder / decoder 130 to perform various aspects described herein may be stored in storage device 140 and then loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during execution of processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, expressions, operations, and computational logic.

[0023] In some embodiments, memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, H.265 and MPEG-H Part 2 are also known), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Team (JVET)).

[0024] Inputs to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals transmitted throughout a broadcast, for example, by a broadcaster, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other embodiments not shown in FIG. 1 include composite video.

[0025] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to (for example) as a channel, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, downconverting received signals to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and re-filtering RF signals transmitted over a wired (e.g., cable) medium to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, omit some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0026] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processor 110, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, as desired, in a separate interface IC or within processor 110. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130 operating in combination with memory and storage elements, to process the data streams as needed for presentation on an output device.

[0027] The various elements of system 100 may be provided within an integrated housing in which the various elements may be interconnected and transmit data to one another using suitable connection arrangements 115, e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards.

[0028] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.

[0029] Data is streamed to system 100, in various embodiments, using a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via communication channel 190 and communication interface 150 adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streamed data is provided to system 100 using a set-top box that delivers data via an HDMI connection in input block 105. Still other embodiments provide streamed data to system 100 using an RF connection in input block 105. As noted above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0030] System 100 may provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. Display 165 in various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 165 can be for a television, a tablet, a laptop, a mobile phone, or other device. Display 165 can also be integrated with other components (e.g., in the case of a smartphone) or separate (e.g., an external monitor for a laptop). Other peripheral devices 185, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (both terms referred to as a DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 to provide functionality based on the output of system 100. For example, a disc player performs the function of playing the output of system 100.

[0031] In various embodiments, control signals are communicated between system 100 and display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that allow inter-device control with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and speakers 175 may be integrated into a single unit with other components of system 100 in an electronic device, such as a television. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0032] Display 165 and speakers 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which display 165 and speakers 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0033] The embodiments may be implemented by computer software implemented by the processor 110, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 120 may be of any type appropriate to the technology environment and may be implemented using any suitable data storage technology, including, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, embedded memory, and removable memory. The processor 110 may be of any type appropriate to the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a multi-core architecture-based processor.

[0034] 2 illustrates an exemplary video encoder 200, such as a VVC (Versatile Video Coding) encoder. FIG. 2 may also illustrate an encoder that has improvements to the VVC standard or that employs techniques similar to VVC.

[0035] Before encoding, the video sequence may undergo encoding pre-processing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and added to the bitstream.

[0036] In the encoder 200, a picture is coded by the encoder elements as described below. The picture to be coded is divided (202) into units, e.g., CUs (Coding Units), for processing. Each unit is coded, e.g., using either intra mode or inter mode. When a unit is coded in intra mode, intra prediction (260) is performed. In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) whether intra mode or inter mode should be used to code the unit, and indicates the intra or inter decision, e.g., by a prediction mode flag. A prediction residual is calculated (210), e.g., by subtracting the predicted block from the original image block.

[0037] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as other syntax elements such as motion vectors and picture partition information, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying the transform or quantization processes.

[0038] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual is combined (255) with the predicted block to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset) / ALF (Adaptive Loop Filter) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer 280.

[0039] Figure 3 illustrates a block diagram of an exemplary video decoder 300. In the decoder 300, a bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs a decoding path that is the inverse of the encoding path described in Figure 2. Furthermore, the encoder 200 generally performs video decoding as part of encoding the video data.

[0040] In particular, the decoder's input includes a video bitstream, which may be generated by the video encoder 200. First, the bitstream is entropy decoded 330 to obtain transform coefficients, prediction modes, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture's partition information (335). The transform coefficients are inversely quantized (340) and inversely transformed (350) to decode the prediction residual. The decoded prediction residual is combined with a predicted block (355) to reconstruct an image block. The predicted block may be obtained from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375) (370). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380). Note that for a given picture, the contents of the reference picture buffer 380 on the decoder 300 side are identical to the contents of the reference picture buffer 280 on the encoder 200 side for the same picture.

[0041] The decoded picture may further undergo post-decoding processing (385), such as an inverse color transform (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0042] In the following, several template-based tools in the Enhanced Compression Model (ECM) are detailed. A template-based tool is configured to output information used to encode a block of a picture. This information can be of various types, such as a prediction of the block to be encoded, one or more prediction modes of the block to be encoded, one or more transforms to be used with the block to be encoded, a reordered list of merging candidates to be used for the block to be encoded, etc. This list is not exhaustive, and the embodiments are not limited to any particular template-based tool or to any particular type of output information.

[0043] Decoder-Side Intra Mode Derivation (DIMD) In ECM-4.0, DIMD derives the indices of two intra prediction modes that are likely to be the two best intra prediction modes for predicting the current luma coding block (CB) in terms of rate distortion from the gradients in the template of the decoded reference samples of the current luma coding block (CB) being coded / decoded. Then, it predicts the current luma coding block (CB) by blending two prediction blocks obtained by applying the derived two intra prediction modes with a prediction block obtained by applying PLANAR, one of the intra prediction modes defined in VVC. The weights involved in the blending are derived from the gradients in this template.

[0044] More specifically, for the current luma CB, the indices of the two intra-prediction modes are derived from the gradients in this template, as illustrated in Figures 4A, 4B, and 4C. First, a histogram of oriented gradients (HOG) with 65 bins corresponding to the 65 directional intra-prediction modes is initialized to 0. Then, for each decoded reference sample in the center row or column of the template, among the three rows of decoded reference samples above the current luma CB and the three columns of decoded reference samples to its left, the following procedure is applied.

[0045] The 3×3 horizontal Sobel filter and the 3×3 vertical Sobel filter are both centered in this decoded reference sample, as shown in FIG. 4A, and the horizontal gradient G HOR and vertical gradient G VER Each of these brings about:

[0046] G HOR and G VER The sign of the horizontal component G is determined in one of four orientation ranges, as illustrated in FIG. 4C. HOR and the vertical component G VER indicates whether a "target" direction can be found that is perpendicular to the gradient G of |G VER |>|G HOR If |, the anchor direction corresponds to the horizontal direction. |G HOR |≧|G VER If |, the anchor direction corresponds to the vertical direction. The "target" direction forms an angle θ with respect to the anchor direction, as shown in Figure 4B.

[0047] By discretizing a scaled version of tan(θ), the index i of the ECM directional intra prediction mode whose direction is closest to the “target” direction is found.

[0048] The HOG bin with index i is represented by |G HOR |+|G VER |Incremented.

[0049] Finally, the indices of the two largest HOG bins are the indices of the two derived intra prediction modes.

[0050] It should be noted that in the above procedure, the fact that the "target" direction is orthogonal to the gradient G is justified by the following principle: when a directional intra prediction mode extrapolates a reference sample along a direction D to a given area, the dominant gradient in this area is most likely to be orthogonal to D.

[0051] Convolutional Cross-Component Model (CCCM) In an Exploration Experiment (EE) on ECM-4.0, a Convolutional Cross-Component Model (CCCM) predicts the current chrominance CB to be coded / decoded by applying a convolution filter to a potentially downsampled version of the reconstructed luminance CB that is collocated with the current chrominance CB. When chroma subsampling is used, this downsampling is performed such that the resolution of the collocated reconstructed luminance CB after downsampling matches the resolution of the chroma grid.

[0052] The CCCM convolution 7-tap filter consists of a 5-tap plus sign spatial component, a nonlinear term, and a bias term. The input to the 5-tap spatial component of the filter consists of the center (C) luma sample co-located with the current chroma sample to be predicted, and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as shown in Figure 5.

[0053] The nonlinear term P is expressed as the square of the central luma sample C, scaled to the sample value range of the content. P=(C * C+midVal)>>bitDepth where bitDepth represents the pixel bit depth and midVal represents the median value of the bit depth range.

[0054] For example, for 10-bit content, it is calculated as follows: P=(C * C+512)>>10

[0055] The bias term B represents a scalar offset between the input and the output. B is set to the median chroma value, e.g., 512 for 10-bit content.

[0056] Calling c0, c1, c2, c3, c4, c5, and c6 the seven coefficients of the 7-tap filter, the current predicted chroma sample "predChromaVal" is expressed as follows: predChromaVal=clip(c0C+c1N+c2S+c3E+c4W+c5P+c6B) where "clip" clips to the range of valid chroma sample values.

[0057] filter coefficient c 0、 c 1、 c 2、 c 3、 c 4、 c 5、 c and c6 are calculated by minimizing the mean squared error (MSE) between predicted chroma samples generated by applying a convolutional 7-tap filter to potentially downsampled versions of the reconstructed luma samples in the luminance reference area (202) and the reconstructed chroma samples in the chrominance reference area (203), as shown in Figures 6A, 6B, and 6C.

[0058] Figure 6A illustrates a reconstructed luminance CB co-located with the current WxH chrominance CB (201) to be coded / decoded. Figure 6B illustrates a downsampled reconstructed luminance CB (200) co-located with this chrominance CB and a luminance reference area (202) for the chroma format 4:2:0, i.e., before coding, where the resolution of each chrominance channel is divided by 2 via subsampling. Figure 6C illustrates the chrominance reference area (203).

[0059] The luma reference area (202) consists of six rows / columns of potentially downsampled reconstructed luma samples above and to the left of the potentially downsampled version of the reconstructed luma CB (200) that is co-located with the current chrominance CB. The chrominance reference area (203) consists of six rows / columns of reconstructed chroma samples above and to the left of the current chrominance CB (201) to be coded / decoded. Each reference area extends one CB width to the right of the CB boundary and one CB height below the CB boundary. Each reference area is adjusted to include only available decoded reference samples. The extension to the black-shaded areas in Figures 6B and 6C is necessary to support side samples of the plus (+)-shaped spatial filter, and is padded when they fall in unavailable areas.

[0060] MSE minimization is performed by calculating the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is ​​LDL decomposed, and the final filter coefficients are calculated using inverse substitution. This process follows the calculation of Adaptive Linear Filtering (ALF) filter coefficients in ECM, except that LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. This calculation uses only integer arithmetic.

[0061] Note that single-model or multi-model variants of CCCM can be used. The multi-model variant uses two models, one derived for samples above the average luma reference value and the other derived for the remaining samples. The multi-model CCCM mode can be selected for coding units (CUs) that contain at least 128 available decoded reference samples.

[0062] It should also be noted that the term "reference area" has been chosen to be consistent with the standard nomenclature of the CCCM, but the reference area of ​​a given CB is equivalent to the template of this CB.

[0063] Template Matching Prediction (TMP) In ECM-4.0, template matching prediction (TMP) is an intra prediction mode that predicts the current W×H luminance CB to be coded / decoded.

[0064] For this purpose, a template (301) of the current luma CB (300) is made up of four rows of decoded reference samples above the current luma CB and four columns of decoded reference samples to its left, as shown in Figure 7. In the search step, for each allowed position within a given search range of the TMP in the current luma channel, a reconstructed WxH luma block candidate (302) whose top-left pixel is located at this allowed position is considered, and the Sum of Absolute Differences (SAD) between the template (303) of the four rows of decoded reference samples above it and the four columns of decoded reference samples to its left and the template (301) of the current luma CB (300) is calculated. The reconstructed luma block selected (the best candidate) is the one with the smallest template matching SAD.

[0065] The selected reconstructed luma block candidate is then used to predict the current luma CB.

[0066] Adaptive Reordering of Merge Candidates with Template Matching (ARMC-TM) In ECM-4.0, merge candidates are adaptively reordered via TM. For a given CU that is inter predicted, the merge mode derives all motion information from spatially and temporally neighboring CUs, called merge candidates. The reordering method applies to normal merge mode, TM merge mode, and affine merge mode (except for SbTMVP candidates). For TM merge mode, merge candidates are reordered before the refinement process. Essentially, when ARMC is used, merge candidates with a shorter template distance to the current block template are placed higher in the list.

[0067] After constructing the merge candidate list, the merge candidates are divided into several subgroups. The subgroup size is set to 5 for normal merge mode and TM merge mode. The subgroup size is set to 3 for affine merge mode. The merge candidates in each subgroup are reordered in ascending order according to their cost values ​​based on template matching. For simplicity, the merge candidates in the last subgroup but not the first subgroup are not reordered.

[0068] The template matching cost of a merge candidate is measured by SAD between the reconstructed samples of the template of the current block and their corresponding reference samples (in the motion, not the intra sense) of the template associated with the reference block (in the motion, not the intra sense). The template includes a set of reconstructed samples surrounding the current block to be coded / decoded. The reference samples (in the motion, not the intra sense) of the template of the reference block are located by the motion information of the merge candidate.

[0069] When a merge candidate utilizes bidirectional prediction, the reference samples (in terms of motion, not in the intra sense) of the merge candidate's template are also generated bidirectionally, as depicted in Figure 8A. Thus, Figure 8A illustrates the identification of reference samples (in terms of motion) of template T of a current block to be encoded / decoded in the case of a bidirectional merge candidate: (1) identification of reference samples (in terms of motion) in reference pictures of reference list 0 of template T surrounding the current block to be encoded / decoded using the motion vector of the merge candidate in reference list 0, and (2) identification of reference samples (in terms of motion) in reference pictures of reference list 1 of template T surrounding the current block to be encoded / decoded using the motion vector of the merge candidate in reference list 1.

[0070] For a sub-block-based merging candidate with a sub-block size equal to W×H, the upper template includes several sub-templates of size W×1, and the left template includes several sub-templates of size 1×H. As illustrated in Figure 8B, the motion information of the sub-blocks in the first row and first column of the current block to be coded / decoded is used to derive the reference samples (for motion) of each sub-template.

[0071] Matrix-based Intra Prediction (MIP) MIP is a linear intra prediction mode with a fixed learning matrix at both the encoder and decoder sides. Predicting the current W×H luma CB in MIP mode includes the following three steps: First, W decoded reference samples above the current luma CB and H decoded reference samples to the left of it are downsampled. Then, the downsampling result is linearly converted into a downscaled prediction. Finally, if necessary, the downscaled prediction is linearly interpolated so that the interpolated prediction has the same size as the current W×H luma CB.

[0072] More precisely, when W=4 and H=4, the downsampling factor is 2. In addition, as shown in FIG. 9A, the size of the MIP matrix in the linear transformation is 16×4 (4 input samples and 16 output samples). When either W=4 and H=8, or W=8 and H=4, or W=8 and H=8, the downsampling factor of the W decoded reference samples is W / 4, and the downsampling factor of the H decoded reference samples is H / 4. In addition, as shown in FIG. 9B, the size of the MIP matrix in the linear transformation is 16×8 (8 input samples and 16 output samples). For all other block sizes, the downsampling factor of the W decoded reference samples is W / 4, and the downsampling factor of the H decoded reference samples is H / 4. In addition, the size of the MIP matrix in the linear transformation is 64×8 (8 input samples and 64 output samples). Regarding the interpolation step, it should be noted that the horizontal interpolation of the downscaling prediction uses some of the H decoded reference samples but not their downsampled versions, and the vertical interpolation of the downscaling prediction uses some of the W decoded reference samples but not their downsampled versions.

[0073] When W=4 and H=4, there are 32 MIP modes. These modes are divided into pairs, and each pair uses the same MIP matrix, but the second mode of each pair swaps the downsampled reference sample above the current luminance CB with the downsampled reference sample to its left. The mapping from MIP mode index to MIP matrix index is depicted as shown in FIG. 9C. When swapping downsampled reference samples is applied, the downscaled prediction is transposed before interpolation. When W=4 and H=8, or W=8 and H=4, or W=8 and H=8, there are 16 MIP modes, and mode pairs still apply, as shown in FIG. 9D. For all other block sizes, 12 MIP modes are used, and mode pairs still apply.

[0074] Template-based neural networks without translational equivariance In parallel with standardization, new template-based tools based on neural networks have been developed. Considering two dimensions of translation, the translational equivariance of neural networks is achieved when the input is

[0075]

number

[0076]

number

[0077]

number

[0078] Apart from template-based tools with translational uniformity, there are template-based tools without translational uniformity. Figures 10A and 10B illustrate examples of templates fed into a typical template-based neural network without translational uniformity. In the first example in Figure 10A, for a given WxH block (310), the template (311) inserted into the neural network is the n aThe decoded pixel in the row and the n pixels to its left l The template is made up of the decoded pixels of the row and the template is extended to the right of this block by W and downward by H. In the extended part of the template, the unavailable decoded pixels are replaced / padded according to the process of unavailable decoded reference sample replacement specified by the HEVC and VVC standards. Thus, the number of unavailable decoded pixels in the template is b ∈[|0,H|] row (313), and n of the unavailable decoded pixels r The rightmost column (312) of the ∈[|0,W|] column is replaced, and n b and n r depends on the encoding / decoding partitioning history. In the second example of FIG. 10B, for a given W×H block (310), the template (314) fed into the neural network is not expanded this time. In both examples, the sequence of calculations in the neural network fed with the template for the W×H block never changes with the availability of decoded pixels in the input template. This exacerbates the tradeoff between the quality of the neural network output and the complexity of its inference.

[0079] For a given block to be coded (respectively decoded), the template, in its general design, does not include the decoded pixel on the upper right side of this block, nor does it include the decoded pixel on the lower left side. If the template includes the pixel on the upper right side and the pixel on the lower left side of this block, the unavailable pixels in these two extensions are typically replaced / padded according to the process of unavailable decoded reference sample replacement specified by the HEVC and VVC standards. Both restricting the template to its common design or extending the template while replacing / padding unavailable pixels reduces the relevance of the output of the template-based tool and, consequently, the coding efficiency of the block coded from this output.

[0080] However, depending on the size of the current block, its position within the current Coding Tree Unit (CTU), and its position within the current frame, decoded pixels on the top right and / or bottom left side of this block may be available. If most of the relevant intensity texture is located on the top right and / or bottom left side of this block, the fact that these decoded pixels are not included in the template can be seen as a significant loss of available information.

[0081] Therefore, it may be advantageous to extend the template towards the top right side of the block and towards its bottom left side. In a first embodiment, this extension towards the top right side of the block can cover as many available decoded pixels as possible, up to a limit of W additional columns of decoded pixels. This extension towards the bottom left side of the block can cover as many available decoded pixels as possible, up to a limit of H additional rows of decoded pixels. In other embodiments, there is no limit to the extension. Finally, various types of extended templates are proposed, for example, templates that completely surround the current block.

[0082] In the template of the block to be coded (respectively decoded), pixels (e.g., decoded pixels) may be unavailable because they have not yet been reconstructed / decoded or are not accessible due to the coding / decoding partitioning history. Pixels may also be inaccessible even if reconstructed due to specific coding constraints, for example, because they belong to a tile different from the tile to which the block to be coded belongs or because they are located outside the frame boundary. In the following, the terms "unavailable pixel" and "unavailable decoded pixel" are used interchangeably.

[0083] In this embodiment, a template-based tool is implemented, and an operation involving the template, such as a vector-matrix multiplication, is further provided with information identifying which decoded pixels in the template are unavailable, in order to skip at least a portion of the operation involving unavailable decoded pixels. In some embodiments, entire modules of calculation involving unavailable decoded pixels may be skipped. As an example, in the case of a template-based tool provided with a template surrounding the current block to be encoded / decoded, the template is expanded toward the upper right and lower left sides of the block, and if the template-based tool includes filters specific to the two expanded template portions and all decoded pixels in these two expanded portions are unavailable, this filtering may be skipped. The portion of the operation may be part of an element-wise multiplication between two tensors, part of a vector-matrix multiplication, or part of a vector reduction by downsampling.

[0084] FIG. 11 illustrates a method 1100 for encoding a block using a template-based tool according to one embodiment. At 1110, information is obtained for identifying which decoded pixels in the template of the block are unavailable. This information may be obtained, for example, in response to a partition history. For example, for a given frame encoded via VVC (with a CTU scan order from top-left to bottom-right and a hierarchical Z scan order for CUs), if the current CTU is partitioned via quadtree partitioning and the current CU is the bottom-right CU resulting from this partitioning, the partition history, i.e., the partition depth (1), the type of partitioning "quadtree," and the index of the current CU resulting from this partitioning (3), directly indicates that all decoded reference samples to the upper right side of the current CB within the current CU are unavailable. If a template of decoded reference samples around the current CB is extracted, this information regarding the unavailability of neighboring decoded reference samples may be the column index of the unavailable decoded pixels in the template.

[0085] At 1120, a template-based tool is applied using the information obtained at 1110 to determine information to be used to encode the block. More precisely, the information obtained at 1110 is implemented in the template-based tool and used to skip at least some of the operations involving unavailable decoded pixels.

[0086] At 1130, the block is encoded using information determined by the template-based tool, such as the predicted block, prediction mode index, transform type, merge candidate ordering, etc. The encoding of the block is performed by determining a residual between the pixels of the block and the prediction and encoding the residual.

[0087] In one example, an encoding device is disclosed that includes one or more processors and at least one memory coupled to the one or more processors, the one or more processors configured to perform the encoding method above.

[0088] 12 illustrates a method 1101 for decoding a block using a template-based tool, according to one embodiment. At 1111, information is obtained for identifying which decoded pixels in the block's template are unavailable. This information may be obtained, for example, in response to a segmentation history.

[0089] At 1121, a template-based tool is applied using the information obtained at 1111 to determine information used to decode the block. The information used to decode the block is identical to the information used to encode the block. More precisely, the information obtained at 1111 is implemented in the template-based tool and used to skip at least some of the operations involving unavailable decoded pixels. Steps 1111 and 1121 are identical to steps 1110 and 1120 on the encoder side.

[0090] At 1131, the pixels of the block are decoded using information determined by the template-based tool, e.g., predicted block, prediction mode index, etc. The decoding of the block is performed by decoding the residual and adding the residual to the prediction to reconstruct the pixels of the block.

[0091] The following examples apply to both the encoding and decoding methods.

[0092] In one example, applying the template-based tool using the acquired information includes skipping a calculation if the calculation involves a pixel identified by the acquired information as unavailable.

[0093] In one example, the method further includes flattening the template before applying the template-based tool.

[0094] In one example, obtaining information to identify which decoded pixels are unavailable in the template of the current block includes obtaining indices of all decoded pixels that are unavailable.

[0095] In one example, obtaining information to identify which decoded pixels are unavailable in the template of the current block includes obtaining indices of all decoded pixels that are available.

[0096] In one example, obtaining information to identify which decoded pixels are unavailable in the template of the current block includes obtaining flags, each flag indicating, for a pixel in the template, whether the pixel is available or not.

[0097] In one example, obtaining information to identify which decoded pixels are unavailable in the template of the current block includes, for a group of adjacent unavailable pixels, obtaining an index of the first unavailable pixel and an index of the last unavailable pixel in the group.

[0098] In one example, obtaining information to identify which decoded pixels are unavailable in the template for the current block includes, for a group of adjacent available pixels, obtaining an index of the first available pixel and an index of the last unavailable pixel in the group.

[0099] In one example, the method further includes spatially rearranging the pixels in the template to increase the number of contiguous decoded pixels available in memory.

[0100] In one example, the unavailable decoded pixels are one of the following: pixels that have not yet been reconstructed, pixels that belong to a tile different from the tile to which the current block belongs, or pixels that are outside the picture boundary.

[0101] In one example, the template-based tool belongs to a set of template-based tools that includes: - decoder-side intra mode derivation, -Convolutional cross-component models, -Adaptive reordering of merge candidates by template matching, Intra prediction based on a template matching prediction matrix, and -Template-based neural network prediction without translational equivariance.

[0102] In one example, the template includes pixels located all around the current block.

[0103] In one example, the template includes a line of pixels located above the current block and columns of pixels located to the right and left of the current block.

[0104] In one example, a decoding device is disclosed that includes one or more processors and at least one memory coupled to the one or more processors, the one or more processors configured to perform the encoding method above.

[0105] In one example, a computer program comprising program code instructions for implementing the steps of the encoding (respectively, decoding) method described above.

[0106] In one example, a computer-readable storage medium storing instructions for encoding or decoding blocks of a picture according to the encoding (respectively, decoding) method described above.

[0107] Additional embodiments are described below in connection with Figures 13A-26D.

[0108] Embodiment 1 The information for identifying which decoded pixels in the template of a block are unavailable may be a set L of indices of the unavailable decoded pixels in the template or in a transformed version of the template, where the transformation may be any transformation such as filtering, reshaping, rotating, flipping, or splitting.

[0109] For example, Figures 13A and 13B illustrate the method of Figures 11 and 12 when the operation in the template-based tool is vector-matrix multiplication and the input template is first flattened. A template (401) for a given W x H block (400) has n b ∈[|0,H|] the unavailable decoded pixel in row (403) and n pixels to the right of it. r∈[|0,W|] column (402) and unavailable decoded pixels. The notation [|a,b|] represents all integers in the range [a;b]. The template is first flattened to yield (404). The product (407) of the flattened template (404) with a weight matrix (406) then uses a set L of indices of unavailable decoded pixels (405) in the flattened template to skip calculations, i.e., multiplications and additions. In this figure, light gray squares indicate unavailable decoded pixels. In the weight matrix (406), light gray areas contain weights that are not used because the calculations are skipped. In other words, the output coefficient for index j is expressed as follows: Σ i∈[|0,s-1|]\L T i w i,j (Formula 1) T i : the coefficient of index i in the flattened template w i,j : index weight (i,j) s: The size of the flattened template

[0110] In a first variant of embodiment 1, the indices of any group of adjacent unavailable decoded pixels may be defined as the indices of the first unavailable decoded pixel and the last unavailable decoded pixel. As a result, each pair of indices in L is transformed into the set of all indices between the two indices of the pair, and the pair of indices is included in the set. The above (Equation 1) remains unchanged except that L is defined differently:

[0111]

number

[0112] In a second variation of the first embodiment, the information for identifying which decoded pixels in a template of a block are unavailable may be a set L of flags, each flag indicating whether an associated decoded pixel in the template or in a transformed version of the template is unavailable. In this embodiment, a flag equal to true indicates that the associated pixel is unavailable, and a flag equal to false indicates that the associated pixel is available. In the example illustrated in Figures 13A and 13B, L may be defined as follows:

[0113]

number

[0114] In fact, the template is a lines, each of which has (n l +2W-n r ) available pixels (hence flag equal to false) followed by n r of unavailable pixels (hence flag equal to true) and then n of available pixels l (hence the flag equals false) and finally (hence the flag equals false) l n b of unavailable pixels (2H-n b ) line, and

[0115] Then the output coefficients for index j are calculated as follows:

[0116]

number

[0117] Embodiment 2 Instead of specifying which pixels in the template are unavailable, the information for identifying which decoded pixels in a block's template are unavailable is a set of indices of available decoded pixels in the template or in a transformed version of the template.

[0118]

number

[0119] For example, Figures 14A and 14B illustrate the method of Figures 11-12 when the operation in the template-based tool is vector-matrix multiplication and the input template is first flattened. A template (501) for a given WxH block (500) has n b ∈[|0, H|] contains unavailable decoded pixels in row (503). The template is first flattened to yield (504). The product (507) of the flattened template (504) with a weight matrix (506) then yields a set of indices of available decoded pixels (505) in the flattened template.

[0120]

number

[0121]

number

[0122] In a first variant of the second embodiment, the index of any group of adjacent available decoded pixels may be defined as the index of the first available decoded pixel and the index of the last available decoded pixel, so that

[0123]

number

[0124]

number

[0125]

number

[0126] In a second variant of the second embodiment, the information for identifying which decoded pixels in the template of the block are unavailable is provided by a set of flags.

[0127]

number

[0128]

number

[0129]

number

[0130] The output coefficient of index j is then calculated as follows:

[0131]

number

[0132] Embodiment 3 In this embodiment, the template may be transformed to increase the number of contiguous decoded pixels available in memory in the transformed version of the template provided to the operation of interest in the template-based tool. The transformation may be, for example, a spatial reorganization of pixels in the template to increase the number of contiguous decoded pixels available in memory. Indeed, the benefit of having as many decoded pixels of the same type available adjacent to each other as possible is that when skipping calculations as described in the previous embodiment, some acceleration methods can be better utilized for non-skipped calculations. For example, assume that AVX-512 is used. Also assume that in the potentially transformed version of the template provided to the operation of interest, each decoded pixel is stored as a 16-bit integer. The greater the number of 32 contiguous packs of decoded pixels available, the better the acceleration of AVX-512.

[0133] This embodiment can be combined with any of the previous embodiments 1 or 2, and any of their variations. An example of the combination of this embodiment with embodiment 1 disclosed with respect to Figures 13A and 13B is shown below.

[0134] 15A-15B provide an example combination of embodiment 3 and embodiment 1, i.e., the information corresponds to the index of the unavailable decoded pixel. Furthermore, the operation of interest in the template-based video coding tool is vector-matrix multiplication, and flattening is performed before the vector-matrix multiplication.

[0135] The template (601) for a given W×H block (600) has n b ∈[|0,H|] the unavailable decoded pixel in row (603) and n pixels to the right of it. r ∈[|0,W|] column (602) and unavailable decoded pixels. The template is first split into two parts at line (604), and the part above line (604) is transposed to obtain (605) and (606). (605) and (606) are then flattened into a single vector (607). Finally, the product (610) of the flattened template (607) with a weight matrix (609) uses a set L of indices of unavailable decoded pixels (608) in the flattened template to skip multiplications and additions. In this example, (607) contains only two distinct groups of available decoded pixels that are contiguous in memory.

[0136] 16A-16B provide another exemplary combination of embodiment 3 and embodiment 1, i.e., the information corresponds to the index of unavailable decoded pixels. Furthermore, the operation of interest in template-based video coding tools is vector-matrix multiplication, and flattening is performed before the vector-matrix multiplication. In this figure, the crosses, circles, and diamonds only help visualize the cascade of flips about the vertical axis and transposition.

[0137] The template (701) for a given W×H block (700) has n b ∈[|0,H|] the unavailable decoded pixel in row (703) and n pixels to the right of it. r∈[|0,W|] column (702) and unavailable decoded pixels. The template is first split into two parts at line (704), and the part above line 704 is flipped about the vertical axis and then transposed to obtain (705) and (706). (705) and (706) are then flattened into a single vector (707). Finally, the product (710) of the flattened template (707) with a weight matrix (709) uses a set L of indices of unavailable decoded pixels (708) in the flattened template to skip multiplications and additions. In this example, (707) contains a single group of available decoded pixels that are contiguous in memory.

[0138] In this embodiment, L is the index n of the last unavailable decoded pixel that belongs to the contiguous group in the first memory of unavailable decoded pixels in (707). a n r −1 and the index n of the first unavailable decoded pixel belonging to a contiguous group of unavailable decoded pixels in (707), this time starting from the end of the vector. b n l -1, which allows for a compact representation of the information for identifying which decoded pixels in the template are unavailable.

[0139] 17A-17B provide another exemplary combination of embodiment 3 and embodiment 1, i.e., the information corresponds to the index of unavailable decoded pixels. Furthermore, the operation of interest in template-based video coding tools is vector-matrix multiplication, which is preceded by column-wise flattening. In this figure, the crosses, circles, and diamonds only help visualize the cascade of flips about the horizontal axis and transposition.

[0140] The template (801) for a given W×H block (800) has n b∈[|0,H|] the unavailable decoded pixel in row (803) and n pixels to the right of it. r ∈[|0,W|] and unavailable decoded pixels in column (802). The template is first split into two parts at line (804), and the left part of line 804 is flipped about the horizontal axis and then transposed to obtain (805) and (806). Next, (805) and (806) are flattened column-wise into a single vector (807). Column-wise flattening means sequentially scanning (806) and (805) column-wise from left to right, and each column from top to bottom, to obtain (807). Finally, the multiplication (810) of the flattened template (807) with the weight matrix (809) uses a set L of indices of unavailable decoded pixels (808) in the flattened template to skip multiplications and additions. In this example, (807) contains a single group of available decoded pixels that are contiguous in memory.

[0141] In this embodiment, L is the index n of the last unavailable decoded pixel that belongs to the contiguous group in the first memory of unavailable decoded pixels in (807). b n l −1 and the index n of the first unavailable decoded pixel belonging to a contiguous group of unavailable decoded pixels in (807), this time starting from the end of the vector. a n r -1, which allows for a compact representation of the information for identifying which decoded pixels in the template are unavailable.

[0142] Embodiment 4 In this embodiment, some of the decoded pixels may be inaccessible while already reconstructed due to specific constraints, e.g., tile independence during encoding / decoding.

[0143] For example, Figures 18A and 18B illustrate the method of Figures 11-12 when the operation in the template-based tool is vector-matrix multiplication, the current frame is divided into tiles, and the input template is first flattened. A template (901) for a given WxH block (900) has n tiles to its right. r ∈[|0,W|] column (902). Furthermore, the leftmost p∈[|0,n l The decoded pixels (903) in the |] column belong to the tile to the left of the tile containing block (900). The boundary between these two tiles is indicated (904). The template is first flattened to yield (905). The product (908) of the flattened template (905) and a weight matrix (907) then skips multiplications and additions using a set L of indices of decoded pixels that are unavailable or inaccessible because the input template overlaps two tiles (906) in the flattened template. In this figure, light gray squares indicate decoded pixels that are unavailable due to the encoding / decoding partitioning history or are inaccessible because the input template overlaps two tiles. In the weight matrix (907), the light gray areas contain weights that are not used because their calculations are skipped. As in embodiment 1, the output coefficient for index j is expressed as follows:

[0144]

number

[0145] All of the embodiments 1 to 3 and their variants can be combined with embodiment 4. In particular, the set of indices of decoded pixels that are available or accessible

[0146]

number

[0147] Embodiment 5 In this embodiment, some of the decoded pixels may be inaccessible while already reconstructed due to specific constraints, e.g., tile independence during encoding / decoding or outside frame boundaries.

[0148] For example, Figures 19A and 19B illustrate the method of Figures 11-12 when the operation in the template-based tool is vector-matrix multiplication, the current frame is divided into tiles, the template for a block extends outside the boundary of the frame containing this block, and the input template is first flattened. A template (1001) for a given WxH block (1000) has n tiles to its right. r ∈[|0,W|] column (1002) contains unavailable decoded pixels. a The decoded pixels (1003) in the |] row belong to the tile above the tile containing block (1000). The boundary between these two tiles is indicated (1005). Furthermore, the leftmost q∈[|0,n l The decoded pixels (1004) of the |] column go outside the left boundary (1006) of the frame containing the block (1000). The template is first flattened to yield (1007). The product (1010) of the flattened template (1007) with a weighting matrix (1009) then yields a set of indices of decoded pixels that are available or accessible given the position of the template relative to the tile boundaries within the flattened template and the frame boundary (1008).

[0149]

number

[0046] In this figure, light grey squares indicate decoded pixels that are unavailable due to the encoding / decoding partitioning history, or are inaccessible because the input template overlaps two tiles or goes outside the frame boundary. In the weight matrix (1009), light grey areas contain weights that are not used because their calculations are skipped. As in embodiment 2, the output coefficients for index j are expressed as follows:

[0150]

number

[0151] All of the embodiments 1 to 3 and their variants can be combined with embodiment 5. In particular, the set L of indices of decoded pixels that are unavailable or inaccessible can be selected from the set L as described above.

[0152]

number

[0153] Embodiment 6 The embodiments described above can be extended so that the template-based tool handles templates of multiple sizes. In this case, the template-based tool:

[0154]

number

[0155]

number

[0156]

number

[0157] Figure 20A shows a set of block sizes.

[0158]

number

[0159]

number

[0160] [Table 1]

[0161] This means that the shape of the "maximum template"

[0162]

number

[0163] For any given W×H block,

[0164]

number

[0165] For example, FIGS. 20B-20C illustrate the method of FIGS. 11-12 when the operation in the template-based tool is vector-matrix multiplication and the input template is first flattened.

[0166] "Maximum template" (1100) is

[0167]

number

[0168] For example, FIGS. 20D-20E illustrate the method of FIGS. 11-12 when the operation in the template-based tool is an element-wise vector-vector product and the input template is first flattened.

[0169] "Maximum template" (1200) is

[0170]

number

[0171]

number

[0172]

number

[0173] The index of a given decoded pixel in the potentially transformed version of the template fed into the vector-vector element-wise multiplication is

[0174]

number

[0175] All of the embodiments 1 to 5 and their variations can be combined with embodiment 6. Instead of being constructed to provide a "maximal template", the template-based tool can provide an "extended maximal template", i.e., a template in which each dimension is

[0176]

number

[0177] The disclosed embodiments for vector-vector element-wise product operations can also be applied to other types of operations, for example, vector-matrix multiplication.

[0178] Embodiment 7 Within a template-based video coding tool, an operation of interest that is supplied with information to identify which decoded pixels are unavailable may need to reinterpret this information, which may depend on how the template is processed before being supplied to the operation.

[0179] To illustrate this, Figures 21A-21B show an adaptation of the example of Figures 13A and 13B, where the template-based tool is a template-based neural network consisting of two convolutional layers and a fully connected layer. The convolution stride of the two layers is equal to 1, and the type of the two convolutions is SAME. In this embodiment, the operation considered in the template-based tool is vector-matrix multiplication in the fully connected layer. The template (3001) of a given W x H block (3000) has n convolutional layers at its bottom. b∈[|0,H|] row (3003) and n pixels to the right of it. r ∈[|0,W|] column (3002) and unavailable decoded pixels. (3002) and (3003) are removed from the template (3001), which is then split into two parts (3004) and (3005). A first convolutional layer with stride 1, SAME type, and n0 kernel takes (3004) and produces a 3D stack (3006). A second convolutional layer with stride 1, SAME type, and n1 kernel takes (3005) and produces a 3D stack (3007). SAME type means that the input to the convolutional layer is padded so that each of the two spatial dimensions of the output is equal to the corresponding dimension in the input divided by the stride. (3006) and (3007) are then flattened into a single vector (3008). In this flattening case, priority is given to the third dimension, then the second dimension. (3008) is fed into a fully connected layer of the neural network. As shown in Figure 21B, in set L (3009), each index of an unavailable decoded pixel in the flattened template must be reinterpreted to incorporate the fact that the template (3001) has been processed (removal of unavailable portions and application of convolutions). For example, index n l +2W-n r is n0(n l +2W-n r ) as another example, l +2W-1 is n0(n l 21A-21B, the calculation skips an amount that removes rows in the weight matrix (3010) with index L after reinterpretation.

[0180] As another example, Figures 22A-22B show the same case as Figures 21A-21B, but using convolution stride 2 instead of convolution stride 1. In Figures 22A-22B, L is not changed. However, the reinterpretation of each index of L is adapted as the processing of the template changes before it is fed into the vector-matrix multiplication in the neural network.

[0181] In templates (3001) and (4001), light grey squares indicate unavailable decoded pixels. In weight matrices (3010) and (4010), light grey areas contain weights that are not used because their calculations are skipped.

[0182] In all embodiments disclosed above, the template is extended by W additional columns of decoded pixels toward the top right side of the block and by H additional rows of decoded pixels toward the bottom left side of the block. However, these embodiments are not limited to these extensions and can be generalized to different extensions, i.e., different sizes and different shapes.

[0183] Embodiment 8 As an example, Figs. 23A and 23B show w = 2W + 4 and e H 13A and 13B using =2H+4. A template (1301) for a given W×H block (1300) has n b ∈[|0,e H -H|] Unavailable decoded pixel in line (1303) and n pixels to the right of it r ∈[|0,e w -W|] column (1302) and unavailable decoded pixels. The template is first flattened to yield (1304). The multiplication (1307) of the flattened template (1304) with a weight matrix (1306) then uses a set L of indices of unavailable decoded pixels (1305) in the flattened template to skip multiplications and additions.

[0184] In all embodiments disclosed above, a template extended towards the top right side of the block to be coded can cover as many available decoded pixels as possible, within a limit of W additional columns of decoded pixels. An extension towards the bottom left side of this block can cover as many available decoded pixels as possible, within a limit of H additional rows of decoded pixels. In other embodiments, there is no limit to the extension. Finally, various types of extended templates are proposed, for example, templates that completely surround the current block.

[0185] In all of the previously disclosed embodiments, the template is expanded towards the top right side of the block and towards the bottom left side of the block, however, these embodiments are not limited to these expansions.

[0186] Embodiment 9 As an example, Figures 24A-24B and 25A-25B show the example of Figures 13A and 13B adapted to a template that is further extended toward the upper-left side of the block and the lower-right side of the block. This form of extended template can advantageously be used in a coding / decoding order different from the conventional coding / decoding order, i.e., left-to-right and top-to-bottom, such as the fixed coding / decoding order of VVC and ECM. Thus, the coding / decoding order can be switched horizontally, from left to right or right to left, at a given macroblock level. However, it should be understood that this type of template can also be used in the conventional left-to-right and top-to-bottom coding / decoding order.

[0187] A first example of using such an expanded template is depicted in Figures 24A-24B. Figure 24A shows an expanded template (1401) that includes a line of pixels above the current block WxH block (1400) and additional columns of pixels on both the left and right sides of block (1400). More precisely, in Figure 24A, the expanded template includes a line of pixels above the current block (1400) (2ew -W)×n a A block and two n on the left and right l ×e H and a block.

[0188] In one example, e w = 2W + 4 and e H = 2H + 4. However, different values ​​can be used. A template (1401) for a given W x H block (1400) has n b ∈[|0,e H -H|] Unavailable decoded pixel in line (1403) and n to the right r ∈[|0,e w -W|] column (1402) and the unavailable decoded pixels in (1404). The encoding / decoding order is from left to right, H n l The decoded pixels of (1400) are always unavailable. The template is first flattened to yield (1405). The height e of the right side of the block (1400) H and width n l Note that the rectangle of is finally flattened. The product (1408) of the flattened template (1405) and the weight matrix (1407) then uses the set L of indices of unavailable decoded pixels (1406) in the flattened template to skip multiplications and additions.

[0189] Another example of using such an expanded template is depicted in Figures 25A-25B. Figure 25A shows an expanded template (1501) that includes a line of pixels above the current block WxH block (1500) and additional columns of pixels on both the left and right sides of block (1500). More precisely, in Figure 25A, the expanded template extends (2e w -W)×n a A block and two n on the left and right l ×e H and a block.

[0190] In this example, e w = 2W + 4 and e H = 2H + 4. The template (1501) of a given W × H block (1500) has n b ∈[|0,e H -H|] Unavailable decoded pixel in line (1503) and n pixels to its left r ∈[|0,e w -W|] column (1502) and the unavailable decoded pixels. The encoding / decoding order is from right to left, and H n l The decoded pixels of (1500) are always unavailable. The template is first flattened to yield (1505). Again, the height e H and width n l The rectangle is finally flattened. The multiplication (1508) of the flattened template (1505) with the weight matrix (1507) then uses a set L of indices of unavailable decoded pixels (1506) in the flattened template to skip multiplications and additions.

[0191] Embodiment 10 The special feature of inter slices is that they contain intra-predicted CUs and inter-predicted CUs. The decoding of a given CTU in an inter slice is broken down into three steps. In the first step, called parsing, all bits of the syntax associated with this CTU are read from the bitstream. In the second step, called decoding, these bits are interpreted. For example, bits associated with the prediction of a given CU are interpreted as intra / inter mode. In the third step, the pixels of the CTU are reconstructed. For an inter-predicted CU, prediction requires only decoded pixels from an already decoded reference frame. In contrast, for an intra-predicted CU, prediction requires only decoded pixels located above and to the left of the current CU. Therefore, for a given CTU in an inter slice, the third step can be performed in parallel for all inter-predicted CUs. However, for any intra-predicted CU, the third step must be performed after reconstructing the pixels in the CUs located above and to the left of it. Knowing this, in some decoders, for a given CTU in an inter slice, the third step is parallelized for all inter-predicted CUs. Then, at a given time step during the third step for the inter-predicted CU, the third step for the intra-predicted CU can begin. For simplicity, for a given CTU in an inter slice, the time step for starting the third step for the intra-predicted CU comes after the third step for all inter-predicted CUs has been completed. Then, for an intra-predicted CU, prediction may access more decoded pixels than those located above and to the left of the CU.

[0192] For example, on the decoder side, for a CTU in an inter slice shown in Figure 26A, for an intra-predicted CU (shown with diagonal lines), prediction may access decoded pixels located all around this CU. For a current CTU in a current inter slice, the time step for starting reconstruction of pixels of the intra-predicted CU occurs after completing reconstruction of pixels of all inter-predicted CUs. Therefore, the dashed lines depict the decoded pixels that may be accessed by the prediction module for an intra-predicted CU.

[0193] As another example, on the decoder side, in a CTU in an inter slice shown in Figure 26B, for the leftmost CU that is intra predicted, prediction may access decoded pixels located above / below and to the left of it. For the leftmost CU that is intra predicted, the dashed lines depict the decoded pixels that may be accessed by the prediction module.

[0194] For inter-slices, given the decoder implementation mentioned above, the previously disclosed embodiments may be advantageously used with template-based tools specific to templates of intra-predicted blocks.

[0195] For example, Figures 26C and 26D depict an adaptation of the example of Figures 13A and 13B when the template-based tool is supplied with a template of a block to be predicted intra-interslice, the operation of interest in the template-based tool is vector-matrix multiplication, and the input template is first flattened. In addition, Figures 26C and 26D show an expanded template (1601) that includes the entire surrounding pixels of the current block WxH block (1600). More precisely, in Figures 26C and 26D, the expanded template is expanded (W+2*n l )×n a block and (W+2*n l )×n aA block and two n on the left and right l ×H block. Different values ​​can be used for the height and width of the blocks surrounding the current block (1600).

[0196] In a template (1601) for an intra-predicted W×H block (1600), e.g., the leftmost intra-predicted CU in Figure 26B, all decoded pixels to the right of the block (1602) are unavailable. The template is first flattened to produce (1603). The product (1606) of the flattened template (1603) with a weight matrix (1605) then skips multiplications and additions using a set L of indices of unavailable decoded pixels (1604) in the flattened template. In the template (1601), light gray squares indicate unavailable decoded pixels. In the weight matrix (1605), light gray areas contain weights that are not used because their calculations are skipped.

[0197] In various embodiments, the operation for which some parts / computations are skipped is vector-matrix multiplication. This is just one example. In all various embodiments, other operations may be considered for computation skipping, such as element-wise vector-vector multiplication, matrix multiplication of tensors, e.g., as implemented by "tf.matmul" in Tensorflow or "torch.matmul" in PyTorch, or the outer product between two vectors, e.g., as implemented by "numpy.outer" in Numpy. TensorFlow, PyTorch, and Numpy are libraries.

[0198] Various methods and other aspects described herein may be used to modify modules, such as intra-prediction modules (260, 360), of video encoder 200 and decoder 300, as shown in Figures 2 and 3. Furthermore, the aspects are not limited to ECM, VVC, or HEVC, but may be applied to, for example, other standards and recommendations, and extensions of any such standards and recommendations. Unless otherwise indicated or technically precluded, the aspects described herein may be used alone or in combination.

[0199] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes and the described aspects are not limited to these specific values.

[0200] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence to generate a final output suitable for display, for example. In various embodiments, such processes include one or more of the operations typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various implementations described herein, such as decoding resampling filter coefficients, resampling a decoded picture, or, for example, obtaining information to identify which pixels are unavailable in a template for a current block of a picture, applying a template-based tool using the obtained information to determine information to be used to decode the current block, and decoding the current block using the determined information.

[0201] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstruction picture process, including entropy decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process as a whole will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.

[0202] Various implementations involve encoding. Similar to the above discussion regarding "decoding," "encoding," as used herein, can encompass all or part of the processes performed on an input video sequence to generate, for example, an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as partitioning, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by encoders of various implementations described herein, such as determining resampling filter coefficients, resampling a decoded picture, or, for example, obtaining information to identify which pixels are unavailable in a template for a current block of a picture, applying a template-based tool using the obtained information to determine information to be used to encode the current block, and encoding the current block using the determined information.

[0203] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.

[0204] This disclosure has described various information, such as syntax, that may be transmitted or stored. This information may be packaged or arranged in a variety of ways, including, for example, ways common in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or slice header), or SEI message. Other ways are also available, including, for example, ways common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, such as that described in the RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport. b. DASH Media Presentation Description (MPD) descriptors, for example as described in DASH and transmitted over HTTP, which are associated with a representation or a set of representations to provide additional characteristics to the content representations. c. RTP header extensions, for example, as used during RTP streaming. d. ISO Base Media File Format, such as that used in OMAF and in some specifications, which uses boxes, which are object-oriented building blocks defined by a unique type identifier and length, also known as "atoms". e. HLS (HTTP Live Streaming) manifests transmitted over HTTP. A manifest can be associated with a version or collection of versions of content, for example to provide characteristics of the version or collection of versions.

[0205] Where a figure is presented as a flow diagram, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flow diagram of the corresponding method / process.

[0206] Some embodiments may refer to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is usually considered, often due to computational complexity constraints. Rate-distortion optimization is usually formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these approaches may be based on an extensive examination of all encoding options, including all considered modes or coding parameter values, but with a thorough evaluation of their coding costs and the associated distortion of the reconstructed signal after coding and decoding. To reduce encoding complexity, more rapid approaches may also be used, particularly with calculation of approximate distortion based on a predicted or prediction residual signal rather than a reconstructed one. A mixture of these two approaches may also be used, such as by using approximate distortion for only some of the considered encoding options and full distortion for others. Other approaches evaluate only a subset of the considered encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding cost and the associated distortion.

[0207] Implementations and aspects described herein may be implemented as, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features may also be embodied in other forms (e.g., an apparatus or a program). For example, an apparatus may be implemented in appropriate hardware, software, and firmware. The method may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0208] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment," or "in one implementation" or "in an implementation" in various places throughout this application, as well as other variations, are not necessarily all referring to the same embodiment.

[0209] Additionally, the application may refer to "determining" various information. Determining / determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0210] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0211] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves in some way, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0212] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded to include as many items as are listed, as would be apparent to one skilled in this and related arts.

[0213] Also, as used herein, the term "signal" specifically refers to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a specific one of multiple resampling filter coefficients or a coded block. Thus, in some embodiments, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder can transmit a specific parameter to a decoder (explicit signaling) so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling without transmission (implicit signaling) can be used to simply allow the decoder to know and select the specific parameter. By avoiding transmitting any actual function, bit savings are realized in various embodiments. It will be understood that signaling can be achieved in a variety of ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. While the above relates to the verb form of the term "signal," the term "signal" can also be used herein as a noun.

[0214] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. For example, such a signal can be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.

[0215] A number of embodiments have been described above, and the features of these embodiments may be provided singly or in any combination across various claim categories and types.

Claims

1. This involves obtaining information that identifies which pixels are unavailable within the current block template of the picture, Applying a template-based tool using the information identifying unavailable pixels to determine the information used to decode the current block, and applying the template-based tool using the acquired information, which includes skipping the calculation if the calculation involves pixels identified as unavailable by the acquired information. A method comprising: decrypting the current block using the determined information.

2. The method according to claim 1, further comprising flattening the template before applying the template-based tool.

3. The method according to claim 1 or 2, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining indices of all unavailable pixels.

4. The method according to claim 1 or 2, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining an index of all available pixels.

5. The method according to claim 1 or 2, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining a flag, each flag indicating whether the pixel in the template is available or unavailable.

6. The method according to claim 1 or 2, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining, for groups of adjacent unavailable pixels, the index of the first unavailable pixel and the index of the last unavailable pixel in the group.

7. The method according to claim 1 or 2, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining, for adjacent groups of available pixels, the index of the first available pixel and the index of the last available pixel in the group.

8. The method according to claim 1 or 2, further comprising spatially rearranging pixels in the template to increase the number of contiguous available pixels in memory.

9. The method according to claim 1 or 2, wherein the unavailable pixel is one of the following: a pixel that has not yet been reconfigured, a pixel that belongs to a different tile from the tile to which the current block belongs, or a pixel that is outside the picture boundary.

10. This involves obtaining information that identifies which pixels are unavailable within the current block template of the picture, Applying a template-based tool using the information identifying unavailable pixels to determine the information used to encode the current block, and applying the template-based tool using the acquired information, which includes skipping the calculation if the calculation involves pixels identified as unavailable by the acquired information. A method comprising: encoding the current block using the determined information.

11. The method according to claim 10, further comprising flattening the template before applying the template-based tool.

12. The method according to claim 10 or 11, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining an index of all unavailable pixels.

13. The method according to claim 10 or 11, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining an index of all available pixels.

14. The method according to claim 10 or 11, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining a flag, each flag indicating whether the pixel in the template is available or unavailable.

15. The method according to claim 10 or 11, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining, for adjacent groups of unavailable pixels, the index of the first unavailable pixel and the index of the last unavailable pixel in the group.

16. The method according to claim 10 or 11, wherein obtaining information to identify which pixels are unavailable within the current block template includes obtaining, for adjacent groups of available pixels, the index of the first available pixel and the index of the last available pixel in the group.

17. The method according to claim 10 or 11, further comprising spatially rearranging pixels in the template to increase the number of contiguous available pixels in memory.

18. The method according to claim 10 or 11, wherein the unavailable pixel is one of the following: a pixel that has not yet been reconfigured, a pixel that belongs to a different tile from the tile to which the current block belongs, or a pixel that is outside the picture boundary.

19. A decoding device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method according to claim 1 or 2.

20. An encoding device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method according to claim 10 or 11.

21. A computer program comprising program code instructions for implementing the method described in claim 1 or 2 when executed by a processor.

22. A computer-readable storage medium storing instructions for implementing the method described in claim 1 or 2.