Encoding and decoding methods using transforms suitable for L-shaped partitions and corresponding devices

By dividing the image blocks into L-shaped partitions and designing a special transformation matrix, the problems of high computational complexity and low compression efficiency in the prior art are solved, and more efficient video encoding and decoding are achieved.

CN120359745APending Publication Date: 2025-07-22INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085298.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-13
Filing Date
2023-11-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing video encoding technology has limited partitioning methods when processing image blocks, resulting in high computational complexity and insufficient compression efficiency. Especially when using non-square or rectangular blocks, the transformation matrix design is not applicable, which affects the encoding and decoding efficiency.

Method used

Partition the image block into at least two partitions, one of which is an L-shaped partition, and design transformations and inverse transformations are designed for L-shaped partitions to reduce computational complexity and improve compression efficiency.

Benefits of technology

Through L-shaped partitioning and specially designed transformation matrix, the computational complexity is reduced, encoding and decoding efficiency is improved, and video compression performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359745A_ABST
    Figure CN120359745A_ABST
Patent Text Reader

Abstract

An encoding method (correspondingly decoding method) is disclosed in which an image block to be encoded (correspondingly decoded) is partitioned into at least two partitions, at least one of which has an L-shape. Various configurations are defined based on the position of the L-shape in the image block. To reduce computational complexity, only a subset of configurations may be allowed. Transforms and inverse transforms are designed to be applied on such L-shaped partitions for encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of European Application No. 22306862.8, filed on December 13, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0003] At least one of the present embodiments generally relates to a method and apparatus for encoding (and correspondingly decoding) a picture block, and more particularly to a method and apparatus for encoding (and correspondingly decoding) a picture block divided into partitions. Background Art

[0004] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancies in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame picture correlations, and then the difference between the original block and the predicted block (usually denoted as prediction error or prediction residual) is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through inverse processes corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention

[0005] In one embodiment, an image block to be encoded (and correspondingly decoded) is divided into at least two partitions, at least one of the partitions having an L-shape. Various configurations are defined based on the position of the L-shape in the image block. To reduce computational complexity, only a subset of the configurations may be allowed. The transform and inverse transform are designed to be applied to such L-shaped partitions for encoding and decoding. Brief Description of the Drawings

[0006] Figure 1 The figure shows a block diagram of a system in which aspects of the present embodiment may be implemented;

[0007] Figure 2 The figure illustrates a block diagram of an embodiment of a video encoder;

[0008] Figure 3 The figure illustrates a block diagram of an embodiment of a video decoder;

[0009] Figure 4 The figure illustrates the principle of directional intra-frame prediction using reference neighbor samples;

[0010] Figure 5 The figure depicts the directional intra-frame modes defined in the common video coding and enhanced compression model;

[0011] Figure 6A and 6B The figure illustrates the horizontal and vertical partitioning of a luminance intra-frame prediction block into sub-partitions;

[0012] Figure 7A Depicts an example of four reference lines to be used by the multi-reference line (MRL) intra prediction process;

[0013] Figure 7B Depicts the set of all coding unit split patterns supported in VVC draft 6;

[0014] Figure 8 Depicts a flowchart of an encoding method according to an embodiment;

[0015] Figure 9 Illustrates partitioning a square block into two partitions with a top-left L-shaped partition according to an embodiment;

[0016] Figure 10 Depicts different configurations for partitioning a square block into two partitions according to an embodiment, one partition being an L-shaped partition;

[0017] Figure 11 Depicts different configurations for partitioning a rectangular block into two partitions according to an embodiment, one partition being an L-shaped partition;

[0018] Figure 12 Depicts different configurations for non-binary partitioning of a square block into two partitions according to an embodiment, one partition being an L-shaped partition;

[0019] Figure 13 Illustrates partitioning a square and rectangular block into three partitions with two L-shaped partitions according to an embodiment;

[0020] Figure 14 Illustrates a prediction process for a negative intra prediction direction in the case of a top-left configuration of an L-shaped partition according to an embodiment;

[0021] Figure 15 Illustrates a prediction process for a positive intra prediction direction in the case of a top-left configuration of an L-shaped partition according to an embodiment;

[0022] Figure 16 Illustrates a prediction process for a horizontal positive intra prediction direction in the case of a bottom-left configuration of an L-shaped partition according to an embodiment;

[0023] Figure 17 Illustrates a prediction process for a positive intra prediction direction in the case of a bottom-right configuration of an L-shaped partition according to an embodiment;

[0024] Figure 18 Illustrates a prediction process for a negative intra prediction direction in the case of a bottom-right configuration of an L-shaped partition according to an embodiment;

[0025] Figure 19Illustrates intra prediction of a planar pattern according to an L-shaped partition according to an embodiment;

[0026] Figure 20 Illustrates the forward transform process of a prediction residual block;

[0027] Figure 21A Depicts a flowchart of a method for an L-shaped block for encoding image data according to an example;

[0028] Figure 21B and 21C Illustrates the forward transform process of an L-shaped block according to an embodiment;

[0029] Figure 22 Illustrates the process of reordering transform coefficients of an 8×8 block;

[0030] Figure 23 and 24 Illustrates various scans of an L-shaped block quantizing transform coefficients;

[0031] Figure 25 Depicts a flowchart of a decoding method according to an embodiment;

[0032] Figure 26A Depicts a flowchart of a method for an L-shaped block for decoding image data according to an example;

[0033] Figure 26B and 26C Illustrates the inverse transform process of an L-shaped block according to an embodiment; and

[0034] Figure 27 Depicts the set of all coding unit splitting patterns according to an embodiment. Detailed Description

[0035] This application describes multiple aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail and are typically described in a way that may sound restrictive, at least to illustrate individual characteristics. However, this is for the purpose of clear description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide additional aspects. In addition, aspects can also be combined and interchanged with aspects described in earlier applications.

[0036] The aspects described and contemplated in this application can be implemented in many different forms. The following Figure 1 , Figure 2 and Figure 3 provide some embodiments, but other embodiments are also contemplated, and Figure 1 , Figure 2 and Figure 3The discussion is not limited to the scope of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing a bitstream generated according to any of the described methods.

[0037] Various methods are described herein, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as for example "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0038] This aspect is not limited to VVC (Versatile Video Coding), ECM (Enhanced Compression Model), or HEVC (High Efficiency Video Coding), and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC, ECM, and HEVC). Unless otherwise indicated or technically precluded, the aspects described in this application can be used alone or in combination.

[0039] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "encode" or "code" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side. Hereinafter, the terms "intra mode" and "intra prediction mode" are used interchangeably. The terms "directional intra prediction mode", "directional prediction mode", "directional intra mode", "directional mode", "angular mode", and "angular intra prediction mode" are used interchangeably.

[0040] Figure 1A block diagram illustrating an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more aspects described in the present application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 may be embodied singly or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in the present application.

[0041] System 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in the present application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., volatile memory devices and / or non-volatile memory devices). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0042] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded video or decoded video, and encoder / decoder module 130 may include its own processor and memory. Encoder / decoder module 130 represents the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software known to those skilled in the art.

[0043] The program code to be loaded onto the processor 110 or the encoder / decoder 130 to execute the various aspects described in the present application can be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 can store one or more of the various items during the execution of the processes described in the present application. The items so stored can include, but are not limited to, input video, decoded video or portions of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0044] In some embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET (Joint Video Exploration Team)).

[0045] As indicated in block 105, input can be provided to the elements of the system 100 through various input devices. Such input devices include, but are not limited to, (i) an RF (Radio Frequency) section that receives RF signals transmitted, for example, by a broadcaster over the air, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) Universal Serial Bus (USB) input terminals, and / or (iv) High-Definition Multimedia Interface (HDMI) input terminals. Figure 1 Other examples not shown include composite video.

[0046] In various embodiments, the input device of block 105 has corresponding input processing elements associated therewith, as is known in the art. For example, the RF section can be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal band to a band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various functions of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or a baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0047] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing - such as Reed-Solomon error correction - can be implemented, as needed, within, for example, a separate input processing IC or within processor 110. Similarly, as needed, various aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in combination with memory and storage elements to process the data stream as needed for presentation on an output device.

[0048] The various elements of system 100 can be provided within an integrated housing, within which the various elements can be interconnected using suitable connection means 115 and data can be transmitted therebetween, the connection means 115 being, for example, an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.

[0049] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented within, for example, a wired and / or wireless medium.

[0050] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, and the set-top box delivers data through an HDMI connection of input block 105. Still other embodiments use an RF connection of input block 105 to provide streaming data to system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0051] System 100 may provide output signals to various output devices, which include a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 may be used for a television, a tablet device, a laptop computer, a phone (mobile phone), or other devices. The display 165 may also be integrated with other components (e.g., as in a smart phone) or be separate (e.g., an external monitor for a laptop computer). In various embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of system 100. For example, a disc player performs the function of playing the output of system 100.

[0052] In various embodiments, control signals are communicated between system 100 and display 165, speaker 175, or other peripheral device 185 using signaling such as AV.Link, CEC, or other communication protocols enabling device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 can be integrated in a single unit with other components of system 100 in an electronic device, such as a television. In various embodiments, display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0053] For example, if the RF portion of input 105 is part of a separate set-top box, display 165 and speaker 175 can alternatively be separated from one or more other components. In various embodiments where display 165 and speaker 175 are external components, the output signal can be provided via a dedicated output connection, which includes, for example, an HDMI port, a USB port, or a COMP output.

[0054] Embodiments can be implemented by computer software implemented by processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments can be implemented by one or more integrated circuits. Memory 120 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, as a non-limiting example, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 110 can be of any type suitable for the technical environment and can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0055] Figure 2 An example video encoder 200, such as a VVC (Versatile Video Coding) encoder, is illustrated. Figure 2 An encoder that improves on the VVC standard or an encoder that employs VVC-like techniques can also be illustrated.

[0056] Before being encoded, the video sequence can undergo pre-encoding processing (201), for example, applying a color transformation to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the preprocessing and attached to the bitstream.

[0057] In encoder 200, as described below, pictures are encoded by encoder elements. The picture to be encoded is partitioned (202) and processed in units such as CUs (Coding Units). Each unit is encoded using, for example, an intra or inter mode. When a unit is encoded in the intra mode, it performs intra prediction (260), for example, using intra prediction tools such as decoder-side intra mode derivation (DIMD). In the inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of the intra mode or inter mode to use to encode the unit and indicates the intra / inter decision, for example, by a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.

[0058] The prediction residual is then transformed (225) and quantized (230). Video coding standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Enhanced Compression Model (ECM 6.0) support different types of block transforms, such as DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform), which have been designed for square or rectangular blocks. These transforms are typically applied separately to the prediction residual blocks obtained after intra or inter prediction.

[0059] The quantized transform coefficients, along with motion vectors and other syntax elements such as picture partition information, are entropy encoded (245) to output a bitstream. The encoder may skip the transform and directly quantize the untransformed residual signal. The encoder may bypass both the transform and quantization, i.e., the residual is directly encoded without applying the transform or quantization process.

[0060] The encoder decodes the encoded blocks for use as a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The image block is reconstructed by combining (255), for example, adding, the decoded prediction residual and the prediction block. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Offset) / ALF (Adaptive Loop Filter) filtering to reduce coding artifacts. The filtered image is stored in the reference picture buffer (280).

[0061] Figure 3 The block diagram of an example video decoder 300 is illustrated. In decoder 300, the bitstream is decoded by decoder elements as described below. Video decoder 300 generally performs a decoding process reciprocal to the encoding process as described in Figure 2 what is described.

[0062] Encoder 200 generally also performs video decoding as part of encoding video data.

[0063] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, prediction modes, motion vectors, and other coding information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition (335) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and the prediction blocks are combined (355), for example, added, to reconstruct the image blocks. The prediction blocks can be obtained (370) from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375). The in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (380). Note that for a given picture, the content of the reference picture buffer 380 on the decoder 300 side is the same as the content of the reference picture buffer 280 on the encoder 200 side for the same picture.

[0064] The decoded picture can further undergo post-decoding processing (385), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse of the remapping process performed in the precoding process (201). The post-decoding processing can use the metadata derived in the precoding process and signaled in the bitstream.

[0065] In VVC and ECM, intra prediction is applied to full intra frames (i.e., frames consisting only of intra blocks) as well as intra blocks in inter frames, where spatial prediction coding units (CUs) are performed from causal neighbor blocks in the same frame, i.e., the top and upper right blocks, the left and lower left blocks, and the upper left block. Based on the decoded pixel values in these blocks, the encoder constructs different predictions for the current block to be encoded (also called the target block) and selects the prediction that results in the best rate-distortion (RD) performance. On the decoder side, based on the decoded pixel values in the causal neighbor blocks, a single prediction is obtained for the target block, i.e., the block to be decoded. The single prediction is the prediction corresponding to the intra prediction mode selected and encoded by the encoder.

[0066] In other words, intra prediction (260, 360) is used to remove the correlations within a local region of the picture. The basic assumption for intra prediction is that the texture of the current picture region is similar to the texture in the local neighborhood (e.g., the picture blocks adjacent to the current region), and thus can be predicted from there. Direct neighbor samples are typically used for prediction, i.e., the samples from the sample row above the current block to be encoded (correspondingly decoded) and the last column of the reconstructed block to the left of the current block. The samples used to predict the current block belong to the causal neighborhood, i.e., they are available (and thus have been reconstructed) when encoding or decoding the current block.

[0067] The reference neighbor samples for predicting the current block depend on the intra prediction mode and may depend on the direction indicated by the intra prediction angle of the corresponding intra prediction mode. Figure 4 An illustration of directional intra prediction using its reference neighbor samples is shown. For example, for horizontal prediction (case (a)), the reference neighbor samples from the left column are directly used; for vertical prediction (case (c)), the reference neighbor samples from the upper row are directly used; for diagonal lower right prediction (case (b)), the reference neighbor samples from the upper left side are applied, and for diagonal lower left prediction (case (d)), the reference neighbor samples from the upper right side are applied.

[0068] In the following sections, various tools for intra prediction in the Enhanced Compression Model (ECM) are detailed.

[0069] To capture arbitrary edge directions present in natural videos, the number of directional intra modes in Versatile Video Coding (VVC) and the Enhanced Compression Model (ECM) is extended from 33, as used in High Efficiency Video Coding (HEVC), to 65, as Figure 5 depicted, and the PLANAR and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and both luminance and chrominance intra prediction.

[0070] In VVC and ECM, the target block (i.e., the block to be encoded or decoded) has the option of intra prediction either by a first method (intra prediction for the entire CU) or by a second method (intra prediction using the Intra Sub-Partition (ISP) of the CU). In the first method, all target pixels are predicted simultaneously in a classical manner based on the reference samples of the entire CU. In the second method, the target CU is divided into, for example, two or four sub-partitions of equal size, and the sub-partitions are encoded (correspondingly decoded) in the prediction mode of the CU. That is, each sub-partition is encoded (correspondingly decoded) separately, where its own reference samples are used to predict its target pixels. When the sub-partitions are encoded (correspondingly decoded) sequentially, the sub-partitions can benefit from the availability of decoded samples from adjacent sub-partitions, which are the direct neighbors of the current sub-partition. In some cases, this can lead to better prediction and compression efficiency than the first method.

[0071] Intra Sub - Partition (ISP)

[0072] Both Versatile Video Coding (VVC) and Enhanced Compression Model (ECM 6.0) support Intra Sub-Partition Prediction (ISP), where a target block can be partitioned vertically or horizontally into two or four sub-partitions depending on the target block size, as shown in Table 1. The sub-partitions are encoded and decoded in sequence, where the target block is treated as a single Coding Unit (CU). All sub-partitions use the prediction mode of the target block (also called the parent coding unit) for intra prediction, and in the case of sequential processing, the decoded pixels in one sub-partition are used as reference samples for the intra prediction of the next sub-partition.

[0073] The sub-partitions have at least 16 pixels. Thus, a block of size 4×4 is not partitioned into sub-partitions, while a block of size 4×8 and 8×4 has only two partitions. Blocks of all other sizes have only four sub-partitions. The sub-partitions can be horizontal or vertical. A block of size 4×8 can have only two vertical partitions each of size 4×4, while a block of size 8×4 can have only two horizontal partitions each of size 4×4. Similarly, as another example, a block of size 4x16 can have four vertical sub-partitions each of size 4x4 or four horizontal sub-partitions each of size 1x16. Figure 6A and Figure 6B illustrates examples of the two possibilities.

[0074] Block Size Number of Sub - Partitions 4×4 1 4×8 and 8×4 2 All Other Cases 4

[0075] Table 1 - Number of Sub-Partitions Depending on Block Size

[0076] For the pixels in each of these sub-partitions, predictions are constructed using the decoded prediction mode of the parent CU. These predicted values are added to the decoded residual values, which are generated by entropy decoding the coefficients sent by the encoder and then dequantizing and inverse-transforming them. The inverse transform is applied at the sub-partition level, just as the forward transform is applied at the encoder. Except for the first sub-partition, the reconstructed pixel values of each sub-partition can be used to generate the prediction of the next sub-partition. The decoded pixels on the last row (horizontal split) or the last column (vertical split) can be used as the top or left reference array for the next sub-partition, respectively.

[0077] The sub-partitions are processed in normal order, regardless of the intra prediction mode and the split utilized. That is, the first sub-partition to be processed is the partition that contains the top-left sample of the CU and then continues sequentially down (horizontal split) or to the right (vertical split). The split type of the CU is signaled using bit ‘0’ (NO_SPLIT), or bits ‘10’ or ‘11’ (for HOR_SPLIT and VER_SPLIT, respectively).

[0078] For each intra-coded block, a flag (e.g., isp_flag) is signaled to indicate whether ISP is to be applied. Under the condition that isp_flag is true, another syntax (e.g., isp_mode) is further signaled to specify whether the split is vertical or horizontal.

[0079] Multi - Reference Line (MRL)

[0080] VVC and ECM also support intra prediction using multiple reference lines (MRL). The target block can choose to use any one of the first, second, and third reference lines that gives the best rate-distortion performance. The motivation for the MRL prediction mode is the observation that non-adjacent reference lines are mainly beneficial for texture patterns with sharp and strongly oriented edges. If the texture pattern is smooth, the MRL prediction mode is expected to be less useful. In Figure 7A Figure 8, an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E respectively. HEVC intra picture prediction uses the closest reference line (i.e., reference line 0). For example, in VVC, MRL intra prediction uses 2 additional lines (reference line 1 and reference line 2).

[0081] The index of the selected reference line is signaled with a flag of one bit (0) (e.g., mrl_idx) to indicate the first reference line, or with two bits (10 or 11) to indicate the second or third reference line respectively. In VVC and ECM, ISP is considered to utilize only the first reference line. Therefore, if the block has an MRL index other than 0, isp_flag is inferred to be 0, and thus it is not sent to the decoder. In this case, intra prediction is performed on the entire CU without any split. Therefore, isp_flag is resolved depending on whether the mrl_idx flag is 0.

[0082] In HEVC, VVC, and ECM, compression is done at the block level, rather than at the whole-image level or at the level of the entire frames of a sequence. Thus, a frame is partitioned into a set of non-overlapping blocks called coding tree units (CTUs), and then each CTU is compressed by sequentially scanning them. A CTU undergoes recursive partitioning into blocks called coding units (CUs), which undergo prediction before a transform is applied to the prediction residuals. In intra-frame, all CUs undergo intra-frame prediction based on previously decoded neighbor pixels in the same frame, while in inter-frame, a CU can have either intra-frame prediction or inter-frame prediction based on pixels in an adjacent region in a previously decoded frame. Due to the quadtree (QT), binary tree (BT), and ternary tree (TT) partitioning structures, these CUs can only have binary square shapes (in HEVC) or binary square or rectangular shapes (in VVC, ECM). More precisely, in VVC, a coding tree unit (CTU) is first partitioned by a quadtree structure, and then each quadtree leaf node can be further partitioned in a binary or ternary manner. As shown in Figure 7B , in addition to NO_SPLIT and quadtree split (QT_SPLIT), there are four split types in VVC: vertical binary split (BT_VER), horizontal binary split (BT_HOR), vertical ternary split (TT_VER), and horizontal ternary split (TT_HOR). The TT_HOR or TT_VER split (horizontal or vertical ternary tree split mode) involves partitioning a parent block into 3 sub-blocks (e.g., CUs), where the corresponding sizes in the direction of the considered spatial partition are equal to 1 / 4, 1 / 2, and 1 / 4 of the parent block size. In the same way, the sub-partitioning defined by the ISP tool in VVC and ECM has been designed for binary coding units (CUs) and is always rectangular or square depending on the target CU size.

[0083] Block transforms such as the discrete cosine transform (DCT) and the discrete sine transform (DST) have been included in standards such as JPEG, HEVC, VVC, ECM, etc. Whether they are directly applied to image data (such as in JPEG) or to prediction residuals (such as in HEVC, VVC, ECM, etc.), they provide high decorrelation and energy compaction properties, which ensure effective compression of visual data. In addition, since they are orthogonal transforms, the inverse transform can be easily obtained by transposing the forward transform matrix, thus eliminating the need for a separate inverse transform matrix. Considering hardware implementation, integer versions of these transforms obtained by scaling and rounding the original floating-point transforms are specified in the above-mentioned standards. In these standards, the transform is designed for square or rectangular blocks of binary size, and the transform operation is performed separately by multiplying with a right transform matrix and another left transform matrix. If other partitioning shapes that do not permit such directly separable operations, such as L-shaped partitioning, are defined, the transform matrix needs to be modified or redefined.

[0084] The CTU recursive partitioning that partitions the CTU into CUs or the ISP partitioning that partitions the CU into sub-partitions may sometimes result in sub-optimal partitioning, for example because it does not correspond to the underlying objects in those CUs. Therefore, extending the CTU recursive partitioning and the ISP partitioning to other types of partitioning such as L-shaped partitioning can improve the compression efficiency. In this context, it is necessary to design the transform and the inverse transform for such new types of partitioning, which will be used to encode or decode the residuals generated by intra / inter prediction.

[0085] In the following sections, the CTU recursive partitioning and the ISP partitioning are modified to improve the compression efficiency. More precisely, a new L-shaped partitioning is introduced. In an example, a parent CU (or a parent block) is divided into at least two partitions (i.e., two sub-CUs or two ISP sub-partitions), where one of the two partitions has an L-shape and the other has a square or rectangle, depending on whether the parent CU is square or rectangular respectively. This can be in the context of the CTU recursive partitioning that partitions the CTU into CUs or in the context of the intra prediction of sub-partitions (ISP). In an example, the L-shaped partition contains three-quarters of the samples, and the square or rectangular partition contains the remaining quarter of the samples of the parent CU. In another example, the parent CU is divided into at least two partitions, where more than one partition is an L-shaped partition. In an example, the partitioning is limited to the binary case, i.e., the side lengths of the L-shaped partition and the other rectangular or square partition are powers of 2. In other examples, it is possible to have partitions with non-binary lengths.

[0086] Figure 8 A flowchart of an encoding method according to an embodiment is depicted.

[0087] In step S100, the current block to be encoded (also referred to as a CU or a parent CU) is partitioned (also referred to as split or divided) into at least two partitions (also referred to as sub-CUs or simpler CUs, ISP sub-partitions or simpler sub-partitions, blocks or sub-blocks), and one of the at least two partitions is an L-shaped partition. In other words, the partition itself can be a CU in the context of CTU recursive partitioning, or a sub-partition in the context of ISP. The current block can be a square block of size NxN or a rectangular block of size NxM, where N is different from M, and N and M are positive integers. In an example, the current block is split into at least two partitions, one of the at least two partitions is an L-shaped partition, and the other is a rectangular or square partition, depending on the shape of the current block. In Figure 9 the example depicted on the left, since the current block to be encoded is square, partition B has a square shape. In Figure 9 another example depicted on the right, partition B has a rectangular shape because the current block to be encoded is rectangular. In a specific instance, it is assumed that the latter partition has at least 8 pixels such that CUs of sizes 4×8 and 8×4 can have this split. Figure 10 and Figure 11 illustrate different configurations of the L-shaped partition, whose names are defined relative to the corners of the current block that includes the L-shaped partition: (a) upper left, (b) lower right, (c) lower left, and (d) upper right. In the above example, the L-shaped partition is obtained by using a half split in the horizontal direction and a half split in the vertical direction. Thus, the L-shaped partition has three-quarters of the pixels, and the other rectangular or square partition contains one-quarter of the pixels of the current block. In the case where the width and height of the current block are powers of 2, the above split is binary, i.e., the lengths of all sides of the L-shaped partition A and the lengths of all sides of the rectangular or square partition B are powers of 2. In fact, the L-shaped partition has six sides. If the two largest sides have lengths of M and N, then two of the remaining sides have lengths equal to M / 2, and the remaining two sides have lengths equal to N / 2. In the case where the parent CU has binary lengths, M and N are powers of 2, and so are (M / 2) and (N / 2). Hereinafter, if the two largest sides of the L-shaped partition have equal lengths, i.e., if M = N, the L-shaped partition is referred to as symmetric; and otherwise, the L-shaped partition is referred to as asymmetric.

[0088] Figure 12 depicts a non-binary split for a square current block, where at least one side of partition A or B is not a power of 2. For a rectangular current block, a similar non-binary split is possible. In yet another example, the current block is partitioned into more than one L-shaped partition, for example, by recursively splitting the square or rectangular partition B, as Figure 13 depicted above. More precisely, in Figure 13Among them, the current block is partitioned into three partitions, and two of the three partitions are L-shaped partitions.

[0089] In step S102, at least two partitions are encoded. In the example, at least two partitions A and B are two sub-blocks obtained by recursively partitioning the CTU of the parent block. A is an L-shaped CU that can be intra- or inter-frame encoded, and B is a square or rectangular CU, or a square or rectangular block recursively partitioned into CUs. In another example, at least two partitions A and B are two ISP sub-partitions of an intra-frame parent CU. In this case, the L-shaped CU is intra-frame encoded. The encoding sequence order of the sub-partitions in the ISP of VVC or ECM is fixed. For horizontal splitting, the sub-partitions are processed from top to bottom, and for vertical splitting, the sub-partitions are processed from left to right. A similar method can be followed here, by first encoding the L-shaped sub-partition A (also referred to as block A hereinafter) and then encoding the sub-partition B (also referred to as block B hereinafter), regardless of the configuration type. In another example, the encoding sequence order of the sub-partitions can depend on the type of configuration, which is depicted in Figure 10 and Figure 11 For the upper-left configuration, the L-shaped block A is processed first, and then block B, because only block A has both its top and left reference lines available. Once block A is processed, the decoded pixels on the top and left of block B can be used as reference pixels for block B. For the lower-right configuration, either encoding sequence order can be followed, but processing block B first will make the decoded pixels available on all the top and left of block A, which is beneficial. For the lower-left and upper-right configurations, it may be preferable to process block A first, and then block B, because block B does not have reference samples on one side. Once block A is processed, block B can use the decoded pixels on its top or left together with the left or top reference samples of the CU depending on the configuration.

[0090] On the encoder side, for the current block, a configuration is selected from a configuration set based on RD optimization. To limit complexity, the number of configurations in the set can be restricted, for example, restricted to one or two configurations. The (one or more) configurations selected into the set can be fixed. As an example, in the case where only one configuration is allowed, the set can include only the top-left configuration, or in the case where two configurations are allowed, the set can include the top-left and bottom-right configurations. In another example applied to ISP, the selected configuration can depend on intra prediction, such as the intra prediction direction of the current block. By convention, in the case where the direction is from the top-right to the bottom-left or from the bottom-left to the top-right, the intra prediction is considered positive, and in the case where the direction is from the top-left to the bottom-right, the intra prediction is considered negative. Using this convention, if the intra prediction direction of the current block is negative, only the top-left or bottom-right configuration can be selected. Otherwise, if the intra prediction direction is positive, only the bottom-left or top-right configuration can be selected (depending on whether the prediction direction is horizontal or vertical, respectively).

[0091] To encode the L-shaped CU A, for example, the predicted residual is obtained by subtracting the predicted L-shaped CU generated by intra or inter prediction from the original L-shaped CU A. In the case of intra prediction, the intra prediction mode is associated with the L-shaped CU and can be a directional intra prediction mode (also called an angular prediction mode) or a non-directional prediction mode (also called a non-angular prediction mode), such as the DC or planar mode. In the case of inter prediction, the L-shaped CU A is predicted from samples in the decoded past or future frames. More precisely, motion estimation and compensation of the reference frames stored in the reference picture buffer are used to predict the L-shaped CU A. The predicted residual is usually but not necessarily transformed and quantized. Refer to Figure 21A and 21BExamples of specific transformation processes for L-shaped blocks are disclosed. The transform coefficients in the three quadrants of the L-shaped CU are quantized, e.g., where the quantization step size is associated with their frequency index (e.g., mapped to or corresponding to their frequency index). Subsequently, the quantized transform coefficients undergo a suitable scanning method before being entropy-coded in a bitstream (also referred to as coded data). The encoder reconstructs the coded L-shaped CU to provide a reference for further prediction, e.g., CU B for intra prediction. For this purpose, the quantized transform coefficients are dequantized and inverse-transformed to obtain the prediction residual. The L-shaped CU is reconstructed by combining (e.g., adding) the prediction residual and the predicted L-shaped CU. In the case where the square or rectangular partition B is a CU, i.e., in the case where it is not further recursively partitioned into multiple CUs, it can be directly encoded in a classical manner (i.e., by prediction, transformation, quantization, possible binarization, and entropy coding). In the case where the square or rectangular partition B is recursively split into multiple CUs, each of these CUs is encoded in the case of an L-shaped CU as disclosed above or in a classical manner in the case of a square or rectangular CU. In other examples, partition B can be encoded before CU A, in which case, in the specific case where the L-shaped CU A is intra-coded, the reconstructed samples from partition B can be used as a reference for encoding the L-shaped CU A.

[0092] To encode the L-shaped ISP sub-partition A (also referred to hereinafter as block A), the prediction residual is thus obtained, e.g., by subtracting the predicted L-shaped block A (produced by intra prediction) from the original L-shaped block A. The same intra prediction mode, i.e., the intra prediction mode selected for the current block (i.e., the parent CU), is used for both sub-partitions A and B. The intra prediction mode can be a directional intra prediction mode (also referred to as an angular prediction mode) or a non-directional prediction mode (also referred to as a non-angular prediction mode) such as DC or planar mode. The prediction residual is typically but not necessarily transformed and quantized. Refer to Figure 21A and 21B Examples of specific transformation processes for L-shaped blocks are disclosed. The transform coefficients in the three quadrants of the L-shaped CU are quantized, e.g., where the quantization step size is associated with their frequency index (e.g., mapped to or corresponding to their frequency index). Subsequently, the quantized coefficients undergo a suitable scanning method before being entropy-coded in a bitstream (also referred to as coded data). The encoder reconstructs the coded L-shaped block to provide a reference for further prediction, e.g., for sub-partition B (also referred to hereinafter as block A). For this purpose, the quantized transform coefficients are dequantized and inverse-transformed to obtain the prediction residual. The L-shaped block is reconstructed by combining (e.g., adding) the prediction residual and the predicted L-shaped block. The square or rectangular sub-partition B is encoded in a classical manner. In other examples, sub-partition B can be encoded before sub-partition A, in which case, the reconstructed samples from sub-partition B can be used as a reference for encoding the L-shaped sub-partition A.

[0093] Additional information (such as syntax elements) can be encoded. In addition to the quantized transform coefficients, this information can also include prediction modes (e.g., (one or more) intra prediction modes), motion vectors in the case of inter coding, and may include partitioning configuration information indicating how the current block is partitioned into at least two partitions. Additionally, this information can include, for example, an indication allowing L-shaped blocks. In an example, a syntax element can be encoded in a slice header to indicate that all CUs in a slice can use L-shaped splitting. In an example, a syntax element can be encoded in a PPS header to indicate that all CUs in a frame can use L-shaped splitting. In an example, a syntax element can be encoded in an SPS header to indicate that all CUs in all frames can use L-shaped splitting.

[0094] Directional Intra Prediction for L - shaped CUs or L - shaped Sub - Partitions

[0095] Hereinafter, we consider the upper left configuration, where L-shaped partition A is encoded first. As previously explained, at least two partitions A and B can be two child CUs generated by recursive partitioning of the CTU of the parent CU, or can be two ISP sub-partitions of an intra parent CU.

[0096] In the case of ISP, Figure 14 illustrates the prediction process for the negative prediction direction, and Figure 15 illustrates the prediction process for the positive prediction direction. In both cases, using the prediction mode of the parent CU, the encoder performs prediction of the samples of L-shaped partition A in the normal way, i.e., using the reference samples of the parent CU located on its top (2M + 1 samples) and its left (2N + 1 samples). L-shaped partition A is then encoded (e.g., by obtaining the transformed and quantized prediction residuals) and reconstructed.

[0097] Once partition A is encoded and reconstructed, the encoder uses the reconstructed samples in partition A to perform prediction of the samples in partition B. The reconstructed samples located at the top and left of partition B are used as reference samples. In Figure 15 the case of the positive prediction direction of, the lower left and upper right reference samples are filled. More precisely, the upper right reference sample is filled by copying the rightmost black pixel P, and the lower left reference sample is filled by copying the bottommost black pixel Q. As Figure 15 depicted above, partition B is predicted from N + 1 left reference samples and M + 1 upper reference samples.

[0098] Hereinafter, we consider the lower left configuration, where L-shaped partition A is encoded first. Figure 16Illustrated is the prediction process in the case where the prediction direction is the positive horizontal direction. In the case of the negative direction, one can simply use the decoded pixels from the L-shaped partition A on the left. In this case, using the prediction mode of the parent CU in the case of the ISP, the encoder performs the prediction of the samples in the L-shaped partition A in the normal way, that is, using the reference samples of the parent CU located at its top (2M + 1 samples) and its left (2N + 1 samples). As in the previous configuration, the L-shaped partition A is encoded and reconstructed. Once the partition A is encoded and reconstructed, the encoder thus uses the reconstructed samples in the partition A to perform the prediction of the samples in the partition B. The top reference sample for the partition B is obtained from the top reference sample of the parent CU. In one example, the decoded samples of the L-shaped partition A used to predict the partition B are the samples located to the left and directly below the partition B. In a particular implementation, so-called top and left reference arrays are used to predict any intra block. Such arrays are defined in VVC and ECM. Thus, in this implementation, the top reference array for the partition B includes the top reference samples of the parent CU (i.e., the M + 1 samples depicted as Figure 16 above). The left reference array includes the decoded samples of the partition A from the boundary ( Figure 16 gray above), i.e., to the left of the left edge of the partition B. If the prediction direction is positive, the remaining reference samples of the left reference array ( Figure 16 black above) are obtained by projecting the decoded samples below the partition B onto the left reference array, as shown in Figure 16 . More precisely, for each pixel position on the lower part of the left array, the position of the decoded sample on the bottom in the prediction direction (i.e., the decoded sample below the partition B) is determined. This sample may not match the decoded sample at an integer position but may be between two decoded samples. In this case, interpolation of the sample is performed (linear interpolation using 2 nearest neighbors, or cubic interpolation using 4 nearest neighbors (higher complexity but more accurate)). This is similar to obtaining the left part of the top reference array in the negative prediction direction in the conventional CU prediction. Thus, as depicted in Figure 16 above, the left reference array includes N + 1 sample values. In the case where the prediction direction is negative, i.e., from top left to bottom right, the availability of the decoded samples below the partition can be used for smoothing in a manner similar to PDPC (Position-Dependent Prediction Combination) defined in VVC. PDPC modifies the original prediction by using a weighted average of the original prediction and the reference samples of the target block to have a gradual intensity change at the top and left of the block. As in the previous configuration, once predicted, the partition B is encoded and reconstructed.

[0099] The case of first encoding the upper right configuration of the L-shaped partition A is similar to the case of first encoding the lower left configuration of the L-shaped partition A. More precisely, the upper right configuration in the case of the positive vertical direction is similar to the lower left configuration in the case of the positive horizontal direction.

[0100] For the lower right configuration, considering that partition A is processed first, followed by partition B, Figure 17 and Figure 18 illustrates the prediction processes for partition A and partition B. Figure 17 Illustrates the prediction process for the positive prediction direction, and Figure 18 illustrates the prediction process for the negative prediction direction. In this case, in the case of ISP, the prediction mode of the parent CU is used, and the encoder performs the prediction of the samples of L-shaped partition A in the usual way, that is, using the reference samples of the parent CU on its top (2M + 1 samples) and on its left (2N + 1 samples). As in the previous configuration, L-shaped partition A is encoded and reconstructed. Once partition A is encoded and reconstructed, the encoder then uses the reconstructed samples in partition A to perform the sample prediction of partition B. For partition B, the decoded samples on all four sides of the partition are available. If the prediction direction is positive, the decoded samples below and to the right of partition B are projected onto the left and top reference arrays, more precisely, onto the lower left and upper right parts of the arrays (in Figure 17 the black on Figure 17 ). The remaining reference samples on the left and top (in Figure 17 the white on Figure 18 are taken from the reference samples of the parent CU. Thus, the left reference array includes N + 1 sample values, and the top reference array includes M + 1 sample values, as depicted in

[0101] In the case of ISP regarding Figures 14 - 17 The disclosed examples can be extended to L-shaped CU A with square or rectangular CU B in the case where CU B is intra-coded. The difference from ISP is that the L-shaped CU can be intra-coded or inter-coded. In the latter case, the samples of the L-shaped CU are predicted from the samples of past or future frames. Additionally, if CU B is inter-coded or further partitioned, these examples will not apply.

[0102] Non - angular Intra Prediction Mode for L - shaped CUs or L - shaped Sub - Partitions

[0103] VVC and ECM include two non-angular intra prediction modes: the planar mode with index 0 and the DC mode with index 1. These two prediction modes model regions of slowly varying intensity in a frame. It is necessary to specify these two modes with an L-shaped partition so that they can be used, for example, with the L-shaped partitions in the ISP or with L-shaped CUs. In the following, we use the top-left configuration to illustrate these two modes. A similar approach can be followed in other configurations. Assume first that the L-shaped partition A (i.e., the L-shaped CU or L-shaped sub-partition) is encoded and reconstructed.

[0104] Intra prediction of the L-shaped partition A is performed using the reference samples of the parent CU in the usual way. If the prediction mode of the CU is DC, the DC value is calculated as usual using the top and left reference samples and the L-shaped partition is filled with this value. More precisely, in the case where the parent CU is square, the DC value is the mean sample value of the reference samples located to the left and above the L-shaped partition A. Otherwise (i.e., the parent CU is not square), the DC value is the mean of the samples on the larger side.

[0105] If the prediction mode is planar, prediction is performed in the usual way as the average of horizontal interpolation and vertical interpolation, where, for horizontal interpolation, the decoded samples in the top-right are repeated at the right edge, and for vertical interpolation, the decoded samples in the bottom-left are repeated at the bottom edge. More precisely, in the planar mode, the predicted sample value is obtained as a weighted average of 4 reference sample values. Here, reference samples in the same row or column as the current sample and reference samples in the bottom-left and top-right positions relative to the L-shaped partition are used. Interpolation is only performed on the L-shaped partition. This is illustrated in Figure 19 the top-left configuration. Horizontal and vertical interpolation are performed up to the edge of the L-shaped partition. For both the DC mode and the planar mode, the subsequent smoothing step using PDPC can be performed in the usual way.

[0106] Transform for L - shaped CUs or L - shaped Sub - Partitions

[0107] Figure 20 The forward transform of the prediction residual of a rectangular or square block (e.g., partition B) is illustrated. When the prediction residual of a sub-partition or CU in the ISP is transform-coded, two orthogonal transforms are typically applied for this purpose.

[0108] The right orthogonal transform T MxM (S1000) is applied to each row of the residual matrix of size NxM to obtain an intermediate data matrix. Then, the left orthogonal transform T t NxN(S1002) to obtain the final transform coefficient matrix. Since the transform specified in the standard is an integer version of the original transform, a scaling step is operated after each transform operation to reduce the coefficients in the working dynamic range. Before CABAC lossless entropy coding, the transform coefficients are quantized and then encoded in binary form, i.e., binarized. For decoding the prediction residuals, the inverse process is followed. After dequantization, the transform coefficients are inverse-transformed with the left and right inverse transform matrices, which are the transposes of the corresponding forward transform matrices. As in the forward transform, a scaling step can be applied after each inverse transform operation.

[0109] One solution for transforming the L-shaped partition A can be to insert zeros in the missing quadrants (i.e., the part corresponding to partition B) in order to create a square or rectangular block and then apply the right and left transforms in the usual way, e.g., as shown in Figure 20 the above figure. However, this solution results in large high-frequency coefficients, whose quantization can cause visible artifacts after the inverse transform operation. In addition, such a transform operation will generate M*N coefficients, while the L-shaped partition A has (3*M*N / 4) samples, thus giving a redundant over-complete representation. The proposed transform and inverse transform methods make it possible to limit or even avoid the above problems.

[0110] The DCT of an N-dimensional real vector x ≡ [x0 x1 x2...x N-1 is defined as X ≡ [X0 X1 X2...X N-1 where

[0111] for k = 0, 1, 2,... N-1.

[0112] The coefficient vector X is also a real vector. Thus, the DCT transform matrix is defined as T ≡ {t n,k} n=0...N-1;k=0...N-1 , where the (n, k)-th element of the matrix is given by

[0113]

[0114] The columns of matrix T are DCT basis vectors. The above transform operation can be equivalently written as X = xT, where T is a matrix of dimension N×N, and x and X are row vectors of size N defined as above.

[0115] Considering only the binary length, N is a power of 2. Now, considering only the even columns of matrix T, i.e., the columns with index k = 2*j, for these the elements of these columns can be expressed as follows:

[0116]

[0117] For

[0118] taking only the first elements from each column, we obtain a matrix of dimension whose elements are given by for and for This is the DCT matrix for dimension Thus, by taking only the first (N / 2) elements of each basis vector, the DCT basis vectors of dimension can be obtained from the even DCT basis vectors of dimension N.

[0119] The DCT vectors defined above are not normalized and, thus, they are orthogonal but not orthonormal. In fact, they are normalized by multiplying by such that the resulting transformation matrix is orthonormal. This allows the inverse transformation to be performed using the same matrix by taking its transpose. In this case, the DCT basis vectors of dimension can be obtained from the even DCT basis vectors of dimension N by taking only the first (N / 2) elements and scaling them by

[0120] The above observations help us to obtain the DCT coefficients of the -dimensional vector by using the even basis vectors of the NxN transformation matrix, by (1) padding the input vector, (2) multiplying by the even columns, and (3) scaling by Let y ≡ [y0 y1 y2... y N / 2-1 denote the input vector of dimension Let T even and T odd denote the matrices including the even and odd columns of the transformation matrix T, respectively, where the transformation matrix T is of dimension N×N. Thus, both T even and T odd have dimension Note that T even and T odd contain the corresponding columns in sequential order. Thus, the DCT coefficient vector can be obtained as follows where Here denotes the row vector of zeros. In the above operations, we have used parentheses to identify the three steps mentioned below:

[0121] (1) Padding with (1) Padding with zeros to obtain the input vector y p ;

[0122] (2) Multiply with the transformation matrix T even ; and

[0123] (3) Scale using to perform scaling.

[0124] A new transformation matrix is obtained by connecting the even transformation matrix and the odd transformation matrix as follows:

[0125] T c ≡ [T even T odd .

[0126] The new transformation matrix has dimensions NxN. The transformed coefficient vector Y can still be obtained by multiplying y p by T c and discarding the last p coefficients produced by multiplying y odd by T and scaling the remaining coefficients according to This results in obtaining the DCT transformation coefficients of the symmetric L-shaped block, as disclosed below. Any transformation with the above characteristics (e.g., the DCT type-II transformation specified in HEVC and VVC) can be used to derive the transformation for the L-shaped block. The characteristic is that the small transformation is included in the large transformation (twice the size) in a certain way. That is, one can obtain the small transformation matrix from the large transformation matrix.

[0127] Figure 21A Depicts a flowchart of a method for encoding an L-shaped block of image data according to an example.

[0128] In step S2000, the L-shaped block of image data is transformed into an L-shaped coefficient block by applying an orthogonal right transformation and an orthogonal left transformation.

[0129] In step S2010, the L-shaped block of coefficients is quantized using quantization weights to obtain an L-shaped block of quantized coefficients, where the quantization weights are associated with (e.g., mapped to) the frequency index of the coefficients. In other words, the coefficients are quantized using quantization weights associated with (e.g., corresponding to) the frequency index (or indices) of the coefficients.

[0130] In step S2020, the L-shaped block of quantized coefficients is encoded.

[0131] Figure 21B and 21CBoth illustrate the transformation (e.g., forward transformation) of the prediction residual (or image data if prediction is not used) of the L-shaped block in the upper-left configuration as applied in step S2000. By rotating the L-shaped block clockwise or counterclockwise, the L-shaped blocks in other configurations (i.e., upper-right, lower-left, and lower-right) can be mapped to (e.g., returned to) the upper-left configuration.

[0132] In step S2100, the missing quadrant of the L-shaped block is filled (or padded) with zeros to obtain a square or rectangular block. For the L-shaped block in the upper-left configuration, the missing quadrant is the lower-right, which is thus filled with zeros as illustrated in Figure 21B the above figure.

[0133] In step S2102, the right transformation (e.g., right forward transformation) corresponding to matrix T c is applied to the obtained block, i.e., the padded block. This right transformation is orthogonal. More precisely, the obtained block is right-multiplied by matrix T c to obtain an intermediate block of coefficients IM1.

[0134] In step S2104, the coefficients in the intermediate block IM1 that are located in the missing quadrant of the L-shaped block (the lower-right quadrant in the case of the upper-left configuration) are replaced with zeros (or set to zero).

[0135] In optional step S2106, in the case of a binary L-shaped block, the lower-left quadrant (in the case of the upper-left configuration) is scaled (i.e., multiplied) by 2. In the case of a non-binary L-shaped block, the scaling factor can be different and depends on the specific transformation and on the top width and bottom width of the L-shaped block. The scaling by 2 in step S2106 can be achieved by a left shift. Thus, the scaling by 2 in step S2106 avoids multiplying by in the inverse transformation. The scaling step S2106 can be included in other scalings, such as scaling due to the use of an integer transformation.

[0136] In step S2108, the left transformation (e.g., left forward transformation) corresponding to matrix is applied to the block obtained after S2104 or S2106, i.e., the scaled block. This left transformation is orthogonal. More precisely, the obtained block is left-multiplied by matrix to obtain the final block of coefficients.

[0137] In step S2110, the coefficients in the final block of coefficients that are located in the missing quadrant of the L-shaped block (the lower-right quadrant in the case of the upper-left configuration) are replaced with zeros (or set to zero) to obtain the L-shaped block of transformed coefficients ( Figure 21B shaded in the above figure).

[0138] The L-shaped block of transform coefficients is finally binary-coded, i.e., the coefficients are quantized and may be binaryized and entropy-coded, for example, by CABAC.

[0139] In Figure 21A and Figure 21B example, the right transform (S2102) is applied first, and then the left transform (S2108) is applied. In another example, the left transform may be applied first, and then the right transform may be applied.

[0140] Quantization and Entropy Coding of L - shaped CUs or L - shaped Sub - Partitions

[0141] With the proposed transform method, the transform coefficient block of the L-shaped block also has the same L-shape because all the elements in its lower right quadrant are zero and thus are not transmitted. After S2110, the transform coefficients of the L-shaped block are encoded. For this purpose, they are first quantized by the encoder before binary coding. After associating (e.g., mapping) the quantization step size (or equivalent quantization weight) with the frequency exponent of the coefficients, the quantizer used in orthogonal transform coding can be used. Since the new transform is obtained by concatenating even and odd basis vectors, the frequency coefficient index has a different order from that obtained by the orthogonal transform matrix used in HEVC, VVC, etc. For example, for an 8×8 block, the coefficient indices are sorted as shown in Figure 22 For any block with binary side lengths, square or rectangular, the sorting of the coefficient indices can be obtained similarly. The coefficients in the block identified by the dashed line are all zero and thus are not transmitted. The coefficient C i,j is quantized with a step size Q i,j where Q i,j denotes the quantization step size used for the (i, j)-th coefficient obtained by using the normal transform. Thus, the quantization weight used to quantize the coefficients is associated with the frequency exponent of the coefficients, in other words, the coefficient is quantized with the quantization weight Q i,j associated (e.g., corresponding) with the frequency exponent of the coefficient.

[0142] After quantization, the coefficients are scanned to map them to a 1D array. The scanning can be performed normally, except that the coefficients in the lower right quadrant are excluded. In video coding standards such as HEVC and VVC, the coefficients are scanned diagonally inside a 4×4 block group called a coefficient group (CG), and the CG itself is scanned diagonally inside a transform unit (TU). The same rule can also apply here. For example, Figure 23 shows the diagonal scanning pattern for a symmetric 8x8 L-shaped block. HEVC and VVC also specify horizontal and vertical scanning patterns for specific intra prediction modes. Similar scanning patterns can be applied to the transform coding of the L-shaped residual block, as in Figure 24As shown. Before lossless entropy coding, the scanned quantized coefficients are encoded in binary form, i.e., they may be binarized, for example, using CABAC. The binary coding of the coefficients based on the significant map can be performed as in VVC, ECM, etc., except that the significant map is calculated only for the CG in three quadrants of the L-shaped block.

[0143] Figure 25 FIG. depicts a flowchart of a decoding method according to an embodiment. The various embodiments / examples of the encoding method disclosed above are also applicable to the decoding method.

[0144] In step S200, the encoded data is obtained. The obtained encoded data is entropy decoded (inverse binarization may also be applied) to obtain information representing the current block (also referred to as CU or parent CU) to be decoded. This information includes, for example, quantized transform coefficients (hereinafter simply referred to as "transform coefficients"), prediction modes (e.g., (one or more) intra prediction modes), motion vectors in the case of inter-frame coding, and possibly partition configuration information that indicates how the current block is partitioned into at least two partitions.

[0145] In step S202, in response to the obtained information, at least two partitions (also referred to as sub-CUs or simpler CUs, ISP sub-partitions or simpler sub-partitions, blocks or sub-blocks) of the current block are reconstructed, one of the at least two partitions being an L-shaped partition. The L-shaped partition itself can be a CU in the context of CTU recursive partitioning or a sub-partition in the context of ISP.

[0146] In an example, at least two partitions A and B are two sub-blocks obtained by recursively partitioning a CTU of a parent block. A is an L-shaped CU that can be intra- or inter-frame coded, and B is a square or rectangular CU, or a square or rectangular block recursively partitioned into CUs. Thus, each partition has its own prediction mode. To decode the L-shaped CU A, a prediction residual is obtained by dequantizing and inverse-transforming the decoded transform coefficients of the L-shaped CU. The image of the L-shaped CU is reconstructed by combining (e.g., adding) the prediction residual and the predicted L-shaped CU. The predicted L-shaped CU is generated by intra- or inter-frame prediction. The prediction on the decoder side is the same as the prediction on the encoder side. In the case of intra-frame prediction, the intra-frame prediction mode is associated with the L-shaped CU and can be a directional intra-frame prediction mode (also referred to as an angular prediction mode) or a non-directional prediction mode (also referred to as a non-angular prediction mode), such as the DC or planar mode. The samples of the reconstructed L-shaped CU can be used as a reference for further prediction, such as for CU B in intra-frame prediction. In the case where the square or rectangular partition B is a CU, i.e., in the case where it is not further recursively partitioned into multiple CUs, the square or rectangular partition B can be directly decoded in a classical manner (i.e., by entropy coding, possible inverse binaryization, prediction, inverse quantization, and inverse transformation). In the case where the square or rectangular partition B is recursively split into multiple CUs, each of these CUs is decoded as disclosed below in the case of an L-shaped CU or in a classical manner in the case of a square or rectangular CU. The same principle applies to all L-shaped CUs in the CTU, while square or rectangular CUs are decoded in a classical manner. In other examples, partition B can be decoded before CU A, in which case, in the specific case where the L-shaped CU A is intra-frame coded, the reconstructed sample partition CU B can be used as a reference for decoding the L-shaped CU A.

[0147] In another example, at least two partitions A and B are two ISP sub - partitions of an internal parent CU. In this case, the L - shaped CU is intra - decoded and the same prediction mode is used for both A and B, i.e., the intra - prediction mode decoded for the parent CU. The prediction on the decoder side is the same as the prediction on the encoder side. To decode the L - shaped ISP sub - partition A (also referred to as block A hereinafter), the prediction residual is obtained by de - quantizing and inverse - transforming the decoded transform coefficients of the L - shaped sub - partition A. The reconstructed image L - shaped sub - partition is obtained by combining (e.g., adding) the prediction residual and the predicted L - shaped sub - partition. The same intra - prediction mode, i.e., the intra - prediction mode decoded for the current block (i.e., the parent CU), is used for both sub - partitions A and B. The intra - prediction mode can be a directional intra - prediction mode (also referred to as an angular prediction mode) or a non - directional prediction mode (also referred to as a non - angular prediction mode), e.g., DC or planar mode. The samples of the reconstructed L - shaped sub - partition A can be used as a reference for further prediction, e.g., for sub - partition B. The square or rectangular sub - partition B is decoded in a classical way. In other examples, sub - partition B can be decoded before sub - partition A, in which case the reconstructed samples from sub - partition B can be used as a reference for decoding the L - shaped sub - partition A.

[0148] Figure 26A A flowchart depicting a method for decoding an L - shaped block of image data according to an example is shown.

[0149] In step S2030, an L - shaped block of values (e.g., values, e.g., integer values, corresponding to the quantization coefficients obtained at the encoder side in S2010) is decoded.

[0150] In step S2040, the L - shaped block of values is de - quantized with a quantization weight to obtain an L - shaped block of reconstructed coefficients, where the quantization weight is associated with (e.g., mapped to) the frequency index of the values. In other words, the values are de - quantized with a quantization weight associated with (e.g., corresponding to) the frequency index (or indices) of the values.

[0151] In step S2050, the obtained L - shaped block of reconstructed coefficients is inverse - transformed into an L - shaped block of image data by applying an orthogonal left inverse transform and an orthogonal right inverse transform.

[0152] Figure 26B and 26C Both illustrate the inverse transform for an L - shaped block of reconstructed coefficients in the upper - left configuration as applied in step S2050. By rotating the L - shaped block clockwise or counter - clockwise, L - shaped blocks in other configurations (i.e., upper - right, lower - left, and lower - right) can be mapped to (e.g., reverted to) the upper - left configuration.

[0153] In step S2200, the missing quadrant of the L-shaped block of the reconstruction coefficients is filled (or padded) with zeros to obtain a square or rectangular block of the reconstruction coefficients. For an L-shaped block in the upper-left configuration, the missing quadrant is the lower-right.

[0154] In step S2202, a left inverse transform (corresponding to matrix T c ) is applied to the obtained block. More precisely, the obtained block is left-multiplied by matrix T c to obtain an intermediate block IM2 of the coefficients.

[0155] In step S2204, the coefficients in the intermediate block IM2 of the coefficients located in the missing quadrant of the L-shaped block (the lower-right quadrant in the case of the upper-left configuration) are replaced with zeros (or set to zero).

[0156] In optional step S2206, the upper-right quadrant is scaled by 2. More generally, in the case of a binary L-shaped block, the elements in the column containing the missing quadrant are multiplied by 2. In the case of a non-binary L-shaped block, the scaling factor can be different and depends on the top width and bottom width of the L-shaped block. In step S2206, the scaling by 2 can be achieved by a left shift. The scaling step S2206 can be included in other scalings, such as the scaling due to the use of an integer transform.

[0157] In step S2208, a right inverse transform (corresponding to matrix ) is applied to the block obtained in S2204 or S2206 (if any). More precisely, the obtained block is right-multiplied by matrix to obtain an inverse transform block of the prediction residual or the image data.

[0158] In step S2210, the coefficients in the final coefficient block located in the missing quadrant of the L-shaped block (the lower-right quadrant in the case of the upper-left configuration) are replaced with zeros (or set to zero) to obtain an L-shaped block of the prediction residual or the image data.

[0159] Examples of the forward and inverse transforms for an 8x8 L-shaped block (i.e., M = N = 8) are given. In VVC or ECM, the forward transform is applied to the prediction residual. However, they can also be applied directly to the image data. For the purpose of illustrating the above transformation process, the input 8x8 L-shaped block is an image data block. In the following example, the DCT transform specified in the HEVC standard is used to derive the new transform T c and The DCT matrix specified in HEVC has integer elements; thus, after each transform operation, there is a scaling that reduces the values within the working dynamic range. In the following, ">>" indicates a right shift and "<<" indicates a left shift.

[0160] The DCT8 matrix in the HEVC standard is

[0161]

[0162] Connect the even and odd columns, so the transformation matrix T c is defined as follows:

[0163]

[0164] The input L-shaped block after zero-padding (i.e., after S2100) is defined as follows:

[0165]

[0166] After step S2102, (right forward transformation (i.e., [ ]*T c ) and scaling (i.e., >>2, because the scaling factor = 2 -(8+3-9) = 2 -2 ), the intermediate block IM1 is as follows:

[0167]

[0168] The above scaling by 2 -2 is because an integer DCT transform is used. The scale factor is 2 -(B+M-9) , where B is the bit depth (which is 8 in this example), and M = log2(N), where N is the transform size. In the above example, N = 8, and thus M = 3.

[0169] After steps S2104 and S2106, the following blocks are obtained:

[0170]

[0171] After step S2108 (left forward transformation (i.e., ) and scaling (i.e., >>9, because the scale factor = 2 -(3+6) = 2 -9 ), the final block of transform coefficients is as follows:

[0172]

[0173] The above scaling by 2 -9 is because an integer DCT transform is used. The scale factor is 2 -(M+6) .

[0174] After step S2110, the L-shaped block of transform coefficients is as follows:

[0175]

[0176] Assuming the L-shaped block of transform coefficients is not quantized, the inverse transform process is as follows.

[0177] After step S2202 (left inverse transformation (i.e., T c *[ ]) and scaling (i.e., >>7, since the scale factor is 2 -7 ), the intermediate block IM2 is as follows:

[0178]

[0179] After steps S2204 and S2206, the obtained blocks are as follows:

[0180]

[0181] In S2208 (right inverse transformation (i.e., ) and scaling (i.e., >>12, since the scale factor is 2 -(20-8) = 2 -12 ), the inverse-transformed blocks are as follows:

[0182]

[0183] The above scaling by 2 -12 is due to the use of integer DCT transformation. The scale factor is 2 -(20-B) .

[0184] After S2210, the L-shaped blocks of the image data are as follows:

[0185]

[0186] The L-shaped block of the image data after S2210 is the same as the input L-shaped block of the image data.

[0187] For an asymmetric upper-left L-shaped block with length M on the left and length N on the top, after separating its even and odd columns in the same way, the right transformation is obtained from a DCT matrix of size NxN, and the left transformation is obtained from a DCT matrix of size MxM. The intermediate steps of scaling and zeroing remain unchanged.

[0188] In the following, different examples of application scenarios and their signaling are disclosed.

[0189] CTU Recursive Partitioning with L - shaped CUs

[0190] In VVC and ECM, the luminance and chrominance components can share the same coding tree, or the luminance and chrominance can each have their own tree (referred to as a dual tree). In the latter case, the luminance tree can be different from the chrominance tree.

[0191] In an example, for a CU partition of a CTU (luma CTU or chroma CTU or both luma and chroma CTUs), an L-shaped partition is added to the existing quadtree (QT), binary tree (BT), and ternary tree (TT) partitions defined in, e.g., VVC or ECM. That is, a CU is allowed to have an L-shape. In the example, to avoid redundancy, an L-shaped CU is not further split. However, the smaller square or rectangular CUs resulting from the L-shaped CU partition can undergo further splitting, including similar recursive L-shaped splitting.

[0192] In an example, to limit complexity, only the top-left split configuration is allowed. Figure 27 A set of all coding unit split patterns according to the example is depicted. In the case where a top-left L-shaped partition is added to the QT, BT, and TT partitions, an example of signaling can be as follows. A first bit is signaled to indicate whether the current block is split. If the first bit is 1, i.e., the current block is indicated as being split, then a second bit is signaled to indicate whether the QT is applicable. If the second bit is 0 (QT is not applicable), then the next bit is signaled to indicate whether L_SPLIT (i.e., L-shaped split) is applicable. If not (i.e., L_SPLIT is not applicable), then the next two bits are signaled to indicate whether BT_VER, TT_VER, BT_HOR, or TT_HOR is applicable. Another example can be as follows. A first bit is signaled to indicate whether the current block is split. If the first bit is 1, i.e., the current block is split, then the next bit b0 is signaled to indicate whether at least one of the QT or L_SPLIT is applicable or neither of them is applicable. If b0 is 1 (i.e., at least one of the QT or L_SPLIT is applicable), then the next bit b1 is signaled to indicate whether the QT or L_SPLIT is applicable (e.g., b1 is set to 1 to indicate the QT, and b1 is set to 0 to indicate the L_SPLIT, or vice versa). Otherwise, i.e., if b0 is 0 (i.e., neither the QT nor the L_SPLIT is applicable), then the next two bits are signaled to indicate whether BT_VER, TT_VER, BT_HOR, or TT_HOR is applicable.

[0193] In another example, as Figure 10 and Figure 11As depicted above, multiple split configurations (e.g., 2, 3, or 4) are allowed. In the case where all four L-shaped partitions are added to the QT, BT, and TT partitions, an example of signaling can be as follows. The first bit is signaled to indicate whether the current block is split. If the first bit is 1, i.e., the current block is indicated as split, then the second bit is signaled to indicate whether QT is applicable. If not (QT is not applicable), then the next bit b0 is signaled to indicate whether the L-shaped split is applicable. If b0 is 1 (i.e., the L-shaped partition is applicable), then the next two bits are signaled to indicate which configuration among top-left, top-right, bottom-left, or bottom-right is applicable. If b0 is 0 (i.e., the L-shaped partition is not applicable), then the next two bits are signaled to indicate whether BT_VER, TT_VER, BT_HOR, or TT_HOR is applicable.

[0194] The transform is applied to the prediction residuals generated by intra or inter prediction. The transform coefficients in three quadrants are quantized with a quantization step size associated with (e.g., mapped to) their frequency indices. Subsequently, the quantized coefficients undergo a suitable scanning method before being binary coded. In the example, to facilitate coefficient scanning based on groups of coefficients of size 4×4, as done in HEVC, VVC, ECM, etc., the minimum size of a CU supporting L-shaped split is assumed to be 8×8.

[0195] The same principle can be applied to chrominance CUs.

[0196] ISP Partitioning with L - shaped Sub - Partitions

[0197] In the example, L-shaped partitions are added in intra prediction with sub-partitions (ISP) for luminance CUs. In VVC or ECM, a CU with intra prediction can be split into two or four vertical or horizontal partitions, where the partitions are processed sequentially for prediction, and the resulting prediction residuals are encoded and decoded. The L-shaped partition allows the CU to be split into a sub-partition with an L-shape and another sub-partition with a square or rectangle. In the example, as in Figure 10 and Figure 11Depicted above, multiple split configurations (e.g., 2, 3, or 4) are allowed. In another example, to limit complexity, only one split is allowed; that is, smaller square or rectangular partitions are not further split. In the example, in addition to the existing horizontal and vertical splits, only one split configuration (the upper-left partition has an L shape) is allowed. The transform is applied to the prediction residual in the L-shaped partition. The transform coefficients in three quadrants are quantized with a quantization step size associated with (e.g., mapped to) its frequency index. Subsequently, the quantized coefficients undergo a suitable scanning method before being binary coded. In the example, to facilitate coefficient scanning based on groups of coefficients of size 4×4, as done in HEVC, VVC, ECM, etc., the minimum size of the parent CU supporting the ISP with L-shaped splits is assumed to be 8×8. The pixels in the L-shaped partition are decoded after adding the predicted value to the prediction residual, which is obtained after applying the inverse transform to the decoded prediction residual coefficients. The decoded pixels are then used as reference samples for intra prediction in smaller square or rectangular partitions.

[0198] In the following example, the signaling of split types in the ISP is detailed.

[0199] In the first example, intra prediction using the ISP in VVC or ECM is extended with partitions including an L shape. Only the upper-left partition configuration is allowed. The encoder checks the RD performance with all possible split types (including no split) and signals the best split with a binary coding scheme. The decoder decodes the split type. The signaling of split types in the ISP is changed. For example, the signaling can be done as '0' for NO_SPLIT, '10' for L_SPLIT, '110' for HOR_SPLIT, and '111' for VER_SPLIT, where L_SPLIT indicates the L-shaped partition. Intra prediction of the L-shaped sub-partition is done using the reference samples of the parent CU. Then, the decoded samples in the upper and left L-shaped sub-partitions are used as reference samples to complete intra prediction of the smaller sub-partitions. In the example, the minimum size of the smaller sub-partitions is assumed to be 8 pixels.

[0200] In a second example, the in - frame prediction using ISP as in VVC or ECM is extended with an L - shaped partition. The number of allowed L - shaped configurations can be 1, 2, 3, or 4. When the number of configurations is 1, only the top - left configuration is allowed. When the number of configurations is 2, the top - left configuration is allowed together with any one of the other three types of configurations. When the number of configurations is 4, all four L - shaped configuration types are allowed. The encoder checks the RD performance with all possible split types (including no split) and signals the best split with an appropriate binary coding scheme. The decoder decodes the split type. The signaling of the split type in the ISP changes according to the number of added L - shaped configurations. For example, when only one L - shaped split is allowed, the signaling can be completed as '0' for NO_SPLIT, '10' for L_SPLIT, '110' for HOR_SPLIT, and '111' for VER_SPLIT, where L_SPLIT indicates the L - shaped partition. Similarly, when all four L - shaped splits are allowed, the signaling can be completed as '0' for NO_SPLIT, '1000' for L_SPLIT_TOP_LEFT, '1001' for L_SPLIT_BOTTOM_RIGHT, '1010' for L_SPLIT_TOP_RIGHT, '1011' for L_SPLIT_BOTTOM_LEFT, '110' for HOR_SPLIT, and '111' for VER_SPLIT, where L_SPLIT_X indicates the type of the L - shaped split, and so on. The in - frame prediction for the L - shaped sub - partition is completed using the reference samples of the parent CU. Then, depending on the split type, the in - frame prediction of the smaller sub - partition is completed using the decoded samples in the L - shaped sub - partition and the reference samples of the parent CU. In the example, the minimum size of the smaller sub - partition is assumed to be 8 pixels.

[0201] In a third example, the intra prediction in the ISP, such as in VVC or ECM, is modified to replace the existing horizontal and vertical splits with L-shaped splits. The number of allowed L-shaped configurations can be 1, 2, or 4. When the number of configurations is 1, only the top-left configuration is allowed. When the number of configurations is 2, the top-left configuration is allowed together with any one of the other three types of configurations. When the number of configurations is 4, all four L-shaped configuration types are allowed. The encoder checks the RD performance with all possible split types (including no split) and signals the best split with a suitable binary coding scheme. The decoder decodes the split type. The split type signaling in the ISP changes according to the number of added L-shaped configurations. For example, when only one L-shaped split is allowed, the signal can be completed as '0' for NO_SPLIT and '1' for L_SPLIT, where L_SPLIT indicates the L-shaped split. Similarly, when all four L-shaped splits are allowed, the signaling can be completed as: '0' for NO_SPLIT, '100' for L_SPLIT_TOP_LEFT, '101' for L_SPLIT_BOTTOM_RIGHT, '110' for L_SPLIT_TOP_RIGHT, '111' for L_SPLIT_BOTTOM_LEFT, where L_SPLIT_X indicates the type of the L-shaped split, and so on. The intra prediction of the L-shaped sub-partitions is completed using the reference samples of the parent CU. Then, depending on the split type, the intra prediction of the smaller sub-partitions is completed using the decoded samples in the L-shaped sub-partitions and the reference samples of the parent CU. The minimum size of the smaller sub-partitions is assumed to be 8 pixels.

[0202] In a fourth example, the intra prediction using the ISP in VVC or ECM is extended to include L-shaped partitions. The number of allowed L-shaped partitions is two. The first partition has one L-shaped sub-partition and one square or rectangular sub-partition. The second partition has two L-shaped sub-partitions and one square or rectangular sub-partition. The second L-shaped sub-partition is obtained by splitting the square or rectangular sub-partition again. The two L-shaped sub-partitions can only have the top-left configuration. These two new partitions can replace the existing horizontal and vertical splits in the ISP, or in addition to these, they can be included in the ISP.

[0203] Accordingly, the signaling scheme is determined. When there are two L-shaped sub-partitions, the intra prediction of the first sub-partition is completed using the reference samples of the parent CU. Then, the intra prediction of the second sub-partition is completed using the decoded samples in the first L-shaped sub-partition on the left and at the top as reference samples. Then, finally, the intra prediction of the smaller sub-partitions is completed using the decoded samples in the second L-shaped sub-partition on the left and at the top as reference samples.

[0204] This aspect is not limited to ECM, VVC or HEVC, and can be applied to, for example, other standards and recommendations, as well as extensions of any such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application can be used alone or in combination.

[0205] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.

[0206] Various implementations involve decoding. As used in this application, "decoding" can include, for example, performing all or part of the process on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also or alternatively includes processes performed by the decoders of the various implementations described in this application, for example, decoding resampling filter coefficients, resampling the decoded picture.

[0207] As another example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstructed picture process including entropy decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a more general decoding process will be clear and is considered well understood by those skilled in the art.

[0208] Various implementations involve encoding. In a manner similar to the discussion above regarding "decoding", as used in this application, "encoding" can include, for example, performing all or part of the process on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as partitioning, differential encoding, transform, quantization, and entropy encoding. In various embodiments, such a process also or alternatively includes processes performed by the encoders of the various implementations described in this application, for example, determining resampling filter coefficients, resampling the decoded picture.

[0209] As another example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Based on the context of the specific description, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to a more general encoding process will be clear and is considered well understood by those skilled in the art.

[0210] The present disclosure has described various pieces of information that can be transmitted or stored, such as, for example, syntax. The information can be packaged or arranged in various ways, including, for example, ways common in video standards, such as putting the information into SPS (Sequence Parameter Set), PPS (Picture Parameter Set), NAL unit (Network Abstraction Layer), header (e.g., NAL unit header or slice header), or SEI message. Other ways are also available, including, for example, ways common for system-level or application-level standards, such as putting the information into one or more of the following:

[0211] a. SDP (Session Description Protocol), which is used to describe the format of a multimedia communication session for the purposes of session announcement and session invitation, for example, as described in the RFCs and used in combination with RTP (Real-Time Transport Protocol) transmission.

[0212] b. DASH MPD (Media Presentation Description) descriptor, for example, as used in DASH and transmitted over HTTP, where the descriptor is associated with a representation or a set of representations to provide additional characteristics to the content representation.

[0213] c. RTP header extension, for example, as used during an RTP stream.

[0214] d. ISO base media file format, for example, as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and a length (also referred to as an "atom" in some specifications).

[0215] e. HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can be associated with, for example, a version or a set of versions of the content to provide characteristics of the version or the set of versions.

[0216] When a figure is represented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0217] Some embodiments relate to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually with a constraint on computational complexity. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate-distortion optimization problem. For example, the method can be based on an extensive test of all encoding options, including all considered modes or encoding parameter values, and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, especially by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, such as by using approximate distortion only for some of the possible encoding options and full distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.

[0218] The implementations and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.

[0219] References to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation” and any other variations that occur throughout the application do not necessarily all refer to the same embodiment.

[0220] Additionally, this application may refer to “determining” various pieces of information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0221] Furthermore, this application may refer to “accessing” various pieces of information. Accessing information can include, for example, one or more of receiving information, (e.g., from a memory) retrieving information, storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0222] Additionally, this application may refer to "receiving" various pieces of information. Like "access", receiving is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Additionally, during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.

[0223] It will be appreciated that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", any use of the following " / " and "and / or" and "at least one of..." is intended to include only selecting the first-listed option (A), or only selecting the second-listed option (B), or selecting both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include only selecting the first-listed option (A), or only selecting the second-listed option (B), or only selecting the third-listed option (C), or only selecting the first and second-listed options (A and B), or only selecting the first and third-listed options (A and C), or only selecting the second and third-listed options (B and C), or selecting all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.

[0224] Furthermore, as used herein, the word "signal" refers, among other things, in particular to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a particular one of a plurality of resampling filter coefficients. Thus, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Therefore, for example, an encoder can transmit (explicitly signal) a particular parameter to a decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It will be appreciated that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0225] As will be clear to those of ordinary skill in the art, an implementation can generate multiple signals, which are formatted to carry information such as can be stored or transmitted. This information can include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.

[0226] Numerous embodiments have been described above. The features of these embodiments can be provided individually or in any combination across various claim categories and types.

[0227] In an example, a method of encoding an L-shaped image data block is disclosed, including:

[0228] transforming the L-shaped image data block into an L-shaped coefficient block by applying an orthogonal right transform and an orthogonal left transform;

[0229] quantizing the L-shaped coefficient block with a quantization weight to obtain an L-shaped quantized coefficient block, where the quantization weight is associated with (e.g., mapped to) the frequency index of the coefficient; and

[0230] encoding the L-shaped quantized coefficient block into encoded data.

[0231] In an example, transforming the L-shaped image data block into an L-shaped coefficient block includes:

[0232] padding the missing quadrant of the L-shaped block with zeros to obtain a first padded block;

[0233] applying the orthogonal right transform to the first padded block to obtain a first coefficient block;

[0234] replacing the coefficients of the first block located in the missing quadrant of the L-shaped block with zeros to obtain a second padded block;

[0235] applying the orthogonal left transform to the second padded block to obtain a second coefficient block;

[0236] replacing the coefficients of the second block located in the missing quadrant of the L-shaped block with zeros to obtain the L-shaped coefficient block.

[0237] In an example, wherein the image data is a prediction residual.

[0238] In an example, applying the orthogonal left transform to the first padding block includes multiplying the first padding block by a transform matrix that concatenates the even and odd columns of a discrete cosine transform.

[0239] In an example, applying the orthogonal right transform to the second padding block includes multiplying the second padding block by the transpose of the transform matrix that concatenates the even and odd columns of a discrete cosine transform.

[0240] In an example, a method of decoding an L-shaped image data block is disclosed, including:

[0241] decoding encoded data into an L-shaped value block (e.g., integer values);

[0242] dequantizing the L-shaped value block with quantization weights to obtain an L-shaped reconstruction coefficient block, where the quantization weights are associated with (e.g., mapped to) the frequency exponents of the values; and

[0243] inverse-transforming the obtained L-shaped reconstruction coefficient block into an L-shaped image data block by applying an orthogonal left inverse transform and an orthogonal right inverse transform.

[0244] In an example, inverse-transforming the obtained L-shaped reconstruction coefficient block includes:

[0245] padding the missing quadrant of the L-shaped reconstruction coefficient block with zeros to obtain a first padding block;

[0246] applying an orthogonal left inverse transform to the first padding block to obtain a first coefficient block;

[0247] replacing the coefficients of the first block in the missing quadrant of the L-shaped block with zeros to obtain a second padding block;

[0248] applying an orthogonal right inverse transform to the second padding block to obtain a second coefficient block; and

[0249] replacing the coefficients of the second block in the missing quadrant of the L-shaped block with zeros to obtain an L-shaped image data block.

[0250] In an example, the image data is a prediction residual.

[0251] In an example, applying an orthogonal left inverse transform to the first padding block includes multiplying the first padding block by a transform matrix that concatenates the even and odd columns of a discrete cosine transform.

[0252] In an example, applying an orthogonal right inverse transform to the second padding block includes multiplying the second padding block by the transpose of the transform matrix that concatenates the even and odd columns of a discrete cosine transform.

Claims

1. A method for encoding an L-shaped image data block, comprising: Transforming the L-shaped image data block into an L-shaped coefficient block by applying an orthogonal right transform and an orthogonal left transform; Quantizing the L-shaped coefficient block with a quantization weight to obtain an L-shaped quantized coefficient block, wherein the quantization weight is associated with the frequency index of the coefficient; And Encoding the L-shaped quantized coefficient block into encoded data.

2. The method according to claim 1, wherein, Transforming the L-shaped image data block into an L-shaped coefficient block includes: Padding the missing quadrant of the L-shaped block with zeros to obtain a first padded block; Applying the orthogonal right transform to the first padded block to obtain a first coefficient block; Replacing the coefficients of the first block located in the missing quadrant of the L-shaped block with zeros to obtain a second padded block; Applying the orthogonal left transform to the second padded block to obtain a second coefficient block; Replacing the coefficients of the second block located in the missing quadrant of the L-shaped block with zeros to obtain an L-shaped coefficient block.

3. The method according to claim 1 or 2, wherein The image data is a prediction residual.

4. The method according to claim 2 or 3, wherein, Applying the orthogonal right transform to the first padded block includes multiplying the first padded block by a transform matrix that concatenates the even and odd columns of a discrete cosine transform.

5. The method according to claim 4, wherein, Applying the orthogonal left transform to the second padded block includes multiplying the second padded block by the transpose of the transform matrix that concatenates the even and odd columns of a discrete cosine transform.

6. A method for decoding an L-shaped image data block, comprising: Decoding the encoded data into an L-shaped value block; Dequantizing the L-shaped value block with a quantization weight to obtain an L-shaped reconstructed coefficient block, wherein the quantization weight is associated with the frequency index of the value; And Inverse-transforming the obtained L-shaped reconstructed coefficient block into an L-shaped image data block by applying an orthogonal left inverse transform and an orthogonal right inverse transform.

7. The method according to claim 6, wherein Inverse-transforming the obtained L-shaped reconstructed coefficient block includes: Padding the missing quadrant of the L-shaped reconstructed coefficient block with zeros to obtain a first padded block; Applying an orthogonal left inverse transform to the first padded block to obtain a first coefficient block; Replacing the coefficients of the first block located in the missing quadrant of the L-shaped block with zeros to obtain a second padded block; Applying an orthogonal right inverse transform to the second padded block to obtain a second coefficient block; and Replacing the coefficients of the second block located in the missing quadrant of the L-shaped block with zeros to obtain an L-shaped image data block.

8. The method according to claim 6 or 7, wherein The image data is a prediction residual.

9. The method according to claim 7 or 8, wherein, Applying the orthogonal left inverse transform to the first padded block includes multiplying the first padded block by a transform matrix that concatenates the even and odd columns of a discrete cosine transform.

10. The method according to claim 9, wherein, Applying the orthogonal right inverse transform to the second padded block includes multiplying the second padded block by the transpose of the transform matrix that concatenates the even and odd columns of a discrete cosine transform.

11. An encoding device, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 1-5.

12. A decoding device, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 6-10.

13. A computer program comprising program code instructions for implementing the method of any one of claims 1 - 10 when executed by a processor.

14. A computer-readable storage medium having stored thereon instructions for implementing the method of any one of claims 1 - 10.