Non-separable conversion for low-latency applications

By selecting non-separable transforms and inverses based on encoding unit characteristics, the complexity of LFNST and NSPT is reduced, allowing their use in low-latency video encoding and maintaining compression efficiency.

JP2026511978APending Publication Date: 2026-04-14INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITALCE PATENT HLDG SAS
Filing Date
2024-03-25
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing video encoding technologies face challenges in utilizing advanced transformation tools like LFNST and NSPT in low-latency applications due to increased complexity, which hinders their implementation.

Method used

A method and apparatus that select non-separable transforms and inverses based on characteristics of encoding units, such as slice type, picture header elements, or causal neighborhood, to optimize their application in low-latency video encoding.

Benefits of technology

This approach reduces the complexity of LFNST and NSPT, enabling their effective use in low-latency video applications while maintaining compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511978000001_ABST
    Figure 2026511978000001_ABST
Patent Text Reader

Abstract

The method includes obtaining the current block of a picture, selecting a transform to apply to the current block from a selectable set of transforms, each of which includes at least one non-separable transform, and applying the selected transform to the current block. One of the at least one non-separable transforms is selectable for the current block based on the characteristics of the encoding unit containing the current block or an encoding unit adjacent to the current block, which is different from the sequence containing the picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims priority to European Patent Application No. 23305508.6, filed on April 6, 2023, the entire content of which is incorporated herein by reference.

[0002] At least one of the embodiments generally relates to a method and apparatus for improving the implementation of non - separable transforms in low - latency video encoding applications.

Background Art

[0003] To achieve high compression efficiency, video encoding schemes typically use prediction and transformation to exploit spatial and temporal redundancy in video content. During encoding, an image of video content is divided into blocks of samples (i.e., pixels), and these blocks are further divided into one or more sub - blocks, which are hereinafter referred to as "original sub - blocks". Next, intra - prediction or inter - prediction is applied to each sub - block to utilize the correlation within the image or the correlation between images. Regardless of the prediction method used (intra or inter), a predictor sub - block is determined for each original sub - block. And a sub - block representing the difference between the original sub - block and the predictor sub - block (often shown as a prediction error sub - block, a prediction residual sub - block, or simply a residual sub - block) is transformed, quantized, and entropy - encoded to generate an encoded video stream. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to transformation, quantization, and entropy - encoding.

[0004] One important aspect of video compression is the transformation of residuals from the spatial domain to the frequency domain. This aspect has been the subject of active research, particularly regarding the definition and implementation of the transformation. For example, recently, new transformation tools such as LFNST (Low-Frequency non-separable transforms) and NSPT (Non-Separable Secondary Transform) have been proposed to complement or replace the linear transformation commonly implemented in the form of DCT.

[0005] LFNST and NSPT have been shown to yield interesting coding gains. However, this coding gain comes with increased complexity, making it difficult to use these transformation tools in low-latency video applications.

[0006] It is desirable to propose solutions that can overcome the above problems. In particular, it is desirable to propose solutions that enable the use of LFNST and NSPT in low-latency video applications. [Overview of the Initiative]

[0007] In a first aspect, one or more of these embodiments provide a method comprising: obtaining the current block of a picture; selecting a transform to apply to the current block from a selectable set of transforms, each of which includes at least one non-separable transform; and applying the selected transform to the current block, wherein one of the at least one non-separable transforms is selectable for the current block based on the characteristics of an encoding unit containing the current block or an encoding unit adjacent to the current block that is different from the sequence containing the picture.

[0008] In a second aspect, one or more of these embodiments provide a method comprising: obtaining the current block of a picture; determining an inverse to apply to the current block from a selectable set of inverses, each of which includes at least one non-separable inverse, and applying the determined inverse to the current block, wherein one of the at least one non-separable inverse is selectable for the current block based on the characteristics of an encoding unit containing the current block or an encoding unit adjacent to the current block that is different from the sequence containing the picture.

[0009] In one embodiment of the first or second aspect, the coding unit is a slice containing the current block, and the characteristic is the slice type of the slice.

[0010] In one embodiment of the first or second aspect, the coding unit is the picture containing the current block, and the characteristic is the value of the picture header level syntax element in the picture header of the picture containing the current block.

[0011] In one embodiment of the first or second aspect, the coding unit is a slice containing the current block, and the characteristic is the value of a slice header level syntax element in the slice header of the slice containing the current block.

[0012] In one embodiment of the first or second aspect, the coding unit is the picture containing the current block, and the characteristic is a picture parameter set level syntax element in a picture parameter set that references the picture containing the current block.

[0013] In one embodiment of the first or second aspect, the coding unit is at least one block in the causal neighborhood of the current block, and the characteristic is a characteristic of at least one block in the causal neighborhood of the current block.

[0014] In one embodiment of the first or second aspect, the coding unit is a slice containing the current block, and the characteristic is the time depth of the slice containing the current block.

[0015] In a third aspect, one or more of these embodiments provide an apparatus comprising an electronic circuit configured to acquire a current block of a picture, select a transform to apply to the current block from a selectable set of transforms, each of which includes at least one non-separable transform, the non-separable transform being selectable for the current block based on the characteristics of a coding unit containing the current block or a coding unit adjacent to the current block, which is different from the sequence containing the picture.

[0016] In a fourth aspect, one or more of these embodiments provide an apparatus comprising an electronic circuit configured to acquire a current block of a picture, determine an inverse to apply to the current block from a selectable set of inverses, including at least one non-separable inverse, and apply the determined inverse to the current block, wherein one of the at least one non-separable inverse is selectable for the current block based on the characteristics of an encoding unit containing the current block or an encoding unit adjacent to the current block that is different from the sequence containing the picture.

[0017] In one embodiment of the third or fourth aspect, the coding unit is a slice containing the current block, and the characteristic is the slice type of the slice.

[0018] In one embodiment of the third or fourth aspect, the coding unit is the picture containing the current block, and the characteristic is the value of the picture header level syntax element in the picture header of the picture containing the current block.

[0019] In one embodiment of the third or fourth aspect, the coding unit is a slice containing the current block, and the characteristic is the value of a slice header level syntax element in the slice header of the slice containing the current block.

[0020] In one embodiment of the third or fourth aspect, the coding unit is the picture containing the current block, and the characteristic is a picture parameter set level syntax element in a picture parameter set that references the picture containing the current block.

[0021] In one embodiment of the third or fourth aspect, the coding unit is at least one block in the causal neighborhood of the current block, and the characteristic is a characteristic of at least one block in the causal neighborhood of the current block.

[0022] In one embodiment of the third or fourth aspect, the coding unit is a slice containing the current block, and the characteristic is the time depth of the slice containing the current block.

[0023] In a fifth aspect, one or more of these embodiments provide a computer program including program code instructions for performing the method according to the first or second aspect.

[0024] In a sixth aspect, one or more of these embodiments provide a non-temporary information storage medium for storing program code instructions for carrying out a method according to the first or second aspect. [Brief explanation of the drawing]

[0025] [Figure 1]Schematically shows an example of the context in which the embodiment is implemented. [Figure 2] Schematically shows an example of partitioning performed by the pixels of the original video picture. [Figure 3] Schematically shows a method for encoding video data. [Figure 4] Schematically shows a method for decoding video data. [Figure 5A] Schematically shows an example of the hardware architecture of a processing module that can implement an encoding module or a decoding module in which various aspects and embodiments are implemented. [Figure 5B] Shows a block diagram of an example of a first system in which various aspects and embodiments are implemented. [Figure 5C] Shows a block diagram of an example of a second system in which various aspects and embodiments are implemented. [Figure 6] Schematically represents an example of an embodiment that enables reduction of the complexity of LFNST and NSPT in an encoding module. [Figure 7] Schematically represents an example of an embodiment that enables reduction of the complexity of LFNST and NSPT in a decoding module. [Figure 8] Shows an example of a GoP (Group Of Picture) structure used in a low-delay encoding configuration.

Mode for Carrying Out the Invention

[0026] The following examples of embodiments are described in the context of video formats similar to VVC (ISO / IEC 23090-3-MPEG-I:VVC (Versatile Video Coding) / ITU-T H.266). However, these embodiments are not limited to video encoding / decoding methods corresponding to VVC. These embodiments are particularly adapted to various video formats, including, for example, HEVC (ISO / IEC 23008-2-MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265), AVC (ISO / IEC 14496-10), EVC (Essential Video Coding / MPEG-5), AV1, AV2, and VP9.

[0027] Figure 1 illustrates an example of a context in which the following embodiments may be implemented.

[0028] In Figure 1, System 11, which can be a camera, storage device, computer, server, or any device capable of distributing a video stream, transmits the video stream to System 13 using a communication channel 12. The video stream is either encoded and transmitted by System 11, or received and / or stored by System 11 and then transmitted. The communication channel 12 is a wired (e.g., Internet or Ethernet®) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.

[0029] System 13, which may be a set-top box for example, receives and decodes a video stream and generates a sequence of decoded pictures.

[0030] The decoded picture sequence is transmitted to the display system 15 using a communication channel 14, which may be a wired or wireless network. The display system 15 then displays the picture.

[0031] In one embodiment, system 13 is included in display system 15. In this case, system 13 and display system 15 are included in a TV, computer, tablet, smartphone, head-mounted display, etc.

[0032] Figures 2, 3, and 4 illustrate examples of video formats.

[0033] Figure 2 shows an example of partitioning performed by a picture 21 of pixels in the original video sequence 20. Here, we assume that each pixel consists of three components: one luminance component and two chrominance components. However, other types of pixels are also possible, including only a luminance component or fewer or more components such as additional depth or transparency components.

[0034] A picture is divided into multiple coding units. First, as shown by reference no. 23 in Figure 2, the picture is divided into a grid of blocks called coding tree units (CTUs). A CTU consists of N×N blocks of luminance samples and corresponding blocks of two chrominance samples. N is generally a power of 2, for example, with a maximum value of "128". Next, the picture is divided into groups of one or more CTUs. For example, it can be divided into one or more tile rows and tile columns, where a tile is a sequence of CTUs covering a rectangular area of ​​the picture. In some cases, a tile can be divided into one or more bricks, each brick consisting of at least one row of CTUs within the tile. Above the concepts of tiles and bricks, there is another coding unit called a slice, which can contain at least one tile or at least one brick of a tile in the picture.

[0035] In the example shown in Figure 2, as indicated by reference numeral 22, the picture 21 is divided into three slices S1, S2, and S3 in raster scan slice mode, each slice containing multiple tiles (not shown), and each tile containing only one brick.

[0036] As represented by reference numeral 24 in Figure 2, a CTU can be divided into a hierarchical tree of one or more subblocks called coding units (CUs). The CTU is the root (i.e., parent node) of the hierarchical tree and can be divided into multiple CUs (i.e., child nodes). Each CU becomes a leaf of the hierarchical tree if it is not further divided into smaller CUs, and becomes the parent node of the smaller CUs (i.e., child nodes) if it is further divided.

[0037] In the example in Figure 2, CTU24 is first divided into "4" square CUs using a quadtree-type partitioning. The top-left CU is a leaf in the hierarchical tree because it is not further divided; that is, it is not the parent node of the other CUs. The top-right CU is further divided into "4" smaller square CUs, again using a quadtree-type partitioning. The bottom-right CU is vertically divided into "2" rectangular CUs using a binary tree-type partitioning. The bottom-left CU is vertically divided into "3" rectangular CUs using a ternary tree-type partitioning.

[0038] During picture encoding, partitioning is performed adaptively, and each CTU is partitioned to optimize the compression efficiency based on the CTU standard.

[0039] HEVC introduced the concepts of prediction units (PUs) and transform units (TUs). In fact, in HEVC, the coding entities used for prediction (i.e., PUs) and transformation (i.e., TUs) can be subdivisions of a CU. For example, as shown in Figure 2, a CU of size 2N×2N can be divided into N×2N or 2N×N PU2411s. Furthermore, this CU can be divided into four N×N TU2412s or sixteen (N / 2)×(N / 2) TU2412s.

[0040] In VVC, except in a few specific cases, the boundaries of TUs and PUs are aligned with the boundaries of CUs. Therefore, a CU typically contains one TU and one PU.

[0041] In this application, the terms “block” or “picture block” may be used to refer to any one of CTU, CU, PU, ​​and TU. Furthermore, the terms “block” or “picture block” may be used to refer to macroblocks, partitions, and subblocks as specified in H.264 / AVC or other video coding standards, and more generally, to refer to arrays of samples of a large number of sizes.

[0042] In this application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, and the terms “image,” “picture,” “subpicture,” “slice,” and “frame” may be used interchangeably. Typically, though not always, the term “reconstructed” is used on the encoder side, and “decoded” is used on the decoder side.

[0043] Figure 3 schematically illustrates how the encoding module encodes a video stream. While variations of this encoding method are possible, the encoding method shown in Figure 3 is described below for clarity, without explaining all possible variations.

[0044] Before encoding, the current original picture of the original video sequence may undergo preprocessing. For example, in step 301, a color conversion may be applied to the current original picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or a remapping may be applied to the components of the current original picture to obtain a signal distribution that is highly compressible (e.g., using histogram equalization of one of the color components). The picture obtained through preprocessing will be referred to below as the preprocessed picture.

[0045] The encoding of the preprocessed picture begins with the division of the preprocessed picture in step 302, as described in relation to Figure 2. The preprocessed picture is divided into CTU, CU, PU, ​​TU, etc. The encoding module determines the encoding mode for each block, either intra-prediction or inter-prediction.

[0046] Intra-prediction in step 303 consists of predicting the pixels of the current block from predicted blocks derived from the pixels of the reconstructed blocks located around the current block to be encoded, according to the intra-prediction method. The result of intra-prediction is a prediction mode indicating which pixels of the surrounding blocks to use, and a residual block calculated as the difference between the current block and the predicted block.

[0047] Interpretation consists of predicting the pixels of the current block from a block of pixels (called a reference block) of a picture that precedes or follows the current picture (this picture is called a reference picture). During the encoding of the current block by the interpretation method, the block of the reference picture closest to the current block is determined by the motion estimation step 304 according to a similarity criterion. During step 304, a motion vector indicating the position of the reference block in the reference picture is determined. This motion vector is used during the motion compensation step 305, in which the residual block is calculated in the form of the difference between the current block and the reference block. In early video compression standards, the unidirectional interpretation mode described above was the only intermode available. As video compression standards have evolved, the family of intermodes has grown considerably and now includes many different intermodes.

[0048] During selection step 306, the coding module selects a prediction mode that optimizes compression performance from among the tested prediction modes (intra-prediction mode, inter-prediction mode) according to the rate / distortion optimization criterion (i.e., the RDO criterion).

[0049] When a prediction mode is selected, the residual block is transformed in step 307. In some embodiments, multiple types of transformations may be applied to the transformed residual block. In fact, in addition to DCT-II, a Multiple Transform Selection (MTS) scheme is used for both inter-prediction blocks and intra-prediction blocks. This scheme uses multiple transformations selected from DCT-VIII / DST-VII.

[0050] Another method for the transformation has been proposed, called the Low-frequency Non-Separable Transform (LFNST). The LFNST is applied between the forward linear transformation (i.e., the normal transformation in step 307) and quantization (corresponding to step 309). The LFNST is applied only to blocks where intra-prediction is selected.

[0051] LFNST defines multiple LFNST sets. Each set contains multiple translation kernels. For example, recent versions of LFNST define 35 LFNST sets, each containing three translation kernels.

[0052] Each LFNST kernel has a given dimension. For example, recent versions of LFNST define three kernels with the following dimensions: - LFNST4: 16×16 -LFNST8:32×64 -LFNST16:32×96

[0053] In LFNST, each LFNST set is associated with a range of INTRA modes. For example, DC and planar modes are associated with a single LFNST set.

[0054] Based on the block's INTRA mode, if an LFNST set is identified for an INTRA prediction block, the coding module has "4" possible options for that block: - No LFNST - LFNST using the first translation kernel of the identified LFNST set - LFNST using the second translation kernel of the identified LFNST set - LFNST using a third translation kernel for the identified LFNST set

[0055] Determining the best option for a block is based on rate-distortion optimization. The complexity of LFNST inherently stems from this rate-distortion optimization, which often hinders its use.

[0056] In recent years, it has been proposed to replace two-stage transformations, namely primary transforms (e.g., DCT-II) and secondary transforms (LFNST), with a single separable transform called an NSPT (Non-SeParable Secondary Transform). An NSPT defines multiple NSPT sets, each containing multiple transformation kernels. For example, in one embodiment, an NSPT defines 35 NSPT sets, each containing 3 transformation kernels. Similar to LFNSTs, each transformation kernel has a given dimension. The determination of the NSPT set for a block is based on the size of the block. Once the NSPT set is determined, the coding module has 4 possible options for the block: - No NSPT - NSPT by the first translation kernel of the identified NSPT set - NSPT by a second translation kernel for the identified NSPT set - NSPT by a third translation kernel for the identified NSPT set

[0057] Determining the best option for a block is based on rate-distortion optimization. Again, the complexity of NSPT inherently stems from this rate-distortion optimization. This complexity often hinders the use of NSPT.

[0058] The transformed blocks are quantized in step 309.

[0059] It should be noted that the encoding module can skip the transformation and apply quantization directly to the untransformed residual signal. When the current block is encoded according to the intra-prediction mode, in step 310, the intra-prediction mode and the transformed and quantized residual block are encoded by the entropy encoder. When the current block is encoded according to inter-prediction, the motion vector of the block is predicted, if necessary, from a prediction vector selected from a set of motion vectors corresponding to reconstructed blocks located in the vicinity of the block being encoded. The motion information is then encoded in step 310 by the entropy encoder in the form of an index for identifying the motion residual and prediction vector. The transformed and quantized residual block is encoded by the entropy encoder in step 310.

[0060] It should be noted that the encoding module can bypass both transformation and quantization; that is, entropy encoding is applied to the residual without the application of any transformation or quantization process. The result of entropy encoding is inserted into the encoded video stream 311.

[0061] Metadata such as SEI (Supplemental Enhancement Information) messages can be added to the encoded video stream 311. For example, an SEI message, as defined in standards such as AVC, HEVC, or VVC, is a data container associated with a video stream and contains metadata that provides information about the video stream.

[0062] After the quantization step 309, the current block is reconstructed so that the pixels corresponding to that block can be used for future predictions. This reconstruction phase is also called the prediction loop. Thus, in step 312, inverse quantization is applied to the transformed and quantized residual block, and in step 313, the inverse transform is applied. In step 314, the predicted block of the block is reconstructed according to the prediction mode used for the block obtained. If the current block is encoded according to the inter-prediction mode, the encoding module applies motion compensation in step 316, if necessary, using the motion vector of the current block to identify the reference block of the current block. If the current block is encoded according to the intra-prediction mode, in step 315, the prediction direction corresponding to the current block is used to reconstruct the predicted block of the current block. The predicted block and the reconstructed residual block are added together to calculate the reconstructed current block.

[0063] Following reconstruction, in step 317, in-loop filtering is applied to the reconstructed blocks with the aim of reducing encoding artifacts. This filtering is called in-loop filtering because it is performed in the prediction loop to obtain the same reference picture in the decoder as in the encoder, thereby avoiding drift between the encoding and decoding processes. In-loop filtering tools include deblocking filtering, SAO (Sample Adaptive Offset), and ALF (Adaptive Loop Filtering).

[0064] Once the block is reconstructed, in step 318 it is inserted into the reconstructed picture stored in memory 319, which is generally called the Decoded Picture Buffer (DPB). The reconstructed picture thus stored can then serve as a reference picture for other pictures to be encoded.

[0065] Figure 4 schematically illustrates how the decoding module decodes the encoded video stream 311, which has been encoded according to the method described in relation to Figure 3. While variations of this decoding method are possible, the decoding method shown in Figure 4 is described below for clarity, without explaining all possible variations.

[0066] Decoding is performed block by block. For the current block, it begins with entropy decoding of the current block in step 410. Entropy decoding allows us to obtain at least the predicted mode of the block.

[0067] If the blocks are encoded according to the interprediction mode, entropy decoding allows obtaining the predicted vector index, motion residual, and residual block at the appropriate time. In step 408, the motion vector is reconstructed for the current block using the predicted vector index and motion residual.

[0068] If the block is encoded according to the intra-prediction mode, entropy decoding allows obtaining the intra-prediction mode and the residual block. Steps 412, 413, 414, 415, 416, and 417 performed by the decoding module are in all respects identical to steps 412, 413, 414, 415, 416, and 417 performed by the encoding module, respectively. If LFNST is applied to the current block on the encoder side, the inverse LFNST transform is applied between the inverse quantization and the inverse linear transform (between steps 412 and 413). Similarly, if NSPT is applied to the current block on the encoder side, in step 413, the inverse NSPT transform replaces the linear and quadratic transforms (e.g., inverse DCT-II and inverse LFNST).

[0069] The decoded blocks are stored in the decoded picture, which is then stored in DPB419 in step 418. When the decoding module decodes a given picture, the picture stored in DPB419 is identical to the picture stored in DPB319 by the encoding module during the encoding of the given image. The decoded picture can also be output by the decoding module, for example, for display.

[0070] The post-processing step 421 may include reverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), reverse mapping which performs the reverse of the remapping process performed in the pre-processing step 301, and post-filtering to improve the reconstructed picture based on filter parameters provided in the SEI message, for example.

[0071] Figure 5A schematically shows an example of a hardware architecture of a processing module 500 that can implement an encoding module or a decoding module that can implement the encoding method of Figure 3 and the decoding method of Figure 4, respectively, modified according to different aspects and embodiments. The encoding module is included in system 11, for example, when the device is responsible for encoding a video stream. The decoding module is included in system 13, for example. The processing module 500 is connected by a communication bus 5005 and includes, in non-limiting examples, one or more microprocessors, general-purpose computers, dedicated computers, and processors based on multi-core architectures, a processor or CPU (Central Processing Unit) 5000, random access memory (RAM) 5001, read-only memory (ROM) 5002, and, without limitation, EEPROM (Electrically Erasable Programmable Read-Only Memory), ROM (Read-Only Memory), PROM (Programmable Read-Only Memory), RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), flash memory, magnetic disk drives, and / or optical disk drives, or SD (secure The storage unit 5003 may include non-volatile memory and / or volatile memory, including a storage medium reader such as a digital card reader and / or a hard disk drive (HDD) and / or a network-accessible storage device, and at least one communication interface 5004 for exchanging data with other modules, devices, or equipment. The communication interface 5004 may include, but is not limited to, a transceiver configured to transmit and receive data over a communication channel. The communication interface 5004 may include, but is not limited to, a modem or a network card.

[0072] If the processing module 500 implements a decoding module, the communication interface 5004 enables, for example, the processing module 500 to receive an encoded video stream and provide a sequence of decoded pictures. If the processing module 500 implements an encoding module, the communication interface 5004 enables, for example, the processing module 500 to receive a sequence of original picture data to be encoded and provide an encoded video stream.

[0073] The processor 5000 can execute instructions loaded into the RAM 5001 from the ROM 5002, from external memory (not shown), from a storage medium, or from a communication network. When the processing module 500 is powered on, the processor 5000 can read instructions from the RAM 5001 and execute them. These instructions form computer programs that cause the processor 5000 to execute, for example, the decoding method described in relation to Figure 4, the encoding method described in relation to Figure 3, and the methods described in relation to Figures 6 and 7. These methods include various aspects and embodiments described below in this specification.

[0074] All or part of the algorithms and steps of the methods shown in Figures 3, 4, 6, and 7 may be implemented in software form by executing a set of instructions by a programmable machine such as a DSP (digital signal processor) or microcontroller, or in hardware form by a machine or dedicated component such as an FPGA (field-programmable gate array) or ASIC (application-specific integrated circuit).

[0075] Thus, microprocessors, general-purpose computers, dedicated computers, processors based on or not based on multicore architectures, DSPs, microcontrollers, FPGAs, and ASICs are electronic circuits configured to at least partially implement the methods shown in Figures 3, 4, 6, and 7.

[0076] Figure 5C shows a block diagram of an example of a system 13 in which various aspects and embodiments are implemented. System 13 can be embodied as a device including various components described below and configured to perform one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and head-mounted displays. The elements of system 13 can be embodied individually or in combination as a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 13 comprises one processing module 500 that implements a decoding module. In various embodiments, system 13 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 13 is configured to implement one or more of the aspects described herein.

[0077] Inputs to the processing module 500 can be provided via various input modules, as shown in block 531. Such input modules include, but are not limited to, (i) a radio frequency (RF) module for receiving RF signals transmitted wirelessly, for example by a broadcasting station; (ii) a component (COMP) input module (or a set of COMP input modules); (iii) a USB (Universal Serial Bus) input module; and / or (iv) an HDMI (High Definition Multimedia Interface) input module. Other examples not shown in Figure 5C include composite video.

[0078] In various embodiments, the input module of block 531 is associated with each input processing element as is known in the Art. For example, an RF module may be associated with an element suitable for (i) selecting a desired frequency (also called selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band in order to select a signal frequency band which may (for example) be called a channel in a particular embodiment, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) multiplexing to select a desired stream of data packets. An RF module in various embodiments includes one or more elements for performing these functions, e.g., frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, e.g., down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one embodiment of a set-top box, the RF module and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments may rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF module includes an antenna.

[0079] Furthermore, the USB and / or HDMI modules may include their respective interface processors for connecting the system 13 to other electronic devices via USB and / or HDMI connections. It should be understood that various forms of input processing, such as Reed-Solomon error correction, can be implemented as needed, for example, within a separate input processing IC or within the processing module 500. Similarly, forms of USB or HDMI interface processing can be implemented as needed, within a separate interface IC or within the processing module 500. The demodulated, error-corrected, and demultiplexed streams are provided to the processing module 500.

[0080] Various elements of system 13 can be housed within an integrated housing. Within the integrated housing, the various elements can be interconnected using internal buses known in the art, such as Inter-IC (I2C) buses, wiring, and printed circuit boards, and data can be transmitted between them. For example, in system 13, processing module 500 is interconnected with other elements of system 13 by bus 5005.

[0081] The communication interface 5004 of the processing module 500 enables the system 13 to communicate on the communication channel 12. As already mentioned above, the communication channel 12 can be implemented, for example, in a wired and / or wireless medium.

[0082] In various embodiments, data is streamed to system 13 or provided in other ways using a wireless network such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 12 and a communication interface 5004 adapted for Wi-Fi communication. In these embodiments, the communication channel 12 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, stream data is provided to system 13 using an RF connection of input block 531. As described above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, for example, a cellular network or a Bluetooth® network.

[0083] System 13 can provide output signals to various output devices, including a display system 55, a speaker 56, and other peripheral devices 57. In various embodiments, the display system 55 includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display system 55 may be for a television, tablet, laptop, mobile phone, head-mounted display, or other device. The display system 55 may also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 57 include one or more of a standalone digital video disc (or digital multi-purpose disc) (DVR for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 57 that provide functions based on the output of System 13. For example, a disc player performs the function of playing back the output of System 13.

[0084] In various embodiments, control signals are communicated between system 13 and the display system 55, speaker 56, or other peripheral devices 57 using signaling such as AV.Link, CEC (Consumer Electronics Control), or other communication protocols that enable inter-device control with or without user intervention. Output devices can be communicably coupled to system 13 via dedicated connections through their respective interfaces 532, 533, and 534. Alternatively, output devices can be connected to system 13 using communication channel 12 via communication interface 5004, or using a dedicated communication channel corresponding to communication channel 54 in Figure 5A via communication interface 5004. The display system 55 and speaker 56 can be integrated into a single unit with other components of system 13 in an electronic device such as a television. In various embodiments, the display interface 532 includes a display driver, such as a timing controller (TCon) chip.

[0085] Alternatively, the display system 55 and speaker 56 can be separated from one or more of the other components. In various embodiments in which the display system 55 and speaker 56 are external components, the output signal can be provided via a dedicated output connection, such as an HDMI port, a USB port, or a COMP output.

[0086] Figure 5B shows a block diagram of an example of a system 11 in which various aspects and embodiments are implemented. System 11 is very similar to system 13. System 11 can be embodied as a device including various components described below and configured to perform one or more of the aspects and embodiments described herein. Such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, cameras, and servers. The elements of system 11 can be embodied individually or in combination as a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 11 comprises one processing module 500 that implements an encoding module. In various embodiments, system 11 is communicably coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 11 is configured to implement one or more of the aspects described herein.

[0087] The input to the processing module 500 can be provided through various input modules, as shown in block 531, which has already been described in relation to Figure 5C.

[0088] Various elements of system 11 can be housed within an integrated housing. Within the integrated housing, the various elements can be interconnected using internal buses known in the art, such as Inter-IC (I2C) buses, wiring, and printed circuit boards, and can transmit data between them. For example, in system 11, processing module 500 is interconnected with other elements of system 11 by bus 5005.

[0089] The communication interface 5004 of the processing module 500 enables the system 11 to communicate on the communication channel 12.

[0090] In various embodiments, data is streamed to system 11 or provided in other ways using a wireless network such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 12 and a communication interface 5004 adapted for Wi-Fi communication. In these embodiments, the communication channel 12 is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the RF connection of input block 531 is used to provide the stream data to system 11.

[0091] As described above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0092] The data provided to system 11 can be provided in different formats. In various embodiments, this data is encoded and conforms to known video compression formats such as AV1, VP9, ​​VVC, HEVC, and AVC. In various embodiments, this data is raw data provided, for example, by a picture and / or audio acquisition module connected to or included in system 11. In this case, the processing module is responsible for encoding this data.

[0093] System 11 can provide output signals to various output devices, such as System 13, which can store and / or decode output signals.

[0094] Various embodiments include decoding. As used in this application, “decoding” may encompass all or part of the processes performed on, for example, a received encoded video stream to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and prediction. In various embodiments, such processes may further, or alternatively, include processes performed by the decoder of the various implementations described in this application to decode, for example, the last significance coefficient of a block from the encoded video stream.

[0095] Whether the term "decryption process" is intended to refer specifically to a subset of operations or to the broader decryption process in general will become clear from the context of the specific explanation and should be easily understood by those skilled in the art.

[0096] Various embodiments include encoding. As with respect to “decoding” as described above, “encoding” as used in this application may encompass all or part of the processes performed on an input video sequence to generate an encoded video stream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, e.g., segmentation, prediction, transformation, quantization, and entropy coding. In various embodiments, such processes also, or alternatively, include processes performed by the encoder in various implementations described in this application to signal the last significance coefficient of a block in the encoded video stream, e.g.

[0097] Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to a broader encoding process in general will become clear from the context of the specific explanation and should be easily understood by those skilled in the art.

[0098] Please note that the syntax element names used herein are descriptive terms; therefore, they do not preclude the use of other syntax element names.

[0099] Please understand that when a diagram is presented as a flow chart, it also provides a block diagram of the corresponding device. Similarly, please understand that when a diagram is presented as a block diagram, it also provides a flow chart of the corresponding method / process.

[0100] Various embodiments refer to rate distortion optimization. In particular, a balance, or trade-off, between rate and distortion is usually considered during the coding process. Rate distortion optimization is typically formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, the approach can be based on extensive testing of all coding options, including all considered mode or coding parameter values, and involves a complete evaluation of their coding costs as well as the associated distortions of the reconstructed signals after coding and decoding. To reduce coding complexity, faster approaches can also be used, in particular, using calculations of distortion approximated based on predicted or predicted residual signals rather than reconstructed ones. For example, a mixture of these two approaches can be used, for example, by using approximate distortion for only some of the possible coding options and full distortion for the others. Other approaches evaluate only a subset of the possible coding options. More generally, many approaches use one of various techniques to perform the optimization, but the optimization does not necessarily fully evaluate both the coding cost and the associated distortions.

[0101] The embodiments and aspects described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if an embodiment of a described feature is described only in the context of a single embodiment (for example, discussed only as a method), it can also be implemented in other forms (for example, apparatus or programs). Apparatus can be implemented, for example, as appropriate hardware, software, and firmware. Methods can be implemented in, for example, processors, which generally refer to processing devices, including computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants (PDAs), and other devices that facilitate the communication of information between end users.

[0102] The phrases "one embodiment," "an embodiment," "one implementation," or "an implementation," as well as references to other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to an embodiment are included in at least one embodiment. Therefore, the phrases "in one embodiment," "in one embodiment," "in one implementation," or "in an implementation," appearing in various places throughout this application, as well as any other variations, do not necessarily all refer to the same embodiment.

[0103] Furthermore, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, retrieving information from memory, or obtaining information from other devices, other modules, or other users.

[0104] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0105] Furthermore, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these. Moreover, "receiving" typically involves, in some way, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information during the operation.

[0106] For example, in the cases of "A / B", "A and / or B", "at least one of A and B", and "one or more of A and B", please understand that the use of " / ", "and / or", "at least one of", and "one or more of" is intended to encompass the selection of only the first enumerated option (A), only the second enumerated option (B), or both options (A and B). As further examples, in the cases of "A, B, and / or C," "at least one of A, B, and C," and "one or more of A, B, and C," such phrasing is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or the selection of all three options (A, B, and C). This can be extended to the same number of items as enumerated, as will be apparent to those skilled in the art in this and related fields.

[0107] Furthermore, as used herein, the term “signal” refers, in particular, to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals the use of certain coding tools. In this way, in one embodiment, the same parameters can be used on both the encoder and decoder sides. Thus, for example, an encoder can send certain parameters to a decoder (explicit signaling) so that the decoder can use the same certain parameters. Conversely, if the decoder already has certain parameters and other parameters, signaling can be used without transmission (implicit signaling) simply to allow the decoder to recognize and select certain parameters. Bit saving is achieved in various embodiments by avoiding the transmission of arbitrary actual functions. It should be understood that signaling can be achieved in various ways. For example, in various embodiments, information is signaled to a corresponding decoder using one or more syntax elements, flags, etc. The above concerns the verb form of the word “signal,” but the word “signal” can also be used as a noun in this specification.

[0108] As will be apparent to those skilled in the art, embodiments can generate a variety of signals formatted to carry information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the embodiments described. For example, a signal may be formatted to carry an encoded video stream and SEI messages of the embodiments described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding an encoded video stream and modulating a carrier wave with an encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored in a processor-readable medium.

[0109] The following proposes various embodiments that reduce the complexity of LFNST and NSPT.

[0110] Figure 6 schematically illustrates an example of an embodiment that enables the reduction of the complexity of LFNST and NSPT in the coding module.

[0111] The process in Figure 6 is executed, for example, by the processing module 500 of system 11. The process in Figure 6 is executed, for example, before the conversion step 307.

[0112] In step 601, the processing module 500 obtains the current block to be transformed. The block is, for example, the residual block resulting from the intra prediction.

[0113] In step 602, the processing module 500 selects a transformation to apply to the current block from a set of selectable transformations, including at least one non-separable transformation, i.e., a set of transformations including LFNST and / or NSPT. For example, the set of transformations includes DCT-II, DCT-VIII, DST-VII, and LFNST and / or NSPT. In one embodiment, each of the selectable transformations is sequentially tested in rate-distortion optimization to determine the best transformation for the current block.

[0114] Figure 7 schematically illustrates an example of an embodiment that enables a reduction in the complexity of LFNST and NSPT in the decoding module.

[0115] The process in Figure 7 is executed, for example, by the processing module 500 of system 13. The process in Figure 7 is executed, for example, before the inverse transform 413 step.

[0116] In step 701, the processing module 500 obtains the current block to be inverted. The current block is, for example, the inverse quantization transformation residual.

[0117] In step 702, the processing module 500 determines which inverse transform to apply to the current block from a selectable set of inverse transforms, which includes at least one non-separable inverse transform, i.e., an inverse LFNST and / or an inverse NSPT. For example, the inverse transforms include an inverse DCT-II, an inverse DCT-VIII, an inverse DST-VII, an inverse LFNST and / or an inverse NSPT.

[0118] In step 703, the processing module applies the determined inverse transform to the current block.

[0119] In the first embodiment, a non-separable transformation is selectable for the current block based on the slice type of the slice containing the current block.

[0120] LFNST and NSPT are applied to residual blocks resulting from intra-prediction, and are therefore particularly well-suited to I-slice (an I-slice where all blocks are encoded intra-). However, using LFNST / NSPT is not very efficient for non-I-slice.

[0121] In the first modification of the first embodiment, LFNST and NSPT are selectable only for blocks contained in I-slices. In this case, LFNST and NSPT are deactivated for non-I-slices.

[0122] In a second modification of the first embodiment, LFNST and NSPT are selectable only for blocks contained in I-slices and P-slices (P-slices are slices in which blocks are encoded in either intra- or inter-unidirectional mode). In this case, LFNST and NSPT are deactivated for B-slices (slices in which biprediction is permitted).

[0123] In a third variation, high-level syntax elements are signaled, for example, in the sequence parameter set (SPS), to indicate that LFNST or NSPT is deactivated for non-I slices (or B slices).

[0124] Figure 8 shows an example of a Group of Picture (GoP) structure used in a low-latency coding configuration. In this example, consecutive pictures are assigned different quantization parameters (QPs), and the rate assignment between different pictures in the GoP is highly non-uniform.

[0125] More precisely, the QP of a picture is set according to a virtual hierarchical structure between pictures. This virtual hierarchical structure assigns a temporal depth value to each picture, and the QP assigned to a given picture depends on the picture's temporal depth value. For example, in Figure 8, picture type B3 represents a B picture with a temporal depth value of "3", and picture type B2 represents a B picture with a temporal depth value of "2". Figure 8 also shows the QP assigned to the temporal depth of each picture relative to the sequence level QP parameter.

[0126] Please note that time depth values ​​are not signaled in the video data and are only used on the encoder side to organize the GOP structure.

[0127] In a second embodiment, the use of non-separable conversion (LFNST or NSPT) is activated / deactivated at the picture level via an LFNST / NSPT flag override mechanism. This override mechanism includes signaling a flag in the picture header that indicates whether the LFNST / NSPT usage flag, which is signaled at the SPS level, is overridden for the current picture relating to the picture header. If overridden, the LFNST / NSPT usage flag is signaled in the picture header to indicate whether LFNST / NSPT is used in the current picture.

[0128] In a variation of the second embodiment, if the SPS-level LFNST / NSPT flag is signaled to be overridden at the picture level, the value of the picture header-level LFNST / NSPT usage flag is not signaled and is inferred as the reciprocal of the SPS-level LFNST / NSPT usage flag value.

[0129] In a variation of the second embodiment, the picture header-level LFNST / NSPT override flag is not signaled within the picture header, but the LFNST / NSPT usage flag is systematically encoded within the picture header.

[0130] In a variation of the second embodiment, the use of non-separable conversion (LFNST or NSPT) and MTS is activated / deactivated at the picture level via an LFNST / NSPT and MTS flag override mechanism. This override mechanism includes signaling a flag in the picture header that indicates whether the LFNST / NSPT and MTS usage flags, which are signaled at the SPS level, are overridden for the current picture relating to the picture header. If overridden, the LFNST / NSPT usage flag is signaled in the picture header to indicate whether LFNST / NSPT is used in the current picture. If overridden, explicit MTS and LFNST / NSPT usage flags are signaled in the slice header.

[0131] In a variation of the second embodiment, the override mechanism (in the case of LFNST / NSPT, or LFNST / NSPT and MTS) is performed at the slice header level. A slice header level flag indicates whether the SPS-level LFNST / NSPT usage flag (or the SPS-level LFNST / NSPT and MTS usage flag) is overridden in a given slice. If overridden, explicit MTS and LFNST / NSPT usage flags are signaled in the picture header.

[0132] In a variation of the second embodiment, the override mechanism (in the case of LFNST / NSPT, or LFNST / NSPT and MTS) is performed at the picture parameter set (PPS) level. The PPS level flag indicates whether the SPS level LFNST / NSPT usage flag (or the SPS level LFNST / NSPT and MTS usage flag) is overridden in the picture referencing this PPS. If it is overridden, the explicit MTS and LFNST / NSPT usage flags are signaled in the PPS.

[0133] As can be seen from the figure, in the second embodiment, non-separable conversions (LFNST and NSPT) can be selected in multiple conversions in steps 602 and 702 if: - The value of the syntactic element at the picture header level within the picture header of the picture containing the current block, or - The value of the syntax element at the slice header level within the slice header of the slice containing the current block, - The value of the PPS level sequence element in PPS that references the picture containing the current block. This indicates that LFNST and / or NSPT are currently permitted for the block.

[0134] In non-I slices, LFNST / NSPT can be restricted to specific blocks depending on the characteristics of the blocks in the causal neighborhood of the current block.

[0135] In the third embodiment, the non-separable transformations (LFNST and NSPT) can be selected from a plurality of transformations in steps 602 and 702 based on the characteristics of at least one block in the causal neighborhood of the current block.

[0136] For example, in the third embodiment, if most of the blocks in the causal neighborhood of the current block are intra-encoded, LFNST and NSPT become selectable for the current block, thereby limiting the use of LFNST / NSPT to only a small portion of blocks where intra-encoding is excessively used. For example, if more than half of the blocks in the causal neighborhood of the current block are intra-encoded, LFNST and NSPT become selectable for the current block.

[0137] In another example, if most of the blocks in the causal neighborhood of the current block are intercoded, then LFNST and NSPT become selectable for the current block. This restricts the use of LFNST / NSPT when encoding the current block is difficult and a significant encoding gain is expected with LFNST.

[0138] In a variation of the third embodiment, a high-level syntax element (for example, in SPS) indicates that the use of LFNST / NSPT is based on the properties of at least one block in the causal neighborhood of the current block.

[0139] In the first, second, and third embodiments, the complexity associated with LFNST and NSPT is limited by controlling the blocks to which LFNST and NSPT may be applied.

[0140] In a fourth embodiment, the complexity of LFNST and NSPT in the coding module is reduced by decreasing the number of possible choices during rate-distortion optimization. This can be achieved by reducing the number of LFNST / NSPT kernels considered (i.e., selectable) for the current block during rate-distortion optimization. As the number of kernels decreases, the corresponding signaling is modified.

[0141] [Table 1]

[0142] Table TAB1 shows the binarization of the LFNST / NSPT conversion indices for the LFNST / NSPT set. In this table, LFNST / NSPT index = 0 indicates that LFNST / NSPT is not used. LFNST / NSPT index = 1 indicates the use of the first LFNST / NSPT kernel. LFNST / NSPT index = 2 indicates the use of the second LFNST / NSPT kernel. LFNST / NSPT index = 3 indicates the use of the third LFNST / NSPT kernel.

[0143] As already mentioned above, if all conversions in the LFNST / NSPT set are selectable, then 2 bits are needed to encode the 4 possible options.

[0144] If the number of kernels is reduced to "1", one bit can be used. Otherwise, as shown in Table TAB1, the truncated unali code can be used for "2" kernels.

[0145] In a modified version of the fourth embodiment, the number of LFNST / NSPT kernels is reduced to "1" or "2" for non-I slices.

[0146] In another variation of the fourth embodiment, the number of LFNST / NSPT kernels is reduced to "1" for B slices and to "2" for P slices.

[0147] In a variation of the fourth embodiment, a high-level syntax element (typically an SPS level flag) is signaled to indicate whether the number of LFNST / NSPT kernels is modified according to the slice type.

[0148] In another modification of the fourth embodiment, the number of kernels used for LFNST / NSPT is reduced to "1" or "2" based on the time depth of the picture or slice containing the current block, as described in relation to Figure 8.

[0149] In this modification, several picture header and / or slice header level syntax elements are introduced to indicate the LFNST / NSPT kernels that are acceptable in the picture or slice.

[0150] These picture header and / or slice header level syntax elements override the SPS level LFNST / NSPT syntax elements, allowing the LFNST / NSPT complexity level to be configured lower than that configured at the SPS level. This higher-level control allows for the adjustment of LFNST / NSPT usage, typically as a function of picture / slice time depth.

[0151] In another variation of the fourth embodiment, the number of kernels used for LFNST / NSPT is reduced based on the characteristics of blocks in the causal neighbors of the current block.

[0152] For example, in the third embodiment, if most of the blocks in the causal neighborhood of the current block are intra-encoded, the number of LFNST and NSPT kernels is reduced to "2" or "1" relative to the current block, thereby limiting the complexity of LFNST / NSTP that would result from excessive use of intra-encoding. For example, if more than half of the blocks in the causal neighborhood of the current block are intra-encoded, the number of LFNST and NSPT kernels is reduced to "2" or "1" relative to the current block.

[0153] In another example, the number of LFNST and NSPT kernels can be reduced to "2" or "1" if most of the blocks in the causal neighborhood of the current block are intercoded, thereby limiting the complexity of LFNST / NSTP when encoding the current block is difficult and a significant encoding gain is expected from LFNST.

[0154] In this modification of the fourth embodiment, the high-level syntax elements (for example, in SPS) indicate that the number of LFNST and NSPT kernels is reduced to "2" or "1" based on the properties of at least one block in the causal neighborhood of the current block.

[0155] Thus, in the fourth embodiment, the number of selectable non-separable transformations depends on the slice type of the slice containing the current block, the time depth of the picture or slice containing the current block, or the characteristics of blocks in the causal neighborhood of the current block.

[0156] Another solution to reduce the complexity of inseparable transformations is to reduce the dimensionality of the LFNST and NSPT kernels. The basic idea here is to use a smaller-dimensional kernel (by zeroing the last N basis vectors of the kernel) on a slice that has higher predictive accuracy than other slices (e.g., an I-slice) (e.g., a B-slice).

[0157] In the fifth embodiment, the dimensions of the selectable LFNST and NSPT kernels for the current block are reduced for the P slice and B slice.

[0158] In a modification of the fifth embodiment, the dimensions of the LFNST and NSPT kernels that can be selected for the current block are reduced only for the B slice.

[0159] In a modification of the fifth embodiment, the dimensions of the selectable LFNST and NSPT kernels for the current block are reduced, according to the time depth of the picture or slice containing the current block. Table TAB2 below provides examples of how the kernel dimensions are reduced at different time depths.

[0160] [Table 2]

[0161] In this modification of the fifth embodiment, several picture header (or slice header) level syntax elements are introduced to demonstrate the use of an LFNST / NSPT kernel with reduced dimensions in the current picture (or slice).

[0162] Therefore, these syntax elements override the SPS-level syntax elements that define the dimensions of the LFNST / NSPT kernel. This override allows control over the dimensions of the LFNST / NSPT kernel at the picture / slice header level, and thus also allows control over the complexity of the LFNST / NSPT.

[0163] In another variation of the fifth embodiment, the dimensions of the LFNST and NSPT kernels selectable for the current block are reduced only for slices with a time depth value of "2" or greater, as shown in Figure 8.

[0164] Thus, in the fifth embodiment, the selectable non-separable transformation depends on the slice type of the slice containing the current block, and the time depth of the picture or slice containing the current block.

[0165] Reducing the complexity of LFNST / NSPT while maintaining good performance compared to a brute-force search for LFNST / NSPT kernels can be achieved by implicitly obtaining the LFNST / NSPT characteristics. This technique, hereafter referred to as "implicit LFNST / NSPT," involves the decoder inferring the best LFNST / NSPT kernel for the current block through known information. Thus, the encoder can benefit from all LFNST / NSPT options while limiting the number of options considered in rate-distortion optimization.

[0166] In the sixth embodiment, the LFNST / NSPT kernel is determined as the modulo of the intra-prediction mode of the current block. For example, the LFNST (or NSPT) kernel represented by the index lfnst_idx (or nspt_idx) is determined as the intra-prediction mode of the current block's intra-mode modulo the total number of LFNST (or NSPT) kernels, num_kernel.

[0167]

number

[0168] Here, "%" represents modulo operation. The number of LFNST / NSPT kernels, num_kernel, is equal to "3" for the last version of LFNST / NSPT, and since "0" means that LFNST / NSPT is not used, +1 is added, so from "1", num_kernel is the index of the LFNST / NSPT kernel from beginning to end.

[0169] The index lfnst_idx (or nspt_idx) is encoded using one bit to signal whether LFNST (or NSPT) is enabled for the current block. On the decoder side, the LFNST (or NSPT) kernel index is reconstructed using intra-prediction mode, similar to the encoder side.

[0170] In the modification of the sixth embodiment, the number of LFNST / NSPT kernels, num_kernel, is reduced from "3" to "2" or "1". The main advantage of this modification is that the encoder checks one or two kernels for each intra-mode, rather than all kernels. This reduces the complexity of the encoder while maintaining high performance, as all kernels are tested for different intra-modes.

[0171] In a variation of the sixth embodiment, a high-level syntax element (typically as an SPS flag within the SPS) is signaled to indicate whether implicit LFNST / NSPT is enabled for a block referencing this SPS.

[0172] In a modification of the sixth embodiment, all LFNST / NSPT options (including no LFNST / NSPT) are distributed across consecutive intramodes. In this case, lfnst_idx (or nspt_idx) is determined as the intra-prediction mode modulo the total number of LFNST kernels, num_kernel+1.

[0173] In a modification of the sixth embodiment, implicit LFNST is enabled only for non-I slices.

[0174] In a modification of the sixth embodiment, implicit LFNST is enabled only for P slices and disabled for B slices.

[0175] In a modification of the sixth embodiment, implicit LFNST is enabled based on the time depth of the picture or slice.

[0176] In a modification of the sixth embodiment, a high-level syntax element (typically an SPS flag within an SPS) is signaled to indicate the implicit use of LFNST for a particular picture or slice type. The implicit use of LFNST can be inferred at the picture level based on the picture's time depth, or at the slice level based on the slice type or slice time depth. Therefore, no additional signaling is required at the picture level or slice level.

[0177] Thus, in the sixth embodiment, the selectable non-separable transformation depends on the intra-prediction mode of the current block.

[0178] Several embodiments have been shown above. Features of these embodiments can be provided individually or in any combination. Furthermore, embodiments may include, individually or in any combination, one or more of the following features, devices, or aspects across various claim categories and types: - A bitstream or signal containing one or more of the described syntax elements, or variations thereof. - To generate, and / or transmit, and / or receive, and / or decode, a bitstream or signal containing one or more of the described syntax elements, or variations thereof. - A TV, set-top box, mobile phone, tablet, or other electronic device that performs at least one of the embodiments described. - A TV, set-top box, mobile phone, tablet, or other electronic device that performs at least one of the embodiments described and displays the resulting picture (for example, using a monitor, screen, or other type of display). - A TV, set-top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal containing an encoded video stream and performs at least one of the embodiments described. - A TV, set-top box, mobile phone, tablet, or other electronic device that wirelessly (for example, using an antenna) receives a signal containing an encoded video stream and performs at least one of the embodiments described. - A server, camera, mobile phone, tablet, or other electronic device that wirelessly (for example, using an antenna) transmits a signal containing an encoded video stream and performs at least one of the embodiments described. - A server, camera, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to transmit a signal containing an encoded video stream and performs at least one of the embodiments described.

Claims

1. To get the current block of the picture, Selecting a transformation to apply to the current block from a set of selectable transformations, which includes at least one non-separable transformation, Applying the selected transformation to the current block, A method including, A method wherein one of the at least one non-separable transforms is selectable for the current block based on the characteristics of an encoding unit containing the current block that is different from the sequence containing the picture, or an encoding unit adjacent to the current block.

2. To get the current block of the picture, In a selectable set of inverse transforms, which includes at least one non-separable inverse transform, the inverse transform applied to the current block is determined. Applying the determined inverse transform to the current block, A method including, A method wherein one of the at least one non-separable inverse transforms is selectable for the current block based on the characteristics of the coding unit containing the current block or a coding unit adjacent to the current block, which is different from the sequence containing the picture.

3. The method of claim 1 or 2, wherein the coding unit is a slice containing the current block, and the characteristic is the slice type of the slice.

4. The method of claim 1 or 2, wherein the encoding unit is the picture containing the current block, and the characteristic is the value of the picture header level syntax element in the picture header of the picture containing the current block.

5. The method of claim 1 or 2, wherein the coding unit is a slice containing the current block, and the characteristic is the value of a slice header level syntax element in the slice header of the slice containing the current block.

6. The method of claim 1 or 2, wherein the coding unit is the picture containing the current block, and the characteristic is a picture parameter set level syntax element in a picture parameter set that references the picture containing the current block.

7. The method of claim 1 or 2, wherein the coding unit is at least one block in the causal neighborhood of the current block, and the characteristic is a characteristic of at least one block in the causal neighborhood of the current block.

8. The method of claim 1 or 2, wherein the coding unit is a slice containing the current block, and the characteristic is the time depth of the slice containing the current block.

9. To get the current block of the picture, Selecting a transformation to apply to the current block from a set of selectable transformations, which includes at least one non-separable transformation, Applying the selected transformation to the current block, A device comprising an electronic circuit configured to perform, A device wherein one of the at least one non-separable transforms is selectable for the current block based on the characteristics of an encoding unit containing the current block that is different from the sequence containing the picture, or an encoding unit adjacent to the current block.

10. To get the current block of the picture, In a selectable set of inverse transforms, which includes at least one non-separable inverse transform, the inverse transform applied to the current block is determined. Applying the determined inverse transform to the current block, A device comprising an electronic circuit configured to perform, An apparatus in which one of the at least one non-separable inverse transforms is selectable for the current block based on the characteristics of the coding unit containing the current block or a coding unit adjacent to the current block, which is different from the sequence containing the picture.

11. The apparatus of claim 9 or 10, wherein the coding unit is a slice containing the current block, and the characteristic is the slice type of the slice.

12. The apparatus of claim 9 or 10, wherein the encoding unit is the picture containing the current block, and the characteristic is the value of a picture header level syntax element in the picture header of the picture containing the current block.

13. The apparatus of claim 9 or 10, wherein the coding unit is a slice containing the current block, and the characteristic is the value of a slice header level syntax element in the slice header of the slice containing the current block.

14. The apparatus of claim 9 or 10, wherein the coding unit is the picture containing the current block, and the characteristic is a picture parameter set level syntax element in a picture parameter set that references the picture containing the current block.

15. The apparatus of claim 9 or 10, wherein the coding unit is at least one block in the causal neighborhood of the current block, and the characteristic is a characteristic of at least one block in the causal neighborhood of the current block.

16. The apparatus of claim 9 or 10, wherein the coding unit is a slice containing the current block, and the characteristic is the time depth of the slice containing the current block.

17. A computer program comprising program code instructions for performing the method according to any one of claims 1 to 8.

18. A non-temporary information storage medium for storing program code instructions for performing the method according to any one of claims 1 to 8.