Low Frequency Non-Separable Transform (LFNST) Signaling

By context encoding of the first bin of the LFNST index and bypass encoding of the second bin, the problem of large signaling overhead related to transformation is solved, and the video data volume is reduced and the encoding efficiency is improved.

CN114223202BActive Publication Date: 2025-07-04QUALCOMM INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202080056777.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-14
Filing Date
2020-08-17
Publication Date
2025-07-04
Estimated Expiration
2040-08-17

AI Technical Summary

Technical Problem

In the existing video encoding technology, the signaling overhead related to transformation is relatively large, resulting in an increase in the amount of video data and affecting the encoding efficiency.

Method used

The first bin of the low-frequency inseparable transform (LFNST) index is used to context code and the second bin is bypass coded to reduce signaling overhead.

Benefits of technology

By reducing signaling overhead, the amount of video data is effectively reduced, while maintaining prediction accuracy and encoding complexity, improving encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114223202B_ABST
    Figure CN114223202B_ABST
Patent Text Reader

Abstract

A method for decoding video data includes: receiving encoded data for a current block; and decoding N bins for a low-frequency non-separable transform (LFNST) index according to the encoded data. The N bins include a first bin and a second bin. Decoding the N bins includes context-decoding each of the N bins. The method further includes: using the N bins to determine the LFNST index; and decoding the encoded data to generate transform coefficients. The method further includes: applying an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and using the residual block and a prediction block for the current block to reconstruct the current block of the video data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of priority of U.S. Application No. 16 / 993,785, filed on Aug. 14, 2020, which claims the benefit of U.S. Provisional Application No. 62 / 889,402, filed on Aug. 20, 2019, and U.S. Provisional Application No. 62 / 903,522, filed on Sep. 20, 2019, the entire contents of each application being incorporated herein by reference. Technical Field

[0002] The present disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital live systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless telephones (so-called "smart phones"), video teleconferencing devices, video streaming devices, etc. Digital video devices implement video encoding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Coding (AVC)), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards. By implementing such video encoding techniques, video devices can send, receive, encode, decode, and / or store digital video information more efficiently.

[0004] Video encoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. For block-based video encoding, a video slice (e.g., a video picture or a portion of a video picture) can be divided into video blocks, which can also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture can be encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture or temporal prediction relative to reference samples in other reference pictures. A picture can be referred to as a frame, and a reference picture can be referred to as a reference frame. Summary of the Invention

[0005] In general, the present disclosure describes techniques for transform signaling for reducing signaling of information related to transforms. Instead of context encoding a first bin for a low-frequency non-separable transform (LFNST) index and bypass encoding a second bin for the LFNST index, the techniques described herein can context encode the first bin and the second bin. By context encoding additional bins for the LFNST index, a video coder (e.g., a video encoder or a video decoder) can determine the LFNST index for a block of a picture of video data with less signaling overhead, thereby potentially reducing the amount of video data transmitted, with little or no loss in prediction accuracy and / or complexity.

[0006] In one example, a method of decoding video data includes: receiving encoded data for a current block; and decoding N bins for an LFNST index based on the encoded data. The N bins include a first bin and a second bin, and decoding the N bins includes context decoding each of the N bins. The method further includes: using the N bins to determine the LFNST index; decoding the encoded data to generate transform coefficients; and applying an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block. The method further includes: reconstructing the current block of the video data using the residual block and a prediction block for the current block.

[0007] In another example, a method of encoding video data includes: generating a residual block for a current block based on a prediction block for the current block; applying a separable transform to the residual block to produce separable transform coefficients; and determining an LFNST index. The method further includes: applying an LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients; and encoding the transform coefficients and N bins for the LFNST index to produce encoded video data. The N bins include a first bin and a second bin, and encoding the N bins includes context encoding each of the N bins. The method further includes: outputting the encoded video data.

[0008] In one example, a device for decoding video data includes one or more processors implemented in circuitry and configured to: receive encoded data for a current block; and decode N bins for an LFNST index based on the encoded data. The N bins include a first bin and a second bin, and to decode the N bins, the one or more processors are configured to perform context decoding on each of the N bins. The one or more processors are further configured to: use the N bins to determine the LFNST index; and decode the encoded data to generate transform coefficients. The one or more processors are further configured to: apply an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and reconstruct the current block of the video data using the residual block and a prediction block for the current block.

[0009] In another example, a device for encoding video data includes one or more processors implemented in circuitry and configured to: generate a residual block for a current block based on a prediction block for the current block; apply a separable transform to the residual block to produce separable transform coefficients; and determine an LFNST index. The one or more processors are further configured to: apply an LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients; and encode the transform coefficients and N bins for the LFNST index to produce encoded video data. The N bins include a first bin and a second bin, and encoding the N bins includes performing context encoding on each of the N bins. The one or more processors are further configured to: output the encoded video data.

[0010] In one example, a device for decoding video data includes: a unit for receiving encoded data for a current block; and a unit for decoding N bins for an LFNST index based on the encoded data. The N bins include a first bin and a second bin, and the unit for decoding the N bins includes a unit for performing context decoding on each of the N bins. The device further includes: a unit for using the N bins to determine the LFNST index; and a unit for decoding the encoded data to generate transform coefficients. The device further includes: a unit for applying an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and a unit for reconstructing the current block of the video data using the residual block and a prediction block for the current block.

[0011] In another example, an apparatus for encoding video data includes: a unit for generating a residual block for a current block based on a predicted block for the current block; and a unit for applying a separable transform to the residual block to produce separable transform coefficients. The apparatus further includes: a unit for determining an LFNST index; and a unit for applying an LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients. The apparatus further includes: a unit for encoding the transform coefficients and N bins for the LFNST index to produce encoded video data. The N bins include a first bin and a second bin, and the unit for encoding the N bins includes a unit for contextually encoding each of the N bins. The apparatus further includes a unit for outputting the encoded video data.

[0012] In one example, a computer-readable storage medium has instructions stored thereon that, when executed, cause one or more processors to perform the following operations: receive encoded data for a current block; and decode N bins for an LFNST index based on the encoded data. The N bins include a first bin and a second bin, and wherein, to decode the N bins, the instructions further cause the one or more processors to perform the following operations: contextually decode each of the N bins. The instructions further cause the one or more processors to perform the following operations: use the N bins to determine the LFNST index; and decode the encoded data to generate transform coefficients. The instructions further cause the one or more processors to perform the following operations: apply an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and reconstruct the current block of the video data using the residual block and the predicted block for the current block.

[0013] In another example, a computer-readable storage medium has instructions stored thereon that, when executed, cause one or more processors to perform the following operations: generate a residual block for a current block based on a predicted block for the current block; and apply a separable transform to the residual block to produce separable transform coefficients. The instructions also cause the one or more processors to perform the following operations: determine an LFNST index; and use the LFNST index to apply LFNST to the separable transform coefficients to produce transform coefficients. The instructions also cause the one or more processors to perform the following operations: encode the transform coefficients and N bins for the LFNST index to produce encoded video data. The N bins include a first bin and a second bin, and in order to encode the N bins, the instructions also cause the one or more processors to perform the following operations: perform context encoding on each of the N bins. The instructions also cause the one or more processors to perform the following operations: output the encoded video data.

[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may execute the techniques of the present disclosure.

[0016] Figure 2A and Figure 2B is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and corresponding coding tree units (CTUs).

[0017] Figure 3 is a block diagram illustrating an example video encoder that may execute the techniques of the present disclosure.

[0018] Figure 4 is a block diagram illustrating an example video decoder that may execute the techniques of the present disclosure.

[0019] Figure 5 is a conceptual diagram illustrating a low-frequency non-separable transform (LFNST) transform on the encoder and decoder sides, where LFNST is a stage between the separable transform and quantization.

[0020] Figure 6 is a conceptual diagram illustrating LFNST coefficients obtained by applying LFNST and zeroing both the Z highest-frequency coefficients in the LFNST region and multiple transform selection (MTS) coefficients outside the LFNST region.

[0021] Figure 7It is a conceptual diagram showing LFNST coefficients obtained by applying LFNST and zeroing out only the MTS coefficients outside the LFNST region.

[0022] Figure 8 It is a conceptual diagram showing 65 angular intra modes.

[0023] Figure 9 It is a conceptual diagram showing angular modes with wide angles in addition to the 65 modes.

[0024] Figure 10 It is a flowchart showing an example method for encoding a current block.

[0025] Figure 11 It is a flowchart showing an example method for decoding a current block of video data.

[0026] Figure 12 It is a flowchart showing an example method for encoding a current block using N bins for LFNST indexing.

[0027] Figure 13 It is a flowchart showing an example method for decoding a current block of video data using N bins for LFNST indexing. Detailed Description

[0028] Generally speaking, the present disclosure describes techniques related to the low-frequency non-separable transform (LFNST). A video encoder can represent a residual block for video data in a form suitable for being signaled by the video encoder and received by a video decoder. It is desirable to reduce the amount of data used to represent the residual block so as to reduce the amount of data sent from the video encoder and received by the video decoder. In video coding, separable transforms have been applied preferentially over non-separable transforms because, compared with non-separable transforms, separable transforms can use fewer operations (e.g., addition, multiplication). A separable transform is a filter that can be written as the product of two or more filters. In contrast, a non-separable filter cannot be written as the product of two or more filters.

[0029] The video encoder can also apply LFNST to increase the energy compaction of the transform coefficient block, rather than relying solely on the separable transform that converts the residual block in the pixel domain to the coefficient block in the frequency domain. For example, LFNST can concentrate the non-zero coefficients of the transform coefficient block closer to the DC coefficient of the transform coefficient block. As a result, there can be fewer transform coefficients between the DC coefficient of the coefficient block and the last valid (i.e., non-zero) transform coefficient of the transform coefficient block, leading to a reduction in the amount of data used to represent the residual block. Similarly, the video decoder can apply the inverse separable transform to transform the transform coefficient block into a residual block. In this way, the data used to represent the residual block can be reduced, thereby reducing the bandwidth and / or storage requirements for the video data and potentially reducing the energy usage of the video decoder and the video encoder.

[0030] The video encoder can generate an LFNST index for LFNST. For example, the video encoder can perform the following operations: when LFNST is disabled, set the LFNST index to "0"; when LFNST is enabled (i.e., not disabled) and the first LFNST type is to be applied, set it to "1"; and when LFNST is enabled (i.e., not disabled) and the second LFNST type is to be applied, set it to "2". In this example, the video encoder can use a first bin (e.g., "0" or "1") and a second bin to represent the LFNST index. In some coding systems, the video encoder can perform context coding on the first bin and bypass coding on the second bin. For example, the video decoder can perform context decoding on the first bin to determine whether LFNST is enabled or disabled for the current block. In this example, the video decoder can perform bypass decoding on the second bin to determine whether LFNST is configured for the first LFNST. Context coding can refer to a situation where a previously decoded value (e.g., for a bin, but other previously decoded values can be possible) can be used to determine the current value of the bin. Examples of context coding can include, for example, an arithmetic coding engine based on CABAC. For example, the video encoder can encode the first bin for the LFNST index based on a context model that is updated based on the previously context decoded value for the first bin for the LFNST index. In contrast, bypass coding can refer to a situation where the previously decoded value for the bin is not used to determine the current value of the bin. For example, in CABAC-based arithmetic coding, the probability that the value of the bin is 0 or 1 is not 50%, and the context can bias this probability to be greater or less than 50%. However, for bypass coding, the probability that the value of the bin is 0 or 1 is 50%, and there is no context to bias this probability.

[0031] According to the techniques of the present disclosure, a video encoder may contextually encode N bins for LFNST indexing. The N bins may include at least a first bin and a second bin, each of which is contextually encoded (e.g., using a CABAC-based arithmetic coding engine). For example, the video encoder may contextually encode the first bin and the second bin for LFNST. Similarly, a video decoder may contextually decode the N bins for LFNST indexing, where the N bins include at least the first bin and the second bin, each of which is contextually encoded. Contextually encoding each of the N bins for LFNST indexing may allow the video decoder to determine the LFNST index of a block of a picture of video data with less signaling overhead, thereby potentially reducing the amount of video data sent, with little or no loss in prediction accuracy and / or complexity.

[0032] Figure 1 FIG. 4 is a block diagram illustrating an example video encoding and decoding system 100 that may perform the techniques of the present disclosure. Generally speaking, the techniques of the present disclosure relate to coding (encoding and / or decoding) video data. Typically, video data includes any data for processing video. Thus, video data may include raw unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0033] As Figure 1 shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 may include any of a variety of devices, including desktop computers, laptop computers (i.e., notebook computers), tablet computers, set-top boxes, cellular phones (such as smart phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and may thus be referred to as wireless communication devices.

[0034] In Figure 1In the example, source device 102 includes video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120, and display device 118. According to the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 can be configured to apply techniques for LFNST signaling (such as, in some examples, based on intra mode). Thus, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, source and destination devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device instead of including an integrated display device.

[0035] In Figure 1 The system 100 shown in is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for LFNST signaling. Source device 102 and destination device 116 are merely examples of such encoding devices, where source device 102 generates encoded video data for transmission to destination device 116. The present disclosure refers to an "encoding" device as a device that performs encoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of encoding devices (specifically, a video encoder and a video decoder), respectively. In some examples, devices 102, 116 may operate in a substantially symmetric manner such that each of devices 102, 116 includes video encoding and decoding components. Thus, system 100 can support one-way or two-way video transmission between video devices 102, 116, e.g., for video streaming, video playback, video broadcast, or video telephony.

[0036] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data for the pictures. The video source 104 of source device 102 may include a video capture device, such as a camera, a video archive unit containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As an additional alternative, video source 104 may generate computer graphics-based data as the source video, or generate a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 may encode the captured, pre-captured, or computer-generated video data. Video encoder 200 may reorder the images from the received order (sometimes referred to as the “display order”) into an encoding order for encoding. Video encoder 200 may generate a bitstream including the encoded video data. Then, source device 102 may output the encoded video data to computer-readable medium 110 via output interface 108 for reception and / or retrieval by, for example, input interface 122 of destination device 116.

[0037] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memories. In some examples, memories 106, 120 may store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 may store software instructions that may be executed, for example, by video encoder 200 and video decoder 300, respectively. Although memories 106 and memory 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, memories 106, 120 may store, for example, the encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106, 120 may be allocated as one or more video buffers, e.g., to store raw, decoded, and / or encoded video data.

[0038] The computer-readable medium 110 can represent any type of medium or device capable of moving the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to directly send the encoded video data to the destination device 116 in real time, for example, via a radio-frequency network or a computer-based network. The output interface 108 can modulate the transmission signal including the encoded video data according to a communication standard such as a wireless communication protocol, and the input interface 122 can demodulate the received transmission information according to a communication standard such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, such as the radio-frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other devices that can be useful for facilitating communication from the source device 102 to the destination device 116.

[0039] In some examples, the computer-readable medium 110 can include a storage device 112. The source device 102 can output the encoded data from the output interface 108 to the storage device 112. Similarly, the destination device 116 can access the encoded data from the storage device 112 via the input interface 122. The storage device 112 can include any of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.

[0040] In some examples, the computer-readable medium 110 can include a file server 114 or another intermediate storage device that can store the encoded video data generated by the source device 102. The source device 102 can output the encoded video data to the file server 114 or another intermediate storage device that can store the encoded video generated by the source device 102. The destination device 116 can access the stored video data from the file server 114 via streaming or downloading. The file server 114 can be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Content Delivery Network (CDN) device, or a Network Attached Storage (NAS) device. The destination device 116 can access the encoded video data from the file server 114 via any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 can be configured to operate according to a streaming protocol, a download transfer protocol, or a combination thereof.

[0041] The output interface 108 and the input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 can be configured to transmit data (such as encoded video data) according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), enhanced LTE, 5G, etc.). In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 can be configured to transmit data (such as encoded video data) according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee TM )、Bluetooth TM standard, etc.). In some examples, the source device 102 and / or the destination device 116 can include respective System-on-Chip (SoC) devices. For example, the source device 102 can include an SoC device for performing the functions ascribed to the video encoder 200 and / or the output interface 108, and the destination device 116 can include an SoC device for performing the functions ascribed to the video decoder 300 and / or the input interface 122.

[0042] The techniques of the present disclosure can be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0043] An input interface 122 of a destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream can include signaling information defined by a video encoder 200 (which is also used by a video decoder 300), such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coding units (e.g., slices, pictures, groups of pictures, sequences, etc.). A display device 118 displays the decoded pictures of the decoded video data to a user. The display device 118 can represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0044] Although not shown in Figure 1 In some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder and can include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream that includes both audio and video in a common data stream. If applicable, the MUX-DEMUX unit can follow the ITU H.223 multiplexer protocol or other protocols (such as the User Datagram Protocol (UDP)).

[0045] Video encoder 200 and video decoder 300 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).

[0046] Video encoder 200 and video decoder 300 may operate according to a video coding standard, such as ITU-T H.265 (also known as the High Efficiency Video Coding (HEVC) standard) or an extension thereof (such as the multi-view or scalable video coding extensions). Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards, such as the ITU-T H.266 standard, also known as Versatile Video Coding (VVC). Drafts of the VVC standard are described in: Bross et al., “Versatile Video Coding (Draft 6)”, Joint Video Team (JVT) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 15th meeting, Gothenburg, Sweden, July 3-12, 2019, JVET-O2001-vB (hereinafter referred to as “VVC Draft 6”); and Bross et al., “Versatile Video Coding (Draft 9)”, Joint Video Team (JVT) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 18th meeting by telephone, April 15-24, 2020, JVET-R2001-v8 (hereinafter referred to as “VVC Draft 9”). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0047] Generally, video encoder 200 and video decoder 300 may perform block-based encoding of pictures. The term "block" generally refers to a structure including data to be processed (e.g., to be encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 may encode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, instead of encoding the red, green, and blue (RGB) data of the samples for a picture, video encoder 200 and video decoder 300 may encode the luminance and chrominance components, where the chrominance components may include both red and blue hue chrominance components. In some examples, video encoder 200 converts the received RGB-formatted data to YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.

[0048] Generally speaking, the present disclosure may relate to encoding (e.g., encoding and decoding) of pictures to include a process of encoding or decoding data of a picture. Similarly, the present disclosure may relate to encoding of blocks of a picture to include a process of encoding or decoding data for a block (e.g., prediction and / or residual encoding). An encoded video bitstream generally includes a series of values for representing encoding decisions (e.g., encoding modes) and syntax elements for partitioning a picture into blocks. Thus, a reference to encoding of a picture or a block should generally be understood as an encoded value of a syntax element for forming the picture or the block.

[0049] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coding device (such as video encoder 200) partitions a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coding device partitions the CTU and CUs into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and a CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coding device may further partition PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of TUs. In HEVC, a PU represents inter-prediction data, while a TU represents residual data. An intra-predicted CU includes intra-prediction information, such as an intra-mode indication.

[0050] As another example, the video encoder 200 and the video decoder 300 may be configured to operate according to VVC. According to VVC, a video encoding device (such as the video encoder 200) divides a picture into a plurality of coding tree units (CTUs). The video encoder 200 may divide a CTU according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple partitioning types, such as the separation between the CU, PU, and TU in HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning, and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0051] In the MTT partitioning structure, quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also referred to as a tri-tree (TT)) may be used to partition a block. Ternary tree or tri-tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, the ternary tree or tri-tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0052] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for the respective chrominance components).

[0053] The video encoder 200 and the video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures per HEVC. For purposes of explanation, a description of the techniques of the present disclosure is given with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video encoding devices configured to use quadtree partitioning or other types of partitioning as well.

[0054] Blocks (e.g., CTUs or CUs) can be grouped in a picture in various ways. As an example, a brick can refer to a rectangular region of CTU rows within a particular tile in the picture. A tile can be a rectangular region of CTUs within a particular tile column and a particular tile row in the picture. A tile column refers to a rectangular region of CTUs that has a height equal to the height of the picture and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular region of CTUs that has a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to the width of the picture.

[0055] In some examples, a tile can be split into multiple bricks, and each brick can include one or more CTU rows within the tile. A tile that is not split into multiple bricks can also be referred to as a brick. However, a brick that is a proper subset of a tile may not be referred to as a tile.

[0056] Bricks in a picture can also be arranged in slices. A slice can be an integer number of bricks of the picture that can be uniquely contained within a single network abstraction layer (NAL) unit. In some examples, a slice includes multiple complete tiles or a contiguous sequence of complete bricks of only one tile.

[0057] This disclosure may interchangeably use "NxN" and "N by N" to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Additionally, a CU does not necessarily need to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include NxM samples, where M does not necessarily equal N.

[0058] Video encoder 200 encodes video data for CU representation prediction and / or residual information and other information. The prediction information indicates how the CU is to be predicted to form a prediction block for the CU. The residual information generally represents the sample-by-sample difference between the samples of the CU before encoding and the prediction block.

[0059] To predict a CU, the video encoder 200 can generally form a prediction block for the CU through inter - prediction or intra - prediction. Inter - prediction generally refers to predicting a CU based on data from previously encoded pictures, while intra - prediction generally refers to predicting a CU based on previously encoded data from the same picture. To perform inter - prediction, the video encoder 200 can use one or more motion vectors to generate the prediction block. The video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, for example, in terms of the difference between the CU and the reference block. The video encoder 200 can calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use uni - directional prediction or bi - directional prediction to predict the current CU.

[0060] VVC can also provide an affine motion compensation mode, which can be considered an inter - prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non - translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0061] To perform intra - prediction, the video encoder 200 can select an intra - prediction mode to generate the prediction block. VVC can provide sixty - seven intra - prediction modes, including various directional modes, as well as a planar mode and a DC mode. Generally, the video encoder 200 selects an intra - prediction mode that describes the neighboring samples of the current block (e.g., the block of the CU) from which the samples of the current block are to be predicted. Assuming the video encoder 200 encodes CTUs and CUs in a raster scan order (from left to right, top to bottom), such samples can generally be above, top - left, or to the left of the current block in the same picture.

[0062] The video encoder 200 encodes data representing the prediction mode for the current block. For example, for an inter - prediction mode, the video encoder 200 can encode data representing which of the various available inter - prediction modes is used and the motion information for the corresponding mode. For uni - directional or bi - directional inter - prediction, for example, the video encoder 200 can use advanced motion vector prediction (AMVP) or the merge mode to encode the motion vectors. The video encoder 200 can use a similar mode to encode the motion vectors for the affine motion compensation mode.

[0063] After prediction such as intra prediction or inter prediction of a block, video encoder 200 may compute residual data for the block. Residual data, such as a residual block, represents the sample-by-sample difference between the block and a predicted block for the block, which is formed using a corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than in the sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. Video encoder 200 produces transform coefficients after applying one or more transforms.

[0064] As described above, after any transform to produce transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process in which the transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thus providing further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a bitwise right shift on the value to be quantized.

[0065] After quantization, video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and thus lower frequency) transform coefficients at the front of the vector and lower energy (and thus higher frequency) transform coefficients at the back of the vector. In some examples, video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy code the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy code the syntax elements representing the transform coefficients in the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy code the values of the syntax elements used to describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.

[0066] To perform CABAC, the video encoder 200 may assign contexts within a context model to symbols to be sent. The context may relate to, for example, whether adjacent values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0067] The video encoder 200 may also generate syntax data (such as block-based syntax data, picture-based syntax data, and sequence-based syntax data) or other syntax data (such as sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS)) for the video decoder 300, for example, in a picture header, a block header, or a slice header. Similarly, the video decoder 300 may decode such syntax data to determine how to decode the corresponding video data.

[0068] In this way, the video encoder 200 may generate a bitstream that includes encoded video data, for example, syntax elements that describe the partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Ultimately, the video decoder 300 may receive the bitstream and decode the encoded video data.

[0069] Generally, the video decoder 300 performs a process opposite to that performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may use CABAC to decode the values of syntax elements for the bitstream in a manner that is substantially similar to, but opposite to, the CABAC encoding process of the video encoder 200. The syntax elements may define partitioning information for partitioning a picture into CTUs and the partitioning of each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. The syntax elements may also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0070] The residual information may be represented by, for example, quantized transform coefficients. The video decoder 300 may inverse-quantize and inverse-transform the quantized transform coefficients of a block to reproduce the residual block for the block. The video decoder 300 uses a signaled prediction mode (intra prediction or inter prediction) and associated prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. The video decoder 300 may then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The video decoder 300 may perform additional processing, such as performing deblocking processing to reduce visual artifacts along the boundaries of the blocks.

[0071] To reduce the amount of data used to represent a residual block, a video encoding device (e.g., video encoder 200 or video decoder 300) may apply LFNST using an LFNST index. The LFNST index may be represented by N bins including at least a first bin and a second bin. For example, the video decoder may context decode the first bin for the LFNST index to determine whether LFNST is enabled or disabled for the current block, and bypass decode the second bin for the LFNST index to determine the type of LFNST to be applied. In this way, the video encoding device may potentially reduce the amount of data used to transmit video data with little or no loss in the prediction accuracy of the video data.

[0072] According to the techniques of the present disclosure, video encoder 200 may be configured to generate a residual block for a current block based on a predicted block for the current block and determine an LFNST index. Video encoder 200 may be configured to apply a separable transform to the residual block to produce separable transform coefficients and use the LFNST index to apply LFNST to the separable transform coefficients to produce transform coefficients. Video encoder 200 may be configured to encode the transform coefficients and the N bins for the LFNST index to produce encoded video data. The N bins may include a first bin and a second bin. To encode the N bins, video encoder 200 may be configured to context encode each of the N bins. Video encoder 200 may be configured to output the encoded data.

[0073] Video decoder 300 may be configured to receive the encoded data for the current block and decode the N bins for the LFNST index according to the encoded data. The N bins may include a first bin and a second bin. To decode the N bins, video decoder 300 may be configured to context decode each of the N bins. Video decoder 300 may be configured to determine the LFNST index using the N bins and decode the encoded data to generate transform coefficients. Video decoder 300 may be configured to use the LFNST index to apply inverse LFNST to the transform coefficients to produce a residual block for the current block and use the residual block and the predicted block for the current block to reconstruct the current block of the video data.

[0074] The present disclosure may generally relate to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the conveyance of values for syntax elements and / or other data used to decode the encoded video data. That is, the video encoder 200 may signal the values for syntax elements in the bitstream. Generally, signaling refers to generating values in the bitstream. As described above, the source device 102 may transmit the bitstream to the destination device 116 substantially in real time or not in real time (such as may occur when storing the syntax elements to the storage device 112 for later retrieval by the destination device 116).

[0075] Figure 2A and Figure 2B FIG. 6 is a conceptual diagram showing an example quadtree binary tree (QTBT) structure 130 and corresponding coding tree units (CTUs) 132. Solid lines represent quadtree splits, and dashed lines represent binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type is used (i.e., horizontal or vertical). In this example, 0 indicates a horizontal split, and 1 indicates a vertical split. For quadtree splits, since a quadtree node splits a block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the split type. Thus, the video encoder 200 may encode, and the video decoder 300 may decode, syntax elements (such as split information) for the region tree level (i.e., the first level) (i.e., solid lines) of the QTBT structure 290, and syntax elements (such as split information) for the prediction tree level (i.e., the second level) (i.e., dashed lines) of the QTBT structure 290. The video encoder 200 may encode video data (such as prediction and transform data) for the CUs represented by the terminal leaf nodes of the QTBT structure 290, and the video decoder 300 may decode the above video data.

[0076] Generally Figure 2B the CTUs 132 may be associated with parameters that define the sizes of the blocks corresponding to the nodes at the first and second levels of the QTBT structure 130. These parameters may include the CTU size (representing the size of the CTU 132 in samples), the minimum quadtree size (MinQTSize, which represents the minimum allowable quadtree leaf node size), the maximum binary tree size (MaxBTSize, which represents the maximum allowable binary tree root node size), the maximum binary tree depth (MaxBTDepth, which represents the maximum allowable binary tree depth), and the minimum binary tree size (MinBTSize, which represents the minimum allowable binary tree leaf node size).

[0077] The root node corresponding to the CTU in the QTBT structure can have four child nodes at the first level of the QTBT structure, and each child node can be divided according to the quadtree division. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 represents such nodes as including a parent node and child nodes with solid branches. If the nodes at the first level are not larger than the maximum allowable binary tree root node size (MaxBTSize), these nodes can be further divided by the corresponding binary tree. The binary tree splitting of a node can be iterated until the nodes generated from the splitting reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents such nodes as having dashed branches. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any further division. As discussed above, the CU can also be referred to as a "video block" or a "block".

[0078] In an example of the QTBT segmentation structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chrominance samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, quadtree division is applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the quadtree leaf node is 128x128, since this size exceeds MaxBTSize (i.e., 64x64 in this example), it will not be further split by the binary tree. Otherwise, the quadtree leaf node will be further divided by the binary tree. Therefore, the quadtree leaf node is also the root node for the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), it means that further vertical splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize means that further horizontal splitting is not allowed for this binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further division.

[0079] Figure 3 is a block diagram showing an example video encoder 200 that can perform the techniques of the present disclosure. Figure 3It is provided for purposes of explanation and should not be considered as limiting the technologies generally illustrated and described in this disclosure. For purposes of explanation, this disclosure describes the video encoder 200 in the context of video coding standards such as the HEVC video coding standard and the developing H.266 video coding standard. However, the technologies of this disclosure are not limited to these video coding standards and are generally applicable to video coding and decoding.

[0080] In Figure 3 the example of, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or in processing circuitry. Additionally, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0081] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104 ( Figure 1 ). The DPB 218 may act as a reference picture memory that stores reference video data for use in predicting subsequent video data by the video encoder 200. The video data memory 230 and the DPB 218 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 may be on-chip (as shown) with other components of the video encoder 200 or off-chip relative to those components.

[0082] In the present disclosure, reference to the video data memory 230 should not be construed as limited to a memory within the video encoder 200 (unless so specifically described), or limited to a memory external to the video encoder 200 (unless so specifically described). Rather, reference to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.

[0083] is shown Figure 3 The various units of to assist in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides a specific function and is pre-set on the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, programmable circuitry can execute software or firmware that causes the programmable circuitry to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuitry can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuitry are generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of the units can be an integrated circuit.

[0084] The video encoder 200 can include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where software executed by programmable circuitry is used to perform the operations of the video encoder 200, the memory 106( Figure 1 ) can store the object code of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 can store such instructions.

[0085] The video data memory 230 is configured to store the received video data. The video encoder 200 can retrieve pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 can be the original video data to be encoded.

[0086] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, a block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0087] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the rate-distortion values obtained for such combinations. The encoding parameters may include dividing a CTU into CUs, the prediction mode for a CU, the transform type for the residual data of a CU, the quantization parameter for the residual data of a CU, etc. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a better rate-distortion value than other tested combinations.

[0088] The mode selection unit 202 may be configured to determine the LFNST index. For example, the mode selection unit 202 may determine whether to enable or disable LFNST for a block based on the compression ratio of the block of video data. The mode selection unit 202 may determine the type of LFNST to be applied to the block. For example, when the first LFNST type has a larger compression ratio than the second LFNST type, the mode selection unit 202 may determine to apply the first LFNST type. In some examples, when the first LFNST type does not have a larger compression ratio than the second LFNST type, the mode selection unit 202 may determine to apply the second LFNST type.

[0089] The mode selection unit 202 may be configured to determine the LFNST index for the block based on whether LFNST is enabled for the block and / or based on whether the first LFNST type or the second LFNST type has been selected for the block. For example, the mode selection unit 202 may perform the following operations: when LFNST is disabled, set the LFNST index to "0"; when LFNST is enabled (i.e., not disabled) and the first LFNST type is to be applied, set the LFNST index to "1"; and when LFNST is enabled (i.e., not disabled) and the second LFNST type is to be applied, set the LFNST index to "2".

[0090] The entropy coding unit 220 can use N bins to represent the LFNST index. In some examples, the N bins can include a first bin (e.g., "0" or "1") and a second bin. The entropy coding unit 220 can perform context coding on each of the N bins. That is, the entropy coding unit 220 can perform context coding on the first bin and the second bin for the LFNST index. Examples of context coding can include, for example, an arithmetic coding engine based on CABAC. For example, the entropy coding unit 220 can perform context coding on the first bin for the LFNST index based on a context model that is updated based on the previously context-coded value for the first bin for the LFNST index. In contrast, bypass coding can refer to a situation where the previously coded value for a bin is not used to determine the current value of that bin. Similarly, the entropy coding unit 220 can perform context coding on the second bin for the LFNST index based on a context model that is updated based on the previously context-coded value for the second bin for the LFNST index. For example, the entropy coding unit 220 can perform the following operations: when the LFNST index is "0", context-code the first bin to indicate "0"; and when the LFNST index is "1" or "2", context-code the first bin to indicate "1". The entropy coding unit 220 can perform the following operations: when the LFNST index is "1", context-code the second bin to indicate "0"; and when the LFNST index is "2", context-code the second bin to indicate "1".

[0091] The video encoder 200 can divide a picture retrieved from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can divide the CTUs of the picture according to a tree structure (such as the QTBT structure or the quadtree structure of HEVC described above). As described above, the video encoder 200 can form one or more CUs by dividing the CTUs according to a tree structure. Such a CU can also typically be referred to as a "video block" or a "block".

[0092] Typically, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., a current CU, or an overlapping portion of a PU and a TU in HEVC). For performing inter prediction on the current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously encoded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 may calculate a value representing how similar a potential reference block will be to the current block, for example, according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 typically may perform these calculations using the per-sample differences between the current block and the considered reference block. The motion estimation unit 222 may identify, from these calculations, the reference block having the lowest value, which indicates the reference block that most closely matches the current block.

[0093] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of the reference block in the reference picture relative to the position of the current block in the current picture. Then, the motion estimation unit 222 may provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. Then, the motion compensation unit 224 may use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve the data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate the values for the prediction block according to one or more interpolation filters. Additionally, for bi-directional inter prediction, the motion compensation unit 224 may retrieve the data of two reference blocks identified by the respective motion vectors, for example, by per-sample averaging or weighted averaging, and combine the retrieved data.

[0094] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block according to the samples adjacent to the current block. For example, for a directional mode, the intra prediction unit 226 may typically mathematically combine the values of the adjacent samples and fill in the calculated values along the defined direction across the current block to produce a prediction block. As another example, for the DC mode, the intra prediction unit 226 may calculate the average value of the adjacent samples of the current block and generate a prediction block to include the obtained average value for each sample of the prediction block.

[0095] The mode selection unit 202 provides a prediction block to the residual generation unit 204. The residual generation unit 204 receives the original, unencoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the per-sample difference between the current block and the prediction block. The resulting per-sample difference defines the residual block for the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, one or more subtractor circuits that perform binary subtraction may be used to form the residual generation unit 204.

[0096] In examples where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As noted above, the size of a CU may refer to the size of the luminance coding block of the CU, and the size of a PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 may support PU sizes of 2Nx2N or NxN for intra prediction, and PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitions for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0097] In examples where the mode selection unit does not further divide a CU into PUs, each CU may be associated with a luminance coding block and a corresponding chrominance coding block. As described above, the size of a CU may refer to the size of the luminance coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0098] For other video coding techniques (such as intra-block copy mode coding, affine mode coding, and linear model (LM) mode coding), as a few examples, the mode selection unit 202 generates a prediction block for the current block being encoded via the corresponding unit associated with the coding technique. In some examples (such as palette mode coding), the mode selection unit 202 may not generate a prediction block and instead generates a syntax element indicating the manner in which the block is to be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.

[0099] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the prediction block and the current block.

[0100] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 can apply various transforms to the residual block to form a transform coefficient block. For example, the transform processing unit 206 can apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 can perform multiple transforms on the residual block, such as a primary transform and a secondary transform (such as a rotation transform). In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0101] The transform processing unit 206 can apply a separable transform to the residual block output by the residual generation unit 204. In this example, the transform processing unit 206 can use an LFNST index to apply the LFNST to the transform coefficients to produce transform coefficients for the current block. For example, the transform processing unit 206 can apply the LFNST configured for the first LFNST type in response to determining that the LFNST is enabled and the LFNST is configured for the first LFNST type (e.g., when the LFNST index is "1"), and apply the LFNST configured for the second LFNST type in response to determining that the LFNST is enabled and the LFNST is configured for the second LFNST type (e.g., when the LFNST index is "2"). In some examples, the first LFNST type can zero both the Z highest frequency coefficients in the LFNST region and the multiple transform selection coefficients outside the LFNST region (see Figure 6 ), and the second LFNST type can zero only the multiple transform selection coefficients outside the LFNST region (see Figure 7 ).

[0102] The quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 can quantize the transform coefficients of the transform coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization can result in information loss, and thus, the quantized transform coefficients can have a lower precision compared to the original transform coefficients generated by the transform processing unit 206.

[0103] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform to the quantized transform coefficient block respectively to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although potentially with a certain degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples from the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0104] The filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce the blocking artifacts along the edges of the CU. In some examples, the operation of the filter unit 216 may be skipped.

[0105] The video encoder 200 stores the reconstructed block in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 may store the reconstructed block into the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 may store the filtered reconstructed block into the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve the reference picture formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter prediction on the blocks of the subsequently encoded pictures. Additionally, the intra prediction unit 226 may use the reconstructed blocks of the current picture in the DPB 218 to perform intra prediction on other blocks in the current picture.

[0106] Generally, the entropy coding unit 220 may perform entropy coding on the syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 may perform entropy coding on the quantized transform coefficient block from the quantization unit 208. As another example, the entropy coding unit 220 may perform entropy coding on the prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit 202. The entropy coding unit 220 may perform one or more entropy coding operations on the syntax elements as another example of video data to generate the entropy-coded data. For example, the entropy coding unit 220 may perform context adaptive variable length coding (CAVLC) operation, CABAC operation, variable-variable (V2V) length coding operation, syntax-based context adaptive binary arithmetic coding (SBAC) operation, probability interval partitioning entropy (PIPE) coding operation, exponential Golomb coding operation, or another entropy coding operation on the data. In some examples, the entropy coding unit 220 may operate in a bypass mode where the syntax elements are not entropy coded.

[0107] Video encoder 200 may output a bitstream, which includes entropy-coded syntax elements required for reconstructed blocks of a slice or a picture. Specifically, entropy coding unit 220 may output the bitstream.

[0108] The above operations are described with respect to blocks. Such descriptions should be understood as operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, luma coding blocks and chroma coding blocks are the luma and chroma components of a CU. In some examples, luma coding blocks and chroma coding blocks are the luma and chroma components of a PU.

[0109] In some examples, it is not necessary to repeat the operations performed on luma coding blocks for chroma coding blocks. As an example, it is not necessary to repeat the operations for identifying motion vectors (MVs) and reference pictures for luma coding blocks to identify MVs and reference pictures for chroma blocks. Rather, the MVs for luma coding blocks may be scaled to determine the MVs for chroma blocks, and the reference pictures may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.

[0110] Video encoder 200 may represent an example of a device configured to generate a residual block based on predicted blocks for a current block using video data for the current block, and determine an LFNST index. Video encoder 200 may be configured to apply a separable transform to the residual block to produce separable transform coefficients, and apply LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients. Video encoder 200 may be configured to encode the transform coefficients and N bins for the LFNST index to produce encoded video data. The N bins may include a first bin and a second bin. To encode the N bins, video encoder 200 may be configured to perform context coding on each of the N bins. Video encoder 200 may be configured to output the encoded data.

[0111] Figure 4 is a block diagram showing an example video decoder 300 that may perform the techniques of the present disclosure. Figure 4 is provided for explanatory purposes and does not limit the techniques generally illustrated and described in the present disclosure. For explanatory purposes, the present disclosure describes video decoder 300 in accordance with the techniques of VVC and HEVC. However, the techniques of the present disclosure may be performed by video coding devices configured for other video coding standards.

[0112] In Figure 4In the example of, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 134 can be implemented in one or more processors or in processing circuitry. Additionally, video decoder 300 can include additional or alternative processors or processing circuitry to perform these and other functions.

[0113] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 can include additional units that perform prediction according to other prediction modes. As an example, the prediction processing unit 304 can include a palette unit, a block copy unit (which can form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, video decoder 300 can include more, fewer, or different functional components.

[0114] The CPB memory 320 can store video data to be decoded by components of the video decoder 300, such as an encoded video bitstream. For example, the video data stored in the CPB memory 320 can be obtained from a computer-readable medium 110( Figure 1 ). The CPB memory 320 can include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 can store video data other than the syntax elements of the encoded pictures, such as temporary data representing the outputs from the various units of the video decoder 300. The DPB 314 generally stores decoded pictures, and the video decoder 300 can output the decoded pictures and / or use the decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 can be formed of any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 can be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 can be on-chip or off-chip relative to the other components of the video decoder 300.

[0115] Additionally or alternatively, in some examples, the video decoder 300 can obtain data from a memory 120( Figure 1)Retrieve the encoded video data. That is, the memory 120 can use the CPB memory 320 to store data as discussed above. Similarly, when some or all of the functions of the video decoder 300 are implemented with software to be executed by the processing circuitry of the video decoder 300, the memory 120 can store the instructions to be executed by the video decoder 300.

[0116] illustrates Figure 4 the respective units illustrated in Figure 3 to assist in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to

[0117] , fixed-function circuitry refers to circuitry that provides a specific function and is pre-set on the operations that can be performed. Programmable circuitry refers to circuitry that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, programmable circuitry can execute software or firmware that causes the programmable circuitry to operate in the manner defined by the instructions of the software or firmware. Fixed-function circuitry can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuitry are generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more units can be integrated circuits.

[0118] The entropy decoding unit 302 can receive the encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. The prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate the decoded video data based on the syntax elements extracted from the bitstream.

[0119] Generally, the video decoder 300 reconstructs pictures on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0120] The entropy decoding unit 302 can perform entropy decoding on the syntax elements of the quantized transform coefficients that define the quantized transform coefficient block and transform information such as quantization parameter (QP) and / or transform mode indication. The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization applied by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse-quantize the quantized transform coefficients. The inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0121] The entropy decoding unit 302 can decode N bins for the LFNST index. In some examples, the N bins can include a first bin and a second bin. The entropy decoding unit 302 can perform context decoding on each of the N bins. That is, the entropy decoding unit 302 can perform context decoding on the first bin and the second bin for the LFNST index. For example, the entropy decoding unit 302 can perform context decoding on the first bin for the LFNST index based on a context model that is updated based on the previously context-decoded value for the first bin for the LFNST index. Similarly, the entropy decoding unit 302 can perform context decoding on the second bin for the LFNST index based on a context model that is updated based on the previously context-decoded value for the second bin for the LFNST index.

[0122] The prediction processing unit 304 can be configured to determine the LFNST index using the N bins. For example, the prediction processing unit 304 can perform the following operations: when the first bin is "0", determine that the LFNST index is "0"; when the first bin is "1" and the second bin is "0", determine that the LFNST index is "1"; and when the first bin is "1" and the second bin is "1", determine that the LFNST index is "2".

[0123] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, inverse integer transform, inverse Karhunen-Loeve transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.

[0124] The inverse transform processing unit 308 may use the LFNST index to apply an inverse LFNST to the transform coefficients to generate a residual block for the current block. For example, the inverse transform processing unit 308 may perform the following operations: when the LFNST index is "1", apply an inverse LFNST configured for the first LFNST type; and when the LFNST index is "2", apply an inverse LFNST configured for the second LFNST type. In some examples, the inverse transform processing unit 308 may apply an inverse separable transform to the transform coefficients after applying the LFNST to generate a residual block.

[0125] In addition, the prediction processing unit 304 generates a prediction block based on the prediction information syntax element entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-predicted, the motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate a reference picture in the DPB 314 from which a reference block is to be retrieved, and a motion vector that identifies the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally perform the inter-prediction process in a manner substantially similar to the manner described with respect to the motion compensation unit 224( Figure 3 )).

[0126] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 can generally perform the intra-prediction process in a manner substantially similar to the manner described with respect to the intra-prediction unit 226( Figure 3 ). The intra-prediction unit 318 may retrieve data of adjacent samples of the current block from the DPB 314.

[0127] The reconstruction unit 310 may use the prediction block and the residual block to reconstruct the current block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0128] The filter unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all examples.

[0129] Video decoder 300 may store the reconstructed blocks in DPB 314. For example, in an example where the operations of filter unit 312 are not performed, reconstruction unit 310 may store the reconstructed blocks into DPB 314. In an example where the operations of filter unit 312 are performed, filter unit 312 may store the filtered reconstructed blocks into DPB 314. As discussed above, DPB 314 may provide reference information (such as the current picture for intra prediction and samples of previously decoded pictures for subsequent motion compensation) to prediction processing unit 304. In addition, video decoder 300 may output the decoded picture from DPB for subsequent presentation on a display device such as Figure 1 display device 118.

[0130] In this way, video decoder 300 represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units that are implemented in circuitry and configured to receive the encoded data for the current block and decode the N bins for the LFNST index according to the encoded data. The N bins include a first bin and a second bin. To decode the N bins, video decoder 300 may be configured to perform context decoding on each of the N bins. Video decoder 300 may be configured to determine to use the N bins to determine the LFNST index and decode the encoded data to generate transform coefficients. Video decoder 300 may be configured to apply inverse LFNST to the transform coefficients using the LFNST index to generate a residual block for the current block and use the residual block and the prediction block for the current block to reconstruct the current block of the video data.

[0131] The present disclosure is related to transform coding, which is part of most modern video compression standards, such as but not limited to M. Wien, High Effciency Video Coding: Coding Tools and Specification, Springer-Verlag, Berlin, 2015. The present disclosure describes various transform signaling processes that may be used in a video coding apparatus to specify the transform selected from among multiple transform candidates for encoding / decoding. A video coding apparatus (e.g., video encoder 200 and video decoder 300) may be configured to apply the techniques described herein based on available side information such as intra mode type to reduce signaling overhead and thus may improve coding efficiency and / or may be used in advanced video codecs including extensions of HEVC and next-generation video coding standards such as VVC.

[0132] In video coding standards prior to HEVC, when using DCT-2 both vertically and horizontally, only a fixed separable transform is used. In HEVC, in addition to DCT-2, DST-7 is also adopted as a fixed separable transform for 4x4 blocks.

[0133] U.S. Patent No. 10,306,229, U.S. Patent Application Publication No. 2018 / 0020218, and U.S. Patent Application No. 16 / 426,749, each incorporated herein by reference, describe a Multiple Transform Selection (MTS) method. MTS was previously known as Adaptive Multiple Transform (AMT), which is only a name change and its technology can be the same.

[0134] An example of MTS described in U.S. Patent Application No. 16 / 426,749 has been adopted in the Joint Exploration Model (JEM-7.0) (JEM software (https: / / jvet.hhi.fraunhofer.de / svn / svn_HMJEMSoftware / tags / HM-16.6-JEM-7.0)) of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, and a simplified version of MTS was later adopted in VVC (hereinafter referred to as "JEM-7.0").

[0135] In JEM-7.0, the Low-Frequency Non-Separable Transform (LFNST) shown in Figure 5 is used to further improve the coding efficiency of MTS, where the implementation of LFNST is based on the example described in U.S. Patent Application Publication No. US 2017 / 0238013, which is incorporated herein by reference (see also U.S. Patent Application Publication No. US2017 / 0094313, U.S. Patent Application Publication No. US 2017 / 0094314, U.S. Patent No. 10,349,085, and Patent Application No. 16 / 364,007, each incorporated herein by reference for alternative or additional designs and other examples). Recently, LFNST has been adopted in JVET-N0193, Reduced Secondary Transform (RST) (CE6-3.1) (available online: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N0193-v5.zip). LFNST was previously known as Non-Separable Secondary Transform (NSST) or Secondary Transform, where LFNST, NSST, and Secondary Transform can be the same.

[0136] Figure 5 is a conceptual diagram showing the low-frequency non-separable transform (LFNST) at the encoder and decoder sides, where the LFNST is a stage between the separable transform and quantization. In Figure 5 the example of, a video encoder 200 (e.g., the transform processing unit 206 of the video encoder 200) may apply a separable transform 440 to the coefficients of a residual block to generate separable transform coefficients, and then may apply an LFNST 442 to the separable transform coefficients to generate transform coefficients. The video encoder 200 (e.g., the quantization unit 208 of the video encoder 200) may apply quantization 444 to the transform coefficients to generate quantized transform coefficients.

[0137] In Figure 5 the example of, a video decoder 300 (or more specifically, the inverse quantization unit 306 of the video decoder 300) may apply inverse quantization 450 to the quantized transform coefficients to generate transform coefficients. In this example, the video decoder 300 (or more specifically, the inverse transform processing unit 308 of the video decoder 300) may apply an inverse LFNST 452 to the transform coefficients to generate separable transform coefficients, and may apply an inverse separable transform 454 to the separable transform coefficients to generate a residual block. The separable transform 440 may convert the residual samples of the residual block from the pixel domain to separable transform coefficients in the frequency domain. The inverse separable transform 454 may convert the separable transform coefficients in the frequency domain to the residual samples of the residual block in the pixel domain. The LFNST 442 may be used for better energy compaction of the transform coefficients.

[0138] Figure 6 is a conceptual diagram showing the LFNST coefficients obtained by applying LFNST and zeroing both the Z highest-frequency coefficients in the LFNST region 460 and the MTS coefficients outside the LFNST region 460. Figure 7 is a conceptual diagram showing the LFNST coefficients obtained by applying LFNST and zeroing only the MTS coefficients outside the LFNST region 462.

[0139] In the current LFNST design in VVC Draft 6, a video encoder (e.g., video encoder 200) may perform a zeroing operation that preserves the K lowest-frequency coefficients transformed by an LFNST of size N (e.g., for 8x8, N = 64), and a video decoder (e.g., video decoder 300) reconstructs separable coefficients (e.g., MTS coefficients) by using only those K coefficients. In VVC Draft 6, such a zeroing process is specified only for LFNSTs of sizes 4x4 and 8x8, where the video decoder 300 implicitly infers (e.g., determines without signaling) that the remaining N - K higher-frequency coefficients are set to zero, and the K LFNST coefficients are used for reconstruction. Figure 6 and Figure 7 shows the transform coefficients obtained after applying the LFNST, where zeroing is performed at the top of a subset of the separable transform coefficients (e.g., MTS coefficients within a sub-block labeled as the "LFNST region"), while the remaining coefficients (e.g., MTS coefficients outside the sub-block labeled as the "LFNST region") are specified to be set to zero. As discussed in U.S. Patent Application Publication No. US 2017 / 0094313, U.S. Patent Application Publication No. US 2017 / 0094314, and Patent Application No. 62 / 849,689 (each incorporated herein by reference), the video encoder 200 or the video decoder 300 may perform the LFNST by first converting a two-dimensional block (e.g., Figure 6 the LFNST region 460 of Figure 7 or the LFNST region 462 of

[0140] to a one-dimensional list (or vector) of coefficients via a predefined scan / sorting and then applying a transform to the subset of coefficients. In VVC Draft 6, the LFNST index may be 0, 1, or 2, such that the LFNST index is binary-coded using 2 bins. That is, the video encoder 200 may encode the N bins for the LFNST index to produce encoded data, where the N bins include a first bin and a second bin (i.e., two bins). Similarly, the video decoder 300 may decode the N bins for the LFNST index to produce decoded data, where the N bins include a first bin and a second bin (i.e., two bins). The first bin is context-coded, and the second bin is bypass-coded. The present disclosure describes various processes for reducing the signaling overhead of LFNST indices / flags. In some examples, the various processes may be intra-mode-dependent (e.g., depending on the type of intra-mode applied to encode or decode the current block). Although the signaling processes disclosed in the present disclosure are customized for LFNST, the signaling process is not limited to LFNST and may be applied to reduce the signaling of transform-related syntax.

[0141] The present disclosure describes the following transform signaling methods that may be used alone or in any combination.

[0142] Video encoder 200 and video decoder 300 may be configured to encode the LFNST flag / index using N bins. For example, video encoder 200 and video decoder 300 may be configured to contextually encode (e.g., encode or decode) each bin (e.g., using a CABAC-based arithmetic coding engine). As an example, video encoder 200 and video decoder 300 may be configured to contextually encode each bin for the LFNST flag / index. That is, video encoder 200 may be configured to encode N bins for the LFNST index, where the N bins include a first bin and a second bin, and where, to encode the N bins, video encoder 200 is configured to contextually encode each of the N bins. Similarly, video decoder 300 may be configured to decode N bins for the LFNST index, where the N bins include a first bin and a second bin, and where, to decode the N bins, video decoder 300 is configured to contextually decode each of the N bins.

[0143] Video encoder 200 and video decoder 300 may be configured to contextually encode a subset of the bins, while the remaining bins among these may be bypass encoded. For example, video encoder 200 and video decoder 300 may be configured to contextually encode each bin in a subset of the bins for the LFNST flag / index, where the subset is less than all of the bins for the LFNST flag / index. In some examples, video encoder 200 and video decoder 300 may be configured to bypass encode each bin.

[0144] Some example techniques for context - based coding of any bin are described below. In some examples, video encoder 200 and video decoder 300 may be configured to use separate contexts associated with the luminance and chrominance channels. For example, video encoder 200 may encode the first of N bins using a first context in response to determining that the first of N bins is assigned to the luminance channel. Similarly, video encoder 200 may encode the second of N bins using the first context in response to determining that the second of N bins is assigned to the luminance channel. In some examples, video encoder 200 may encode the first of N bins using a second context different from the first context in response to determining that the first of N bins is assigned to the chrominance channel. Similarly, video encoder 200 may encode the second of N bins using a second context different from the first context in response to determining that the second of N bins is assigned to the chrominance channel.

[0145] Video decoder 300 may decode the first of N bins using the first context in response to determining that the first of N bins is assigned to the luminance channel. Similarly, video decoder 300 may encode the second of N bins using the first context in response to determining that the second of N bins is assigned to the luminance channel. In some examples, video decoder 300 may encode the first of N bins using a second context different from the first context in response to determining that the first of N bins is assigned to the chrominance channel. Similarly, video decoder 300 may encode the second of N bins using a second context different from the first context to decode the second of N bins.

[0146] Video encoder 200 and video decoder 300 may be configured to use separate contexts for different types of block - splitting modes (e.g., dual - tree, single - tree, or other TU / CU - level splitting).

[0147] Video encoder 200 and video decoder 300 may be configured to use separate contexts for different intra - frame modes (e.g., see Figure 8 and Figure 9 ). Figure 8 is a conceptual diagram showing 65 angular intra - frame modes 470. Figure 9 is a conceptual diagram showing angular modes with wide angles 472 (- 1 to - 10 and 67 to 76) in addition to the 65 modes.

[0148] In some examples, video encoder 200 and video decoder 300 may be configured to use one context for the planar mode and another context for the remaining modes among these modes. In some examples, video encoder 200 and video decoder 300 may be configured to share a single context for the planar mode and the DC mode. For example, video encoder 200 and video decoder 300 may be configured to use a specific context for both the planar mode and the DC mode.

[0149] Video encoder 200 and video decoder 300 may be configured to use a subset of angular modes that share a single context. Angular modes with similar orientations may share a single context, where the similarity may be a measure of the proximity between angular modes. As an example, a threshold value T may specify the number of adjacent modes grouped to share a single context. For example, T = 3 may indicate that the 4 modes closest to the horizontal mode (i.e., Figure 9 mode 18 in

[0150] can share a context (i.e., modes from 15 to 21 can share the same context).

[0151] To encode bins, video encoder 200 and video decoder 300 may be configured to use a single context for all intra modes, but the bin values may be predicted or modified based on mode information. That is, video encoder 200 may be configured to use the same context for all intra modes and, using this context, encode the bin values based on the intra mode information for the corresponding bin among N bins. Video decoder 300 may be configured to use the same context for all intra modes and predict the bin values based on the intra mode information for the corresponding bin among N bins.

[0152] The video encoder 200 and the video decoder 300 can be configured to flip the bin value 0 or 1 for certain modes (e.g., 0 can be flipped to the value 1). For example, the video encoder 200 and the video decoder 300 can be configured to flip the bin value 0 (relative to 1) to 1 (relative to 0) for odd intra-frame modes, while for even modes, the video encoder 200 and the video decoder 300 can encode (e.g., decode or encode) the value without any flipping.

[0153] For example, the video encoder 200 can be configured to determine the value for the intra-frame mode used to encode the current block and to determine the value for the last bin (e.g., the second bin) among the N bins. In this example, the video encoder 200 can be configured to contextually encode the value of the last bin among the N bins based on the value for the intra-frame mode.

[0154] The video decoder 300 can be configured to determine the value for the intra-frame mode used to decode the current block and to contextually decode the value of the last bin among the N bins to generate a decoded value. In this example, the video decoder 300 can be configured to determine the value of the last bin based on the decoded value and the value for the intra-frame mode.

[0155] For example, the video encoder 200 can be configured to determine whether the value for the intra-frame mode is odd. In response to determining that the value for the intra-frame mode is not odd, the video encoder 200 can contextually encode the value of the last bin. For example, the video encoder 200 can perform the following operations in response to determining that the value for the intra-frame mode is not odd: when the value of the last bin is "0", contextually encode the encoded value as "0"; and when the value of the last bin is "1", contextually encode the encoded value as "1". However, in response to determining that the value for the intra-frame mode is odd, the video encoder 200 can contextually encode the value opposite to the value of the last bin. For example, the video encoder 200 can perform the following operations in response to determining that the value for the intra-frame mode is odd: when the value of the last bin is "0", contextually encode the encoded value as "1"; and when the value of the last bin is "1", contextually encode the encoded value as "0".

[0156] The video decoder 300 may be configured to determine whether the value for the intra mode is odd. In response to determining that the value for the intra mode is not odd, the video decoder 300 may predict the value of the last bin as the decoded value. For example, the video decoder 300 may perform the following operations in response to determining that the value for the intra mode is not odd: when the decoded value for the last bin is "0", predict the value of the last bin as "0"; and when the decoded value for the last bin is "1", predict the value of the last bin as "1". However, in response to determining that the value for the intra mode is odd, the video decoder 300 may predict the value of the last bin as the value opposite to the decoded value. For example, the video decoder 300 may perform the following operations in response to determining that the value for the intra mode is odd: when the decoded value for the last bin is "0", predict the value of the last bin as "1"; and when the decoded value for the last bin is "1", predict the value of the last bin as "0".

[0157] The video encoder 200 may be configured to determine whether the value for the intra mode is even. In response to determining that the value for the intra mode is not even, the video encoder 200 may context-encode the encoded value as the value of the last bin. For example, the video encoder 200 may perform the following operations in response to determining that the value for the intra mode is not even: when the value of the last bin is "0", context-encode the encoded value as "0"; and when the value of the last bin is "1", context-encode the encoded value as "1". However, in response to determining that the value for the intra mode is even, the video encoder 200 may context-encode the encoded value as the value opposite to the value of the last bin. For example, the video encoder 200 may perform the following operations in response to determining that the value for the intra mode is even: when the value of the last bin is "0", context-encode the encoded value as "1"; and when the value of the last bin is "1", context-encode the encoded value as "0".

[0158] The video decoder 300 may be configured to determine whether the value for the intra mode is even. In response to determining that the value for the intra mode is not even, the video decoder 300 may predict the value of the last bin as the decoded value. For example, the video decoder 300 may perform the following operations in response to determining that the value for the intra mode is not even: when the decoded value for the last bin is "0", predict the value of the last bin as "0"; and when the decoded value for the last bin is "1", predict the value of the last bin as "1". However, in response to determining that the value for the intra mode is even, the video decoder 300 may predict the value of the last bin as the value opposite to the decoded value. For example, the video decoder 300 may perform the following operations in response to determining that the value for the intra mode is even: when the decoded value for the last bin is "0", predict the value of the last bin as "1"; and when the decoded value for the last bin is "1", predict the value of the last bin as "0".

[0159] In some examples, the video encoder 200 and the video decoder 300 may be configured to introduce an LFNST index predictor based on intra mode information. For example, the video encoder 200 and the video decoder 300 may be configured to create a lookup table to map intra prediction modes to the most likely LFNST index or LFNST kernel used as a predictor. To indicate the LFNST kernel, instead of directly signaling the LFNST index, the video encoder 200 and the video decoder 300 may be configured to signal other information indicating whether the index is equal to the predictor. In the case of a binary choice (e.g., only two choices for LFNST are possible), the video encoder 200 and the video decoder 300 may be configured to derive the actual LFNST kernel based on the other information indicating whether the index is equal to the predictor. If the index is equal to the predictor, the video encoder 200 and the video decoder 300 may be configured to determine that the kernel is equal to the kernel from the lookup table. Otherwise, the video encoder 200 and the video decoder 300 may determine that the kernel is equal to another (not the predictor) kernel. This example may be extended to more than 2 choices (e.g., more than two choices for LFNST). In such an example, if the LFNST index or kernel is not equal to the predictor, the video encoder 200 may be configured to signal information indicating the LFNST selection from the set of selections where the predictor selection is removed.

[0160] In VVC, the above techniques (either individually or in any combination) may be applied to encode the second (last) bin of the LFNST index (which is currently bypass-encoded in VVC draft 6).

[0161] As a specific example for VVC Draft 6 (e.g., the example described in U.S. Patent Application Publication No. US 2017 / 0238013, which is incorporated herein by reference), a signaling method can be as follows:

[0162] Video encoder 200 and / or video decoder 300 can be configured to context-encode the last bin (last_lfnst_bin) of the LFNST index.

[0163] Video encoder 200 and / or video decoder 300 can be configured to predict the value of the last bin as follows according to the intra mode (e.g., predModeIntra):

[0164] if (predModeIntra % 2 == 0)

[0165] predicted_last_lfnst_bin = last_lfnst_bin

[0166] else

[0167] predicted_last_lfnst_bin =!(last_lfnst_bin)

[0168] where "!" represents the flip / inversion operation, such that in the case where predModeIntra is odd, it flips / inverts the signaled last_lfnst_bin value to obtain the predicted last bin predicted_last_lfnst_bin. Otherwise, if predModeIntra is even, the signaled last bin is used as it is.

[0169] In another example, video encoder 200 and / or video decoder 300 can be configured to predict and / or flip the bin based on the following:

[0170] if (predModeIntra % 2 == 1)

[0171] predicted_last_lfnst_bin = last_lfnst_bin

[0172] else

[0173] predicted_last_lfnst_bin =!(last_lfnst_bin)

[0174] Wherein, if predModeIntra is even, the bin value (e.g., last_1fnst_bin) can be flipped, which is different from the previous example where the bin value was flipped if predModeIntra was odd.

[0175] Figure 10 is a flowchart illustrating an example method for encoding a current block. The current block may include a current CU. Although described with respect to video encoder 200 ( Figure 1 and Figure 3 ), it should be understood that other devices may be configured to perform methods similar to those of Figure 10 .

[0176] In this example, the mode selection unit 202 of video encoder 200 may predict the current block (502). For example, the mode selection unit 202 may form a prediction block for the current block. The residual generation unit 204 of video encoder 200 may calculate a residual block (504) for the current block. To calculate the residual block, video encoder 200 may calculate the difference between the original, unencoded block and the prediction block for the current block. The transform processing unit 206 and quantization unit 208 of video encoder 200 may transform and quantize the coefficients of the residual block (506). For example, the transform processing unit 206 may apply LFNST using N bins for the LFNST index, where each of the N bins is context-encoded. The entropy encoding unit 220 of video encoder 200 may scan the quantized transform coefficients of the residual block (508). During or after the scan, video encoder 200 may entropy-encode the transform coefficients (510). For example, the entropy encoding unit 220 may use CAVLC or CABAC to encode the transform coefficients. In some examples, the entropy encoding unit 220 may context-encode each of the N bins for the LFNST index. Then, the entropy encoding unit 220 may output the entropy-encoded data of the block (512).

[0177] Figure 11 is a flowchart illustrating an example method for decoding a current block of video data. The current block may include a current CU. Although described with respect to video decoder 300 ( Figure 1 and Figure 4 ), it should be understood that other devices may be configured to perform methods similar to those of Figure 11 .

[0178] The entropy decoding unit 302 of the video decoder 300 may receive entropy - encoded data for the current block (such as entropy - encoded prediction information and entropy - encoded data for the coefficients of the residual block corresponding to the current block) (522). The entropy decoding unit 302 may perform entropy decoding on the entropy - encoded data to determine the prediction information for the current block and reproduce the coefficients of the residual block (524). For example, the entropy decoding unit 302 may perform context decoding on each of the N bins for the LFNST index.

[0179] The prediction processing unit 304 of the video decoder 300 may predict the current block (526), for example, using an intra - frame or inter - frame prediction mode as indicated by the prediction information for the current block, to calculate the prediction block for the current block. Then, the entropy decoding unit 302 may perform an inverse scan on the reproduced coefficients (528) to create a block of quantized transform coefficients. The inverse quantization unit 306 and the inverse transform processing unit 308 may perform inverse quantization and inverse transform on the transform coefficients to generate a residual block (530). For example, the inverse transform processing unit 308 may apply an inverse LFNST using the N bins for the LFNST index, where each of the N bins is context - decoded. The reconstruction unit 310 of the video decoder 300 may decode the current block by combining the prediction block and the residual block (532).

[0180] Figure 12 is a flowchart showing an example method for encoding a current block using N bins for the LFNST index. Although described with respect to the video encoder 200 ( Figure 1 and Figure 3 ), other devices may be configured to perform methods similar to the Figure 12 method.

[0181] The video encoder (e.g., the video encoder 200, or more specifically, e.g., the residual generation unit 204 of the video encoder 200) may generate a residual block for the current block based on the prediction block for the current block (550). The video encoder (e.g., the video encoder 200, or more specifically, e.g., the transform processing unit 206 of the video encoder 200) may apply a separable transform to the residual block to produce separable transform coefficients (552). The video encoder may determine the LFNST index (554). For example, the mode selection unit 202 may operate as follows: when LFNST is disabled, determine the LFNST index as "0"; when LFNST is enabled (i.e., not disabled) and the first LFNST type is to be applied, determine the LFNST index as "1"; and when LFNST is enabled (i.e., not disabled) and the second LFNST type is to be applied, determine the LFNST index as "2".

[0182] A video encoder may use an LFNST index to apply an LFNST to separable transform coefficients to generate transform coefficients (556). For example, the transform processing unit 206 may perform the following operations: when the LFNST index is "1", apply the LFNST configured for the first LFNST type to the separable transform coefficients; and when the LFNST index is "2", apply the LFNST configured for the second LFNST type to the separable transform coefficients.

[0183] The video encoder may encode the transform coefficients and N bins for the LFNST index to generate encoded video data (558). The N bins may include a first bin and a second bin. To encode the N bins, the video encoder may perform context encoding on each of the N bins. For example, the entropy encoding unit 220 may perform the following operations: when the LFNST index is "0", perform context encoding on the first bin indicating "0"; and when the LFNST index is "1" or "2", perform context encoding on the first bin indicating "1". The entropy encoding unit 220 may perform the following operations: when the LFNST index is "1", perform context encoding on the second bin indicating "0"; and when the LFNST index is "2", perform context encoding on the second bin indicating "1". The video encoder may output the encoded video data (560).

[0184] In some examples, the video encoder may allocate a context model for performing context encoding on the first bin based on the luminance channel. For example, in response to determining that the first bin among the N bins is assigned to the luminance channel, the video encoder may encode the first bin among the N bins using a first context. In response to determining that the second bin among the N bins is assigned to the chrominance channel, the video encoder may encode the second bin among the N bins using a second context different from the first context.

[0185] Figure 13 is a flowchart illustrating an example method for decoding a current block of video data using N bins for an LFNST index. Although described with respect to video decoder 300 ( Figure 1 and Figure 4 ), other devices may be configured to perform methods similar to the Figure 13 method.

[0186] A video decoder (e.g., video decoder 300, or more specifically, e.g., the entropy decoding unit 302 of video encoder 200) may receive the encoded data (570) for a current block. The video decoder (e.g., video decoder 300, or more specifically, e.g., the entropy decoding unit 302 of video encoder 200) may decode N bins for the LFNST index based on the encoded data (572). The N bins may include a first bin and a second bin. To decode the N bins, the video decoder may perform context decoding on each of the N bins.

[0187] The video decoder (e.g., video decoder 300, or more specifically, e.g., the prediction processing unit 304 of video encoder 200) may use the N bins to determine the LFNST index (574). For example, when the entropy decoding unit 302 context decodes the first bin as "0", the prediction processing unit 304 may determine the LFNST index as "0"; when the entropy decoding unit 302 context decodes the first bin as "1" and the second bin as "0", the prediction processing unit 304 may determine the LFNST index as "1"; and when the entropy decoding unit 302 context decodes the first bin as "1" and the second bin as "1", the prediction processing unit 304 may determine the LFNST index as "2".

[0188] The video decoder (e.g., video decoder 300, or more specifically, e.g., the entropy decoding unit 302 of video encoder 200) may decode the encoded data to generate transform coefficients (576). For example, the entropy decoding unit 302 may generate quantized transform coefficients based on the received encoded data, and the inverse quantization unit 306 may generate transform coefficients based on the quantized transform coefficients.

[0189] The video decoder (e.g., video decoder 300, or more specifically, e.g., the inverse transform processing unit 308 of video encoder 200) may use the LFNST index to apply an inverse LFNST to the transform coefficients to produce a residual block for the current block (578). For example, the inverse transform processing unit 308 may perform the following operations: when the LFNST index is "1", apply an inverse LFNST configured for a first LFNST type to the transform coefficients; and when the LFNST index is "2", apply an inverse LFNST configured for a second LFNST type to the transform coefficients. The inverse transform processing unit 308 may apply an inverse separable transform to the separable transform coefficients output by the inverse LFNST to produce a residual block for the current block. The video decoder (e.g., video decoder 300, or more specifically, e.g., the reconstruction unit 310 of video encoder 200) may use the residual block and the prediction block for the current block to reconstruct the current block of the video data (580).

[0190] In some examples, the video decoder may determine a context model for context encoding the first bin based on whether the current block is for a luminance channel or a chrominance channel. For example, in response to determining that the first bin among the N bins is assigned to the luminance channel, the video decoder may use a first context to decode the first bin among the N bins. In response to determining that the second bin among the N bins is assigned to the chrominance channel, the video decoder may use a second context different from the first context to decode the second bin among the N bins.

[0191] The following provides a non-limiting illustrative list of examples of the techniques of the present disclosure.

[0192] Example 1. A method for decoding video data, the method comprising: receiving entropy-encoded data for a current block; entropy-decoding N bins for a low-frequency non-separable transform (LFNST) flag / index according to the entropy-encoded data, where N is an integer value greater than 0; using the N bins to determine the LFNST flag / index; entropy-decoding the entropy-encoded data to generate coefficients; applying LFNST to the coefficients using the LFNST flag / index to produce a residual block; and reconstructing the current block of the video data using the residual block and a prediction block for the current block.

[0193] Example 2. A method for encoding video data, the method comprising: generating a residual block using video data for a current block and a prediction block for the current block; determining a low-frequency non-separable transform (LFNST) flag / index; applying LFNST to the residual block using the LFNST flag / index to produce coefficients; entropy-encoding the coefficients and N bins for the LFNST flag / index to produce entropy-encoded video data, where N is an integer value greater than 0; and outputting the entropy-encoded video data.

[0194] Example 3. The method according to Example 1 or 2, wherein entropy-encoding or decoding the N bins comprises: context-encoding or decoding each of the N bins.

[0195] Example 4. The method according to Example 1 or 2, wherein entropy-encoding or decoding the N bins comprises: determining a subset of the N bins from the N bins, the subset of the N bins including fewer bins than the N bins; based on the subset of the N bins, determining the remainder of the bins from the N bins, wherein the subset of the N bins and the remainder of the bins form the N bins; context-encoding or decoding the subset of the N bins; and context-encoding or decoding the remainder of the bins from the N bins.

[0196] Example 5. The method according to Example 1 or 2, wherein entropy encoding or decoding the N bins includes: bypass encoding or decoding each of the N bins.

[0197] Example 6. The method according to any one of Examples 1-4, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin of the N bins is assigned to a luminance channel, context encoding or decoding the first bin of the N bins using a first context; and in response to determining that a second bin of the N bins is assigned to a chrominance channel, context encoding or decoding the second bin of the N bins using a second context different from the first context.

[0198] Example 7. The method according to any one of Examples 1-4 and 6, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin of the N bins is assigned to a first block partitioning mode, context encoding or decoding the first bin of the N bins using a first context; and in response to determining that a second bin of the N bins is assigned to a second block partitioning mode, context encoding or decoding the second bin of the N bins using a second context different from the first context.

[0199] Example 8. The method according to any one of Examples 1-4, 6 and 7, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin of the N bins is assigned to a first intra mode, context encoding or decoding the first bin of the N bins using a first context; and in response to determining that a second bin of the N bins is assigned to a second intra mode, context encoding or decoding the second bin of the N bins using a second context different from the first context.

[0200] Example 9. The method according to any one of Examples 1-4 and 6-8, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin of the N bins is assigned to a planar mode, context encoding or decoding the first bin of the N bins using a first context; and in response to determining that a second bin of the N bins is assigned to the planar mode, context encoding or decoding the second bin of the N bins using a second context different from the first context.

[0201] Example 10. The method according to any one of Examples 1-4 and 6-9, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin among the N bins is assigned to a planar mode, context encoding or decoding the first bin among the N bins using a first context; and in response to determining that a second bin among the N bins is assigned to a DC mode, context encoding or decoding the second bin among the N bins using the first context.

[0202] Example 11. The method according to any one of Examples 1-4 and 6-10, wherein entropy encoding or decoding the N bins includes: in response to determining that a first bin among the N bins is assigned to a first angular mode within a first subset of angular modes, context encoding or decoding the first bin among the N bins using a first context assigned to the first subset of angular modes; in response to determining that a second bin among the N bins is assigned to a second angular mode within the first subset of angular modes, context encoding or decoding the second bin among the N bins using the first context; and in response to determining that a third bin among the N bins is assigned to a third angular mode within a second subset of angular modes different from the first subset of angular modes, context encoding or decoding the third bin among the N bins using a second context assigned to the second subset of angular modes that is different from the first context.

[0203] Example 12. The method according to any one of Examples 1-4 and 6-11, wherein entropy encoding or decoding the N bins includes: using a single context for all intra modes; and predicting or modifying the bin values based on the mode information for the respective bins among the N bins.

[0204] Example 13. The method according to any one of Examples 1-4 or 6-12, wherein determining the LFNST flag / index includes: determining a value for an intra mode for decoding the current block; context decoding the value of the last bin among the N bins to generate a decoded value; and predicting the value of the last bin based on the decoded value and the value for the intra mode.

[0205] Example 14. The method according to Example 13, wherein predicting the value of the last bin includes: determining whether the value for the intra mode is odd; in response to determining that the value for the intra mode is not odd, predicting the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is odd, predicting the value of the last bin as a value opposite to the decoded value.

[0206] Example 15. The method according to Example 13, wherein predicting the value of the last bin includes: determining whether the value for the intra mode is even; in response to determining that the value for the intra mode is not even, predicting the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is even, predicting the value of the last bin as a value opposite to the decoded value.

[0207] Example 16. An apparatus for encoding or decoding video data, the apparatus comprising: one or more units for performing the method according to any one of Examples 1-15.

[0208] Example 17. The apparatus according to Example 16, wherein the one or more units include one or more processors implemented in a circuit.

[0209] Example 18. The apparatus according to any one of Examples 16 and 17, further comprising: a memory for storing the video data.

[0210] Example 19. The apparatus according to any one of Examples 16-18, further comprising: a display configured to display the decoded video data.

[0211] Example 20. The apparatus according to any one of Examples 16-19, wherein the apparatus includes one or more of the following: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0212] Example 21. The apparatus according to any one of Examples 16-20, wherein the apparatus includes a video decoder.

[0213] Example 22. The apparatus according to any one of Examples 16-21, wherein the apparatus includes a video encoder.

[0214] Example 23. A computer-readable storage medium having instructions stored thereon, which when executed cause one or more processors to perform the method according to any one of Examples 1-15.

[0215] It should be recognized that, according to examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or completely omitted (e.g., not all described actions or events are necessary for implementing the techniques). Additionally, in certain examples, actions or events may be performed, for example, concurrently by multithreading, interrupt processing, or multiple processors rather than sequentially.

[0216] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium including any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0217] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0218] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the terms "processor" and "processing circuitry" as used herein may refer to any one of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Further, the techniques may be implemented entirely within one or more circuits or logic elements.

[0219] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but need not be implemented by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.

[0220] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprising: Receiving encoded data for a current block; Decoding N bins for a Low Frequency Non-Separable Transform (LFNST) index according to the encoded data, where the N bins include a first bin and a second bin, and where decoding the N bins includes context decoding each of the N bins; Using the N bins to determine the LFNST index, where determining the LFNST index includes: determining a value for an intra mode for decoding the current block; context decoding a value of a last one of the N bins to generate a decoded value; and determining the value of the last bin based on the decoded value and the value for the intra mode; Wherein determining the value of the last bin includes: Determining whether the value for the intra mode is odd; in response to determining that the value for the intra mode is not odd, predicting the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is odd, predicting the value of the last bin as a value opposite to the decoded value; or Determining whether the value for the intra mode is even; in response to determining that the value for the intra mode is not even, predicting the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is even, predicting the value of the last bin as a value opposite to the decoded value; Decoding the encoded data to generate transform coefficients; Applying an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and Reconstructing the current block of the video data using the residual block and a prediction block for the current block.

2. The method according to claim 1, wherein Performing context decoding includes: In response to determining that the first bin of the N bins is assigned to a luminance channel, context decoding the first bin of the N bins using a first context; and In response to determining that the second bin of the N bins is assigned to a chrominance channel, context decoding the second bin of the N bins using a second context different from the first context.

3. The method according to claim 1, wherein, Performing context decoding includes: Predicting a bin value based on intra mode information for a corresponding one of the N bins using a first context that is the same for all intra modes for the luminance channel or using a second context that is the same for all intra modes for the chrominance channel.

4. A method for encoding video data, the method comprising: Generating a residual block for a current block based on a prediction block for the current block; Applying a separable transform to the residual block to produce separable transform coefficients; Determining a Low Frequency Non-Separable Transform (LFNST) index; Applying an LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients; Encode the transform coefficients and the N bins for the LFNST index to produce encoded video data, where the N bins include a first bin and a second bin, and where encoding the N bins includes context encoding each of the N bins, where context encoding each bin includes: determining a value for an intra mode for encoding the current block; and context encoding the value of the last bin among the N bins based on the value for the intra mode; where context encoding the value of the last bin includes: determining whether the value for the intra mode is odd; in response to determining that the value for the intra mode is not odd, context encoding the value of the last bin; and in response to determining that the value for the intra mode is odd, context encoding a value opposite to the value of the last bin; or determining whether the value for the intra mode is even; in response to determining that the value for the intra mode is not even, context encoding the value of the last bin; and in response to determining that the value for the intra mode is even, context encoding a value opposite to the value of the last bin; and output the encoded video data.

5. The method according to claim 4, wherein Performing context encoding includes: in response to determining that the first bin among the N bins is assigned to a luminance channel, context encoding the first bin among the N bins using a first context; and in response to determining that the second bin among the N bins is assigned to a chrominance channel, context encoding the second bin among the N bins using a second context different from the first context.

6. The method according to claim 4, wherein Performing context encoding includes: using a first context that is the same for all intra modes for the luminance channel, or using a second context that is the same for all intra modes for the chrominance channel, context encoding the bin values based on the intra mode information for the respective bins among the N bins.

7. A device for decoding video data, the device including one or more processors implemented in circuitry and configured to perform the following operations: receive encoded data for a current block; Decode N bins for a low-frequency non-separable transform (LFNST) index based on the encoded data, wherein, the N bins include a first bin and a second bin, and where, to decode the N bins, the one or more processors are configured to context decode each of the N bins; use the N bins to determine the LFNST index, where, to determine the LFNST index, the one or more processors are configured to: determine a value for an intra mode for decoding the current block; context decode the value of the last bin among the N bins to generate a decoded value; and determine the value of the last bin based on the decoded value and the value for the intra mode; where, to determine the value of the last bin, the one or more processors are configured to: Determine whether the value for the intra mode is odd; in response to determining that the value for the intra mode is not odd, predict the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is odd, predict the value of the last bin as the value opposite to the decoded value; or Determine whether the value for the intra mode is even; in response to determining that the value for the intra mode is not even, predict the value of the last bin as the decoded value; and in response to determining that the value for the intra mode is even, predict the value of the last bin as the value opposite to the decoded value; Decode the encoded data to generate transform coefficients; Apply an inverse LFNST to the transform coefficients using the LFNST index to produce a residual block for the current block; and Reconstruct the current block of the video data using the residual block and the prediction block for the current block.

8. The device according to claim 7, wherein, For context decoding, the one or more processors are configured to: In response to determining that the first bin among the N bins is assigned to the luminance channel, perform context decoding on the first bin among the N bins using a first context; And In response to determining that the second bin among the N bins is assigned to the chrominance channel, perform context decoding on the second bin among the N bins using a second context different from the first context.

9. The device according to claim 7, wherein For context decoding, the one or more processors are configured to: Predict a bin value based on intra mode information for the corresponding bin among the N bins using a first context that is the same for all intra modes for the luminance channel or using a second context that is the same for all intra modes for the chrominance channel.

10. The apparatus according to claim 7, further comprising: A display configured to display the video data.

11. The device according to claim 7, wherein, The device includes one or more of the following: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

12. A device for encoding video data, the device including one or more processors implemented in circuitry and configured to perform the following operations: Generate a residual block for the current block based on a prediction block for the current block; Apply a separable transform to the residual block to produce separable transform coefficients; Determine a low-frequency non-separable transform (LFNST) index; Apply an LFNST to the separable transform coefficients using the LFNST index to produce transform coefficients; Encode the transform coefficients and the N bins for the LFNST index to produce encoded video data, where the N bins include a first bin and a second bin, and where, to encode the N bins, the one or more processors are configured to contextually encode each of the N bins, where, to contextually encode each bin, the one or more processors are configured to: determine a value for an intra mode for encoding the current block; and contextually encode the value of the last bin among the N bins based on the value for the intra mode; where, to contextually encode the value of the last bin, the one or more processors are configured to: determine whether the value for the intra mode is odd; in response to determining that the value for the intra mode is not odd, contextually encode the value of the last bin; and in response to determining that the value for the intra mode is odd, contextually encode a value opposite to the value of the last bin; or determine whether the value for the intra mode is even; in response to determining that the value for the intra mode is not even, contextually encode the value of the last bin; and in response to determining that the value for the intra mode is even, contextually encode a value opposite to the value of the last bin; and output the encoded video data.

13. The device according to claim 12, wherein, To perform contextual encoding, the one or more processors are configured to: in response to determining that the first bin among the N bins is assigned to a luminance channel, contextually encode the first bin among the N bins using a first context; and in response to determining that the second bin among the N bins is assigned to a chrominance channel, contextually encode the second bin among the N bins using a second context different from the first context.

14. The device according to claim 12, wherein, To perform contextual encoding, the one or more processors are configured to: contextually encode bin values based on intra mode information for the respective bin among the N bins, using a first context that is the same for all intra modes for the luminance channel or using a second context that is the same for all intra modes for the chrominance channel.

Citation Information

Patent Citations

  • Enhanced multiple transforms for prediction residual

    US10306229B2

  • Efficient parameter storage for compact multi-pass transforms

    US10349085B2

  • Coding adaptive multiple transform information for video coding

    US10986340B2

  • Non-separable secondary transform for video coding

    US20170094313A1

  • Non-separable secondary transform for video coding with reorganizing

    US20170094314A1