Zeroing Pattern Based Low-Frequency Inseparable Transform Signal Notification for Video Coding

By determining the last effective coefficient position of the transform block in video decoding and inferring the LFNST index value, the problem of large LFNST index signal notification overhead is solved, and the decoding efficiency is improved.

CN113812157BActive Publication Date: 2025-07-04QUALCOMM INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202080034606.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-13
Filing Date
2020-05-14
Publication Date
2025-07-04
Estimated Expiration
2040-05-14

AI Technical Summary

Technical Problem

In the existing video decoding technology, the signal notification overhead of low-frequency inseparable transform (LFNST) index is relatively large, affecting the decoding efficiency.

Method used

By determining the position of the last effective coefficient in the transform block, the value of the LFNST index is inferred based on its relative relationship with the zeroed area, thereby reducing signal notification and improving decoding efficiency.

Benefits of technology

Reduces signaling overhead and improves video decoding efficiency, especially in advanced video codecs using LFNST, such as extensions of HEVC and VVC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113812157B_ABST
    Figure CN113812157B_ABST
Patent Text Reader

Abstract

A video decoder is configured to determine the position of the last significant coefficient in a transform block of video data. The video decoder may then determine a value of a low-frequency non-separable transform (LFNST) index for the transform block based on the position of the last significant coefficient relative to a zeroed region of the transform block, where the zeroed region of the transform block includes both a first region of the transform block within the LFNST region and a second region of the transform block outside the LFNST region. The video decoder may then perform an inverse transform on the transform block according to the value of the LFNST index.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Patent Application No. 15 / 931,271, filed on May 13, 2020, which claims the benefit of U.S. Provisional Application No. 62 / 849,689, filed on May 17, 2019, the entire contents of each of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to video encoding and video decoding. Background Art

[0003] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called "smart phones", video teleconferencing devices, video streaming devices, etc. Digital video devices implement video decoding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC) and extensions of such standards. By implementing such video decoding techniques, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004] Video decoding techniques include spatial (intra) prediction and / or temporal (inter) prediction to reduce or eliminate redundancy inherent in a video sequence. For block-based video decoding, a video segment (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) segment of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) segment of a picture can use spatial prediction relative to reference samples in adjacent blocks in the same picture, or temporal prediction relative to reference samples in other reference pictures. A picture can be referred to as a frame, and a reference picture can be referred to as a reference frame. Summary of the Invention

[0005] In general, the present disclosure describes techniques for transform decoding, which is an essential element of modern video compression standards (High Efficiency Video Coding: Coding Tools and Specification by M. Wien, Springer-Verlag, Berlin, 2015). The techniques of the present disclosure include various transform signal notification methods that can be used in a video codec to specify a transform selected from among multiple transform candidates for decoding. Specifically, the present disclosure describes techniques for inferring the value of a low-frequency non-separable transform (LFNST) index from multiple values. Inferring means determining the value from among multiple values without receiving a syntax element in the encoded video bitstream that indicates the value.

[0006] The value of the LFNST index indicates whether the LFNST is to be applied to a transform block and, when applied, indicates the type of LFNST to be applied. The LFNST is a non-separable transform applied to an LFNST region of a transform block. The LFNST region can be a subset of the transform coefficients of the transform block and can include the low-frequency components of the transform block (e.g., the upper left corner of the transform block). In some applications, when the LFNST is applied, certain transform coefficients within the LFNST region are set to zero (e.g., zeroed-out). Additionally, transform coefficients outside the LFNST region in the transform block can also be zeroed out.

[0007] Before determining the value of the LFNST index for a transform block, the video decoder can be configured to determine the position of the last significant coefficient in the transform block. When the transform coefficients of the transform block are sorted / scanned according to a scan order, the last significant coefficient in the transform block can refer to the last non-zero transform coefficient of the transform block. For example, the video decoder can receive and decode a syntax element that indicates the position (e.g., the X and Y coordinates in the transform block) of the last significant (i.e., non-zero) coefficient along a predetermined scan order. In the case where the position of the last significant coefficient is determined to be in a portion of the transform block that would be zeroed out if the video encoder applied the LFNST (either in the LFNST region or outside the LFNST region), the video decoder can infer that the value of the LNFST index is zero (i.e., the LFNST is not applied). That is, in the case where the video decoder determines that there is a non-zero coefficient in a position in the transform block that would be zeroed out if the LFNST were applied (e.g., the transform coefficient would have a zero value), it can determine that the LFNST is not applied.

[0008] In this way, in the case where the position of the last significant coefficient in the transform block is in the part that would be set to zero if LFNST is applied (either in the LFNST region or outside the LFNST region), the video encoder does not need to generate and signal a syntax element indicating the value of the LFNST index. Therefore, the signaling overhead can be reduced, and the decoding efficiency can be improved. Since the techniques proposed in this disclosure can reduce the signaling overhead, the techniques of this disclosure can improve the decoding efficiency and can be used in advanced video codecs that use LFNST, including extensions of HEVC and next-generation video coding standards such as Versatile Video Coding (VVC) or H.266.

[0009] In one example, this disclosure describes a method for decoding video data, the method including: determining a position of a last significant coefficient in a transform block of the video data; determining a value of an LFNST index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, where the zeroing region of the transform block includes both a first region within an LFNST region of the transform block and a second region outside the LFNST region of the transform block; and performing an inverse transform on the transform block according to the value of the LFNST index.

[0010] In another example, this disclosure describes an apparatus configured to decode video data, the apparatus including: a memory configured to store a transform block of video data; and one or more processors in communication with the memory, the one or more processors configured to: determine a position of a last significant coefficient in a transform block of the video data; determine a value of an LFNST index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, where the zeroing region of the transform block includes both a first region within an LFNST region of the transform block and a second region outside the LFNST region of the transform block; and perform an inverse transform on the transform block according to the value of the LFNST index.

[0011] In another example, this disclosure describes an apparatus configured to decode video data, the apparatus including: a unit configured to determine a position of a last significant coefficient in a transform block of the video data; a unit configured to determine a value of an LFNST index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, where the zeroing region of the transform block includes both a first region within an LFNST region of the transform block and a second region outside the LFNST region of the transform block; and a unit configured to perform an inverse transform on the transform block according to the value of the LFNST index.

[0012] In another example, the present disclosure describes a non - transitory computer - readable storage medium storing instructions that, when executed, cause one or more processors configured to decode video data: to determine the position of the last significant coefficient in a transform block of the video data; to determine the value of the LFNST index of the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, where the zeroing region of the transform block includes both a first region of the transform block within the LFNST region and a second region of the transform block outside the LFNST region; and to perform an inverse transform on the transform block according to the value of the LFNST index.

[0013] One or more example details are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a block diagram illustrating an exemplary video encoding and decoding system that can execute the techniques of the present disclosure.

[0015] Figure 2A and Figure 2B is a conceptual diagram illustrating an exemplary quadtree - binary tree (QTBT) structure and a corresponding coding tree unit (CTU).

[0016] Figure 3 is a block diagram illustrating an exemplary video encoder that can execute the techniques of the present disclosure.

[0017] Figure 4 is a block diagram illustrating an exemplary video decoder that can execute the techniques of the present disclosure.

[0018] Figure 5 is a block diagram illustrating an exemplary low - frequency non - separable transform (LFNST) at an encoder and a decoder.

[0019] Figure 6 is a conceptual diagram illustrating transform coefficients obtained after applying LFNST to a transform block with zeroing.

[0020] Figure 7 is a conceptual diagram illustrating transform coefficients obtained after applying LFNST to a transform block without zeroing.

[0021] Figure 8 is a conceptual diagram illustrating transform coefficients obtained after applying an exemplary LFNST to a transform block with zeroing.

[0022] Figure 9It is a conceptual diagram showing transform coefficients obtained after applying an exemplary LFNST to a transform block without zeroing.

[0023] Figure 10 It is a flowchart showing an exemplary encoding method of the present disclosure.

[0024] Figure 11 It is a flowchart showing an exemplary decoding method of the present disclosure.

[0025] Figure 12 It is a flowchart showing another exemplary decoding method of the present disclosure. Detailed Description

[0026] The technology of the present disclosure includes various transform signal notification methods that can be used in a video codec to specify a transform selected from among multiple transform candidates for decoding. More specifically, the present disclosure describes techniques for inferring the value of a low-frequency non-separable transform (LFNST) index. Inferring means determining the value without receiving a syntax element indicating the value in the encoded video bitstream.

[0027] The value of the LFNST index indicates whether the LFNST is applied to the transform block, and when applied, indicates the type of LFNST to be applied. The LFNST is a non-separable transform applied to the LFNST region of the transform block. The LFNST region can be a subset of the transform coefficients of the transform block and can include the low-frequency components of the transform block (e.g., the upper left corner of the transform block). In some applications, when the LFNST is applied, certain transform coefficients within the LFNST region are set to zero (e.g., zeroed). Additionally, transform coefficients outside the LFNST region in the transform block can also be zeroed.

[0028] Before determining the value of the LFNST index for a transform block, the video decoder can be configured to determine the position of the last significant coefficient in the transform block. For example, the video decoder can receive and decode a syntax element indicating the position (e.g., the X and Y coordinates in the transform block) of the last significant (i.e., non-zero) coefficient along a predetermined scan order. If the position of the last significant coefficient is determined to be in a part of the transform block that would be zeroed if the video encoder applied the LFNST (either in the LFNST region or outside the LFNST region), then the video decoder can infer that the value of the LNFST index is zero (i.e., the LFNST is not applied). That is, in the case where the video decoder determines that there is a non-zero coefficient in a position in the transform block that would be zeroed if the LFNST were applied (e.g., the transform coefficient would have a zero value), it can determine that the LFNST is not applied.

[0029] In this way, in the case where the portion in the transform block where the position of the last valid coefficient would be set to zero if LFNST is applied (in the LFNST region or outside the LFNST region), the video encoder does not need to generate and signal a syntax element indicating the value of the LFNST index. Thus, the signaling overhead can be reduced and the decoding efficiency can be improved.

[0030] Figure 1 FIG. 4 is a block diagram showing an exemplary video encoding and decoding system 100 in which the techniques of the present disclosure may be implemented. The techniques of the present disclosure generally relate to decoding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data may include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0031] As Figure 1 shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 may include any of a variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablets, set-top boxes, cellular phones (such as smart phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and may thus be referred to as wireless communication devices.

[0032] In Figure 1 the example of FIG. 4, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. In accordance with the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 may be configured to apply techniques for transform decoding. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source device and the destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device rather than include an integrated display device.

[0033] As Figure 1The system 100 shown is merely an example. In general, any digital video encoding and / or decoding device can perform techniques for transform coding. The source device 102 and the destination device 116 are merely examples of such coding devices, where the source device 102 generates coded video data for transmission to the destination device 116. This disclosure refers to a "coding" device as a device that performs the coding (encoding and / or decoding) of data. Thus, the video encoder 200 and the video decoder 300 represent examples of coding devices, specifically, a video encoder and a video decoder. In some examples, the source device 102 and the destination device 116 may operate in a substantially symmetric manner such that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Thus, the system 100 may support one-way or two-way video transmission between the source device 102 and the destination device 116, e.g., for video streaming, video playback, video broadcasting, or video telephony.

[0034] In general, the video source 104 represents the source of video data (i.e., raw, uncoded video data) and provides a sequence of pictures (also referred to as "frames") of the video data to the video encoder 200, which encodes the data of the pictures. The video source 104 of the source device 102 may include a video capture device (such as a camera), a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 104 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may reorder the pictures from the received order (sometimes referred to as the "display order") to a coded order for encoding. The video encoder 200 may generate a bitstream including the coded video data. Then, the source device 102 may output the coded video data to a computer-readable medium 110 via the output interface 108 for reception and / or extraction by, e.g., the input interface 122 of the destination device 116.

[0035] The memories 106 of the source device 102 and 120 of the destination device 116 represent general memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from the video source 104 and raw decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable, for example, by the video encoder 200 and the video decoder 300, respectively. Although the memories 106, 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store, for example, encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, components of the memories 106, 120 may be allocated as one or more video buffers, e.g., for storing raw, decoded, and / or encoded video data.

[0036] The computer-readable medium 110 may represent any type of medium or device capable of transferring the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium for enabling the source device 102 to directly send the encoded video data to the destination device 116 in real time, e.g., via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, the output interface 108 may modulate the transmission signal including the encoded video data, and the input interface 122 may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device useful for facilitating communication from the source device 102 to the destination device 116.

[0037] In some examples, the source device 102 may output the encoded data from the output interface 108 to the storage device 112. Similarly, the destination device 116 may access the encoded data from the storage device 112 via the input interface 122. The storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.

[0038] In some examples, the source device 102 may output the encoded video data to a file server 114 or to another intermediate storage device that may store the encoded video data generated by the source device 102. The destination device 116 may access the video data stored in the file server 114 via streaming or downloading. The file server 114 may be any type of server device capable of storing the encoded video data and sending the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 may access the encoded video data in the file server 114 via any standard data connection, including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 may be configured to operate according to a streaming protocol, a download transfer protocol, or a combination thereof.

[0039] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any one of the various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded video data) according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples where the output interface 108 and the input interface 122 include a wireless transmitter and / or a wireless receiver, the output interface 108 and the input interface 122 may be configured to transmit data (such as encoded video data) according to other wireless standards such as the IEEE 802.11 specifications, the IEEE 802.15 specifications (e.g., ZigBee TM )), Bluetooth TM standards, etc.). In some examples, the source device 102 and / or the destination device 116 may include respective System-on-Chip (SoC) devices. For example, the source device 102 may include an SoC device to perform the functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device to perform the functions attributed to the video decoder 300 and / or the input interface 122.

[0040] The techniques of the present disclosure can be applied to video coding for supporting any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (such as Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0041] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream can include signaling information defined by the video encoder 200, which is also used by the video decoder 300, such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coding units (e.g., slices, pictures, groups of pictures, sequences, etc.). A display device 118 displays the decoded pictures of the decoded video data to a user. The display device 118 can represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0042] Although not shown in Figure 1 In some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder and can include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream that includes both audio and video in a common data stream. If applicable, the MUX-DEMUX unit can conform to the ITU H.223 multiplexer protocol or other protocols, such as the User Datagram Protocol (UDP).

[0043] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device. Devices including the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).

[0044] Video encoder 200 and video decoder 300 may operate according to a video coding standard such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC) or its extensions (such as multi-view and / or scalable video coding extensions). Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as the Joint Exploration Test Model (JEM) or ITU-T H.266, also known as the Versatile Video Coding (VVC) standard. A draft of the VVC standard is described in “Versatile Video Coding (Draft 5)” by Bross et al., Joint Video Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 14th Meeting, Geneva, Switzerland, March 19 - 27, 2019, JVET-N1001-v5 (hereinafter referred to as “VVC Draft 5”). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0045] Generally, video encoder 200 and video decoder 300 may perform block-based coding of pictures. The term “block” generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 may code video data represented in the YUV (e.g., Y, Cb, Cr) format. That is, instead of coding the red, green, and blue (RGB) data of the samples of a picture, video encoder 200 and video decoder 300 may code the luminance and chrominance components, where the chrominance components may include the red-difference and blue-difference chrominance components. In some examples, video encoder 200 converts the received RGB format data into a YUV representation before encoding, and video decoder 300 converts the YUV representation into the RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.

[0046] The present disclosure generally may relate to coding (e.g., encoding and decoding) of pictures to include a process of encoding or decoding data of a picture. Similarly, the present disclosure may relate to coding of blocks of a picture to include a process of encoding or decoding data of a block, e.g., prediction and / or residual coding. An encoded video bitstream generally includes a series of values for syntax elements that represent coding decisions (e.g., coding modes) and the partitioning of a picture into blocks. Thus, a reference to a syntax element for a picture or a block should generally be understood as coding the values of the syntax elements used to form the picture or the block.

[0047] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coder divides a CTU and a CU into four equal non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further divide PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the division of TUs. In HEVC, a PU represents inter-prediction data, while a TU represents residual data. A CU predicted intra-frame includes intra-frame prediction information, such as an intra-mode indicator.

[0048] As another example, video encoder 200 and video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) divides a picture into a plurality of coding tree units (CTUs). Video encoder 200 may divide a CTU according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partitioning types, such as the separation between the CUs, PUs, and TUs of HEVC. The QTBT structure includes two levels: a first level divided according to quadtree partitioning, and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0049] In the MTT partitioning structure, quadtree (QT) partitioning, binary tree (BT) partitioning, and / or one or more types of ternary tree (TT) (also referred to as trinary tree (TT)) partitioning may be used to partition a block. Ternary tree or trinary tree partitioning is a partitioning for dividing a block into three sub-blocks. In some examples, ternary tree or trinary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0050] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance component and the chrominance components, while in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for the two chrominance components (or two QTBT / MTT structures for each chrominance component).

[0051] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures according to HEVC. For illustrative purposes, a description of the techniques of the present disclosure is given with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video decoders configured to use quadtree partitioning, MTT partitioning, or other types of partitioning.

[0052] In some examples, a CTU includes: a coding tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of a picture having three sample arrays, or a CTB of samples of a monochrome picture or a picture decoded using three separate color planes and syntax structures for decoding samples. A CTB may be an NxN block of samples of some N value such that dividing the component into CTBs is a partition. A component is an array or a single sample from one of the three arrays (luminance and two chrominances) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array of samples of a picture in monochrome format. In some examples, a coding block is an M×N block of samples of some M and N values such that dividing the CTB into coding blocks is a partition.

[0053] Blocks (e.g., CTUs or CUs) in a picture can be grouped in various ways. As an example, a brick may refer to a rectangular region of a row of CTUs within a particular tile in a picture. A tile may be a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs having a height equal to the height of the picture and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular region of CTUs having a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to the width of the picture.

[0054] In some examples, a tile may be divided into multiple bricks, and each brick may include one or more rows of CTUs within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. However, a brick that is a proper subset of a tile may not be referred to as a tile.

[0055] Bricks in a picture may also be arranged in slices. A slice may be an integer number of bricks that are exclusively included in a single network abstraction layer (NAL) unit in a picture. In some examples, a slice includes multiple complete tiles or a consecutive sequence of complete bricks of only one tile.

[0056] The present disclosure interchangeably uses "N×N" and "N by N" to refer to the sample size of a block (e.g., a CU or other video block) in terms of vertical and horizontal dimensions, such as 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU can be arranged in rows and columns. Additionally, a CU does not have to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can include N×M samples, where M does not necessarily equal N.

[0057] Video encoder 200 encodes video data of a CU representing prediction information and / or residual information, as well as other information. The prediction information indicates how the CU is to be predicted to form a prediction block for the CU. The residual information generally represents the sample-by-sample difference between the CU samples before encoding and the prediction block.

[0058] To predict a CU, video encoder 200 can generally form a prediction block for the CU by inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from previously decoded pictures, while intra-frame prediction generally refers to predicting the CU based on previously decoded data in the same picture. To perform inter-frame prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search, for example, according to the difference between the CU and a reference block to identify a reference block that closely matches the CU. Video encoder 200 can use the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, video encoder 200 can use uni-directional prediction or bi-directional prediction to predict the current CU.

[0059] Some examples of VVC also provide an affine motion compensation mode, which can be regarded as an inter-frame prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).

[0060] To perform intra prediction, video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of VVC provide 67 intra prediction modes, including various directional modes, as well as planar mode and DC mode. Generally, video encoder 200 selects an intra prediction mode, which describes the neighboring samples of the current block for predicting the samples of the current block (e.g., the block of a CU). Assuming that video encoder 200 decodes CTUs and CUs in raster scan order (from left to right, top to bottom), such samples are typically above, top-left, or to the left of the current block in the same picture as the current block.

[0061] Video encoder 200 encodes data representing the prediction mode of the current block. For example, for an inter prediction mode, video encoder 200 may encode data indicating which one of the various available inter prediction modes is used and the motion information for the corresponding mode. For uni-directional or bi-directional inter prediction, for example, video encoder 200 may use the Advanced Motion Vector Prediction (AMVP) mode or the merge mode to encode the motion vectors. Video encoder 200 may use a similar mode to encode the motion vectors for the affine motion compensation mode.

[0062] After prediction (such as intra prediction or inter prediction of a block), video encoder 200 may compute the residual data of the block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block of the block formed using the corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, video encoder 200 may apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. Video encoder 200 produces transform coefficients after applying one or more transforms.

[0063] As described above, after any transform used to produce transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to possibly reduce the amount of data used to represent the transform coefficients to provide further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, video encoder 200 may round an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0064] After quantization, video encoder 200 may scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher-energy (and thus lower-frequency) transform coefficients at the front of the vector and lower-energy (and thus higher-frequency) transform coefficients at the back of the vector. In some examples, video encoder 200 may use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients in the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy encode the values of syntax elements that describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.

[0065] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be sent. The context may relate to, for example, whether the neighboring values of the symbol are zero values. Probability determination may be based on the context assigned to the symbol.

[0066] Video encoder 200 may also generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, or other syntax data, such as sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS), for video decoder 300, for example, in a picture header, a block header, or a slice header. Video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.

[0067] In this way, video encoder 200 may generate a bitstream including encoded video data that describes, for example, syntax elements that partition a picture into blocks (e.g., CUs) and prediction information and / or residual information for the blocks. Finally, video decoder 300 may receive the bitstream and decode the encoded video data.

[0068] Typically, video decoder 300 performs the inverse process of the process performed by video encoder 200 to decode the encoded video data in the bitstream. For example, video decoder 300 may use CABAC to decode the values of syntax elements in the bitstream in a way that is inverse but substantially similar to the CABAC encoding process of video encoder 200. The syntax elements may define partitioning information for partitioning a picture into CTUs, and each CTU is partitioned according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. The syntax elements may further define prediction information and residual information for blocks (e.g., CUs) of video data.

[0069] The residual information may be represented by, for example, quantized transform coefficients. Video decoder 300 may perform inverse quantization and inverse transformation on the quantized transform coefficients of a block to reproduce the residual block of the block. Video decoder 300 uses the signaled prediction mode (intra prediction or inter prediction) and related prediction information (e.g., motion information for inter prediction) to form a prediction block for the block. Video decoder 300 may then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 may perform additional processing, such as performing deblocking processing to reduce visual artifacts along the boundaries of the blocks.

[0070] According to the techniques of the present disclosure, video encoder 200 and video decoder 300 may be configured to: based on the pattern of zero coefficients defined by the specification in a video data block, not signal / infer the value of a low-frequency non-separable transform index or flag and transform the video data block according to the low-frequency non-separable transform index or flag. For example, video decoder 300 may be configured to: determine the position of the last significant coefficient in a transform block of video data, determine the value of the LFNST index for the transform block based on the relative relationship between the position of the last significant coefficient and the zeroing region of the transform block, where the zeroing region of the transform block includes both a first region of the transform block within the LFNST region and a second region of the transform block outside the LFNST region, and perform inverse transformation on the transform block according to the value of the LFNST index.

[0071] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" generally may refer to the communication of the value of a syntax element and / or other data used to decode the encoded video data. That is, video encoder 200 may signal the value of a syntax element in the bitstream. Generally, signaling refers to generating a value in the bitstream. As described above, source device 102 may transmit the bitstream to destination device 116 substantially in real time or non-real time (such as may occur when storing the syntax elements in storage device 112 for later extraction by destination device 116).

[0072] Figure 2A and Figure 2B is a conceptual diagram showing an exemplary quadtree binary tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quadtree splits, while dashed lines represent binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type (i.e., horizontal or vertical) is used. In this example, 0 represents a horizontal split, and 1 represents a vertical split. For quadtree splits, since a quadtree node splits a block horizontally and vertically into 4 sub-blocks of equal size, there is no need to indicate the split type. Thus, the video encoder 200 can encode and the video decoder 300 can decode syntax elements (such as split information) for the region tree level (i.e., solid lines) of the QTBT structure 130 and syntax elements (such as split information) for the prediction tree level (e.g., dashed lines) of the QTBT structure 130. The video encoder 200 can encode and the video decoder 300 can decode video data for the CUs represented by the terminal leaf nodes of the QTBT structure 130, such as prediction data and transform data.

[0073] Generally speaking, Figure 2B the CTU 132 can be associated with parameters that define the block sizes corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters can include the CTU size (representing the size of the CTU 132 in samples), the minimum quadtree size (MinQTSize, representing the minimum allowable quadtree leaf node size), the maximum binary tree size (MaxBTSize, representing the maximum allowable binary tree root node size), the maximum binary tree depth (MaxBTDepth, representing the maximum allowable binary tree depth), and the minimum binary tree size (MinBTSize, representing the minimum allowable binary tree leaf node size).

[0074] The root node of the QTBT structure corresponding to the CTU can have four child nodes at the first level of the QTBT structure, and each child node can be divided according to the quadtree partitioning. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure 130 represents such a node as including a parent node and child nodes with solid lines for the branches. If the node at the first level is not greater than the allowed maximum binary tree root node size (MaxBTSize), it can be further divided by the corresponding binary tree. The binary tree splitting of a node can be iterated until the nodes generated by the splitting reach the allowed minimum binary tree leaf node size (MinBTSize) or the allowed maximum binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents such a node as having dashed lines for the branches. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra prediction or inter prediction) and transformation without any further partitioning. As described above, a CU can also be referred to as a "video block" or a "block".

[0075] In an example of the QTBT partitioning structure, the CTU size is set to 128×128 (luminance samples and two corresponding 64×64 chrominance samples), the MinQTSize is set to 16×16, the MaxBTSize is set to 64×64, the MinBTSize (for width and height) is set to 4, and the MaxBTDepth is set to 4. The quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quadtree node is 128×128, it will not be further split by the binary tree because its size exceeds the MaxBTSize (64×64 in this example). Otherwise, the leaf quadtree node will be further divided by the binary tree. Thus, the quadtree leaf nodes are also the root nodes of the binary tree, and the depth of the binary tree is 0. When the depth of the binary tree reaches the MaxBTDepth (4 in this example), further splitting is not allowed. When the width of the binary tree node is equal to the MinBTSize (4 in this example), it means that further horizontal splitting is not allowed. Similarly, the height of the binary tree node being equal to the MinBTSize means that further vertical splitting is not allowed for this binary tree node. As described above, the leaf nodes of the binary tree are called CUs, and they are further processed for prediction and transformation without further partitioning.

[0076] Figure 3 is a block diagram showing an exemplary video encoder 200 that can perform the techniques of the present disclosure. Provided Figure 3This is for illustrative purposes and should not be considered a limitation of the various techniques widely exemplified and described in this disclosure. For illustrative purposes, this disclosure describes the video encoder 200 in the context of video coding standards such as the H.265 (HEVC) video coding standard and the developing H.266 (VCC) video coding standard. However, the techniques of this disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.

[0077] In Figure 3 the example of, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or processing circuits. Additionally, the video encoder 200 may include additional or alternative processors or processing circuits to perform these functions and other functions.

[0078] The video data memory 230 may store video data to be encoded by components of the video encoder 200. The video encoder 200 may receive the video data stored in the video data memory 230 from, for example, a video source 104 ( Figure 1 ). The DPB 218 may serve as a reference picture memory that stores reference video data for use by the video encoder 200 in predicting subsequent video data. The video data memory 230 and the DPB 218 may be formed from any one of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The video data memory 230 and the DPB 218 may be provided by the same storage device or separate storage devices. In various examples, the video data memory 230 may be on-chip with other components of the video encoder 200, as shown, or off-chip relative to those components.

[0079] In the present disclosure, a reference to the video data memory 230 should not be construed as being limited to a memory internal to the video encoder 200 (unless so specifically described), or a memory external to the video encoder 200 (unless so specifically described). Rather, a reference to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., the video data of the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage of the outputs of the various units of the video encoder 200.

[0080] illustrates Figure 3 the various units to assist in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and has pre-set operations that can be executed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be executed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuit are generally immutable. In some examples, one or more units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more units can be integrated circuits.

[0081] The video encoder 200 can include an arithmetic logic unit (ALU), a basic function unit (EFU), digital circuits, analog circuits, and / or a programmable core formed by programmable circuits. In an example where software executed by programmable circuits is used to perform the operations of the video encoder 200, the memory 106 ( Figure 1 ) can store the object code, i.e., the instructions, of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 can store such object code.

[0082] The video data memory 230 is configured to store received video data. The video encoder 200 can extract pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 can be the original video data to be encoded.

[0083] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include other functional units for performing video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, a block copy unit (which may be a component of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and so on.

[0084] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include: the division of CTUs into CUs, the prediction mode for the CUs, the transform type for the residual data of the CUs, the quantization parameter for the residual data of the CUs, and so on. The mode selection unit 202 may ultimately select the combination of encoding parameters with a rate-distortion value superior to other tested combinations.

[0085] The video encoder 200 may divide the pictures extracted from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within slices. The mode selection unit 202 may divide the CTUs of the pictures according to a tree structure (such as the QTBT structure, the MTT structure, or the quadtree structure of HEVC) as described above. As described above, the video encoder 200 may form one or more CUs by dividing the CTUs according to the tree structure. Such CUs are also commonly referred to as "video blocks" or "blocks".

[0086] Generally, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for the current block (e.g., the current CU, or the overlapping portion of the PU and TU in HEVC). For inter prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more reference blocks in one or more reference pictures (e.g., one or more previously encoded pictures stored in the DPB 218) that closely match. Specifically, the motion estimation unit 222 may calculate a value representing the similarity between a potential reference block and the current block, for example, according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), and so on. The motion estimation unit 222 generally may perform these calculations using the per-sample differences between the current block and the considered reference block. The motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.

[0087] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in the current picture. Then, the motion estimation unit 222 may provide the motion vectors to the motion compensation unit 224. For example, for uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors.

[0088] The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to extract data of the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate values of the prediction block according to one or more interpolation filters. Further, for bi-directional inter prediction, the motion compensation unit 224 may extract data of two reference blocks identified by the respective motion vectors and combine the extracted data, for example, by per-sample averaging or weighted averaging.

[0089] As another example, for intra prediction or intra prediction coding, the intra prediction unit 226 may generate a prediction block according to samples adjacent to the current block. For example, for the directional mode, the intra prediction unit 226 may generally mathematically combine values of adjacent samples and fill the calculated values across the current block along a defined direction to produce the prediction block. As another example, for the DC mode, the intra prediction unit 226 may calculate an average value of adjacent samples of the current block and generate the prediction block to include the obtained average value for each sample of the prediction block.

[0090] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives an original un-coded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the per-sample difference between the current block and the prediction block. The resulting per-sample difference defines the residual block of the current block. In some examples, the residual generation unit 204 may also determine differences between sample values in the residual block to generate the residual block using residual differential pulse coding modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form the residual generation unit 204.

[0091] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luminance prediction unit and a corresponding chrominance prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of a CU may refer to the size of the luminance decoding block of the CU, and the size of a PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a specific CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra prediction, and 2N×2N, 2N×N, N×2N, N×N, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitions of PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.

[0092] In an example where the mode selection unit 202 does not further divide a CU into PUs, each CU may be associated with a luminance decoding block and a corresponding chrominance decoding block. As described above, the size of a CU may refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0093] For other video decoding techniques (such as, as several examples, intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), the mode selection unit 202 generates a prediction block for the current block being encoded through corresponding units associated with the decoding technique. In some examples, such as palette mode decoding, the mode selection unit 202 may not generate a prediction block, but rather generate syntax elements that indicate the manner in which the block is to be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy encoding unit 220 for encoding.

[0094] As described above, the residual generation unit 204 receives the video data of the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the per-sample difference between the prediction block and the current block.

[0095] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form a transform coefficient block. For example, the transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 may perform multiple transforms on the residual block, such as a primary transform and a secondary transform, such as a rotation transform. In some examples, the transform processing unit 206 does not apply a transform to the residual block.

[0096] As will be explained in more detail below, in some examples, the transform processing unit 206 may be configured to apply both a low-frequency non-separable transform (LFNST) and one or more separable transforms (e.g., using multi-transform selection (MTS) techniques) to a transform block of video data. The transform processing unit 206 may first apply one or more separable transforms before applying the LFNST. In some examples, the transform processing unit 206 applies the LFNST to a subset of the transform coefficients of the transform block obtained after applying the separable transform. The subset of the transform coefficients of the transform block to which the LFNST is applied may be referred to as the LFNST region. The LFNST region may be the upper left portion of the transform block, which represents the lowest frequency transform coefficients of the transform block.

[0097] In conjunction with applying the LFNST, the transform processing unit 206 may further be configured to apply a zeroing process to a portion of the resulting transform coefficients in the LFNST region. The zeroing process simply makes the value of each transform coefficient in a specific region have a zero value. In one example, the transform processing unit 206 may zero the transform coefficients in the higher frequency region (e.g., the lower right corner) of the LFNST region. Additionally, in some examples, the transform processing unit 206 may also zero the transform coefficients outside the LFNST region in the transform block (e.g., the transform coefficients in the so-called MTS region).

[0098] If the transform processing unit 206 has applied the LFNST to the transform block, the video encoder 200 may generate a concurrent signal to notify the LFNST index syntax element. The value of the LFNST index syntax element may indicate a specific transform among the multiple transforms used when performing the LFNST. In other examples, the LFNST index may indicate that the LFNST is not applied (e.g., the LFNST index value is 0). The video encoder 200 may be configured to generate the LFNST index when the LFNST is applied. When the LFNST is not applied, the video encoder 200 may be configured to determine whether to signal the LFNST index.

[0099] For example, in a case where the position of the last significant (e.g., non-zero) transform coefficient is in a position in the transform block that would typically be zeroed if LFNST were applied, the video encoder 200 may determine not to signal the LFNST index. This is because the video encoder 200 will also generate and signal in the encoded video bitstream one or more syntax elements indicating the position of the last significant coefficient. Since the video decoder 300 will first receive and decode the position of the last significant coefficient, if the position of the last significant coefficient is in the zeroing region of the transform block, the video decoder 300 does not need to receive the LFNST index indicating that LFNST is not performed. Instead, the video decoder 300 may infer (e.g., determine in the absence of an explicit syntax element) that the value of the LFNST index is zero and that LFNST is not applied based on the position of the last significant coefficient. If the video encoder 200 does not apply LFNST, but the position of the last significant coefficient is not in the zeroing region, then in some examples, the video encoder 200 signals the LFNST index.

[0100] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients in the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and thus, the quantized transform coefficients may have lower precision than the original transform coefficients produced by the transform processing unit 206.

[0101] The inverse quantization unit 210 and the inverse transform processing unit 212 may respectively apply inverse quantization and inverse transform to the quantized transform coefficient block to reconstruct a residual block from the transform coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although there may be some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0102] The filter unit 216 may perform one or more filter operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In some examples, the operation of the filter unit 216 may be skipped.

[0103] Video encoder 200 stores the reconstructed blocks in DPB 218. For example, in an example where the operation of filter unit 216 is not required, reconstruction unit 214 may store the reconstructed blocks into DPB 218. In an example where the operation of filter unit 216 is required, filter unit 216 may store the filtered reconstructed blocks into DPB 218. Motion estimation unit 222 and motion compensation unit 224 may extract reference pictures from DPB 218, which are formed by reconstructed (and possibly filtered) blocks, for performing inter prediction on blocks of subsequent encoded pictures. Additionally, intra prediction unit 226 may use the reconstructed blocks in DPB 218 of the current picture to perform intra prediction on other blocks in the current picture.

[0104] Generally, entropy coding unit 220 may perform entropy coding on syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 may perform entropy coding on the quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 may perform entropy coding on prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from mode selection unit 202. Entropy coding unit 220 may perform one or more entropy coding operations on the syntax elements (which are another example of video data) to generate the entropy-coded data. For example, entropy coding unit 220 may perform context-adaptive variable length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another entropy coding operation on the data. In some examples, entropy coding unit 220 may operate in a bypass mode, in which no entropy coding is performed on the syntax elements.

[0105] Video encoder 200 may output a bitstream, which includes the entropy-coded syntax elements required for reconstructing the blocks of a segment or picture. Specifically, entropy coding unit 220 may output the bitstream.

[0106] The above operations are described for blocks. Such a description should be understood as being for the operations of luma decoding blocks and / or chroma decoding blocks. As described above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of a CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of a PU.

[0107] In some examples, operations performed on a luminance coding block need not be repeated for a chrominance coding block. As an example, operations for identifying a motion vector (MV) and a reference picture for a luminance coding block need not be repeated to identify the MV and reference picture for a chrominance block. Instead, the MV for the luminance coding block can be scaled to determine the MV for the chrominance block, and the reference picture can be the same. As another example, for a luminance coding block and a chrominance coding block, the intra prediction process can be the same.

[0108] As will be explained in more detail below, video encoder 200 represents an example of a device configured to encode video data, including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: infer (e.g., not encode or signal) the value of a low-frequency non-separable transform index or flag based on a pattern of zero coefficients defined in a video data block, and transform the video data block according to the low-frequency non-separable transform index or flag.

[0109] Figure 4 is a block diagram illustrating an exemplary video decoder 300 that may perform the techniques of the present disclosure. Provided Figure 4 is for illustrative purposes only and is not a limitation on the techniques widely illustrated and described in the present disclosure. For illustrative purposes, the present disclosure describes a video decoder 300 according to techniques of JEM, VVC, and HEVC. However, the techniques of the present disclosure may be performed by a video coding device configured according to other video coding standards.

[0110] In Figure 4 example, video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any one or all of CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 may be implemented in one or more processors or processing circuitry. Additionally, video decoder 300 may include additional or alternative processors or processing circuitry to perform these functions and other functions.

[0111] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units for performing prediction according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, a block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and so on. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0112] The CPB memory 320 may store video data to be decoded by components of the video decoder 300, such as an encoded video bitstream. The video data stored in the CPB memory 320 may be obtained, for example, from a computer-readable medium 110( Figure 1 ). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Moreover, the CPB memory 320 may store video data other than the syntax elements of decoded pictures, such as temporary data representing the outputs from the respective units of the video decoder 300. The DPB 314 generally stores decoded pictures, and the video decoder 300 may output the decoded pictures and / or use the decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed of any of a variety of storage devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The CPB memory 320 and the DPB 314 may be provided by the same storage device or separate storage devices. In various examples, the CPB memory 320 may be on-chip or off-chip relative to the other components of the video decoder 300.

[0113] Additionally or alternatively, in some examples, the video decoder 300 may extract decoded video data from the memory 120( Figure 1 ). That is, the memory 120 may store data in the CPB memory 320 as described above. Similarly, when some or all of the functions of the video decoder 300 are implemented in software to be executed by the processing circuitry of the video decoder 300, the memory 120 may store instructions to be executed by the video decoder 300.

[0114] Shown Figure 4 The various units shown to assist in understanding the operations performed by the video decoder 300. These units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Similar to Figure 3, A fixed - function circuit refers to a circuit that provides a specific function and has pre - set executable operations. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functions in executable operations. For example, a programmable circuit can execute software or firmware, and the software or firmware causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed - function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed - function circuit are generally immutable. In some examples, one or more units can be different circuit blocks (fixed - function or programmable), and in some examples, one or more units can be integrated circuits.

[0115] Video decoder 300 can include an ALU, EFU, digital circuits, analog circuits, and / or a programmable core formed by programmable circuits. In an example where the operations of video decoder 300 are performed by software executed on the programmable circuit, on - chip or off - chip memory can store the instructions (e.g., object code) of the software that video decoder 300 receives and executes.

[0116] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reproduce syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0117] Generally, video decoder 300 reconstructs pictures on a block - by - block basis. Video decoder 300 can perform reconstruction operations on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the “current block”).

[0118] Entropy decoding unit 302 can perform entropy decoding on: the syntax elements that define the quantized transform coefficients of the quantized transform coefficient block, and transform information such as quantization parameter (QP) and / or transform mode indication. Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization that inverse quantization unit 306 is to apply. Inverse quantization unit 306 can, for example, perform a bit - shift - left operation to inverse - quantize the quantized transform coefficients. Inverse quantization unit 306 can thereby form a transform coefficient block including transform coefficients.

[0119] After the inverse quantization unit 306 forms a transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.

[0120] As will be explained in more detail below, in some examples, the inverse transform processing unit 308 may be configured to apply both an inverse low-frequency non-separable transform (LFNST) and one or more inverse separable transforms (e.g., using multi-transform selection (MTS) techniques) to a transform block of video data. The inverse transform processing unit 308 may first apply the inverse LFNST before applying one or more separable inverse transforms. In some examples, the inverse transform processing unit 308 applies the inverse LFNST to a subset of the transform coefficients of the transform block obtained after inverse quantization. The subset of the transform coefficients of the transform block to which the inverse LFNST is applied may be referred to as the LFNST region. The LFNST region may be the upper left portion of the transform block, which represents the lowest frequency transform coefficients of the transform block.

[0121] As referred to above Figure 3 As explained, the transform processing unit 206 of the video encoder 200 may be configured to apply a zeroing process to a portion of the resulting transform coefficients in the LFNST region. The zeroing process simply causes the value of each transform coefficient in a specific region to have a zero value. In one example, the transform processing unit 206 may zero the transform coefficients in the higher frequency region (e.g., the lower right corner) of the LFNST region. Additionally, in some examples, the transform processing unit 206 may also zero the transform coefficients outside the LFNST region in the transform block (e.g., the coefficients in the so-called MTS region). Thus, the inverse transform processing unit 308 may be configured to zero (or ensure that zeroing has been performed) the transform coefficients in a specific region of the transform block when applying the LFNST.

[0122] As referred to above Figure 3 As discussed, if the transform processing unit 206 has applied the LFNST to a transform block, the video encoder 200 may generate a concurrent signal notifying the LFNST index syntax element. The value of the LFNST index syntax element from among multiple values may indicate a specific transform among multiple transforms used when performing the LFNST. In other examples, the LFNST index may indicate that the LFNST has not been applied (e.g., the LFNST index value is 0). The video encoder 200 may be configured to generate the LFNST index when applying the LFNST. When the LFNST is not applied, the video encoder 200 may be configured to determine whether to signal the LFNST index. Similarly, referring to Figure 4, the inverse transform processing unit 308 of the video decoder 300 may be configured not to receive the LFNST index in the encoded video bitstream in some cases. Instead, the inverse transform processing unit 308 of the video decoder 300 may infer the value of the LFNST index in some cases.

[0123] For example, in the case where the position of the last valid (e.g., non-zero) transform coefficient in the transform block is in a position that would typically be set to zero if LFNST were applied, the video encoder 200 may determine not to signal the LFNST index. This is because the video encoder 200 will also generate and signal in the encoded video bitstream one or more syntax elements indicating the position of the last valid coefficient. Since the video decoder 300 will first receive and decode the position of the last valid coefficient, if the position of the last valid coefficient is in the zeroing region of the transform block, the video decoder 300 does not need to receive the LFNST index indicating that LFNST is not performed. Instead, the inverse transform processing unit 308 of the video decoder 300 may infer (e.g., determine in the absence of an explicit syntax element) that the value of the LFNST index is zero and that LFNST is not applied.

[0124] Furthermore, the prediction processing unit 304 generates a prediction block according to the prediction information syntax element entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter prediction, the motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate: the reference picture in the DPB 314 from which the reference block is extracted, and the motion vector, which identifies the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally perform the inter prediction processing in a manner substantially similar to the manner described for the motion compensation unit 224 ( Figure 3 )).

[0125] As another example, if the prediction information syntax element indicates that the current block is intra prediction, the intra prediction unit 318 may generate a prediction block according to the intra prediction mode indicated by the prediction information syntax element. Similarly, the intra prediction unit 318 can generally perform the intra prediction processing in a manner substantially similar to the manner described for the intra prediction unit 226 ( Figure 3 ). The intra prediction unit 318 can extract the data of the neighboring samples of the current block from the DPB 314.

[0126] The reconstruction unit 310 can use the prediction block and the residual block to reconstruct the current block. For example, the reconstruction unit 310 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0127] The filter unit 312 can perform one or more filter operations on the reconstructed block. For example, the filter unit 312 can perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all examples.

[0128] The video decoder 300 can store the reconstructed block in the DPB 314. For example, in an example where the operations of the filter unit 312 are not required, the reconstruction unit 310 can store the reconstructed block in the DPB 314. In an example where the operations of the filter unit 312 are required, the filter unit 312 can store the filtered reconstructed block in the DPB 314. As described above, the DPB 314 can provide reference information to the prediction processing unit 304, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation. In addition, the video decoder 300 can output the decoded picture (e.g., decoded video) from the DPB 314 for subsequent presentation on a display device (such as Figure 1 display device 118).

[0129] In this way, as will be explained in more detail below, the video decoder 300 represents an example of a video decoding device that includes: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: infer (e.g., without decoding) the value of a low-frequency non-separable transform index or flag based on the pattern of zero coefficients defined in a video data block, and perform an inverse transform on the video data block according to the low-frequency non-separable transform index or flag.

[0130] In one example, the video decoder 300 can be configured to: determine the position of the last significant coefficient in a transform block of video data; determine the value of a low-frequency non-separable transform (LFNST) index for the transform block based on the relative relationship between the position of the last significant coefficient and the zeroing region of the transform block, where the zeroing region of the transform block includes both a first region of the transform block within the LFNST region and a second region of the transform block outside the LFNST region; and perform an inverse transform on the transform block according to the value of the LFNST index.

[0131] Overview of Transformation-related Tools

[0132] In exemplary video coding standards prior to HEVC, only fixed separable transforms or fixed separable inverse transforms were used in video coding and video decoding, where type 2 discrete cosine transform (DCT-2) was used vertically and horizontally. In HEVC, in addition to DCT-2, type 7 discrete sine transform (DST-7) is also adopted as a fixed separable transform for 4×4 blocks.

[0133] The following co-pending U.S. patents and U.S. patent applications describe multi-transform selection (MTS) techniques: U.S. Patent No. 10,306,229, published on May 28, 2019, U.S. Patent Publication No. 2018 / 0020218, published on January 18, 2018, and U.S. Patent Publication No. 2019 / 0373261, published on December 5, 2019. Note that MTS was previously known as adaptive multi-transform (AMT). The MTS technique is generally the same as the previously described AMT technique. An example of MTS described in U.S. Patent Publication No. 2019 / 0373261 is adopted in the Joint Exploration Model 7.0 (JEM-7.0) of the Joint Video Exploration Team (JVET) (e.g., see http: / / www.hhi.fraunhofer.de / fields-of-competence / image- processing / research-groups / image-video-coding / hevc-high-efficiency-video- coding / transform-coding-using-the-residual-quadtree-rqt.html ), and a simplified version of MTS is later adopted in VVC.

[0134] Generally, when encoding or decoding a transform block of transform coefficients using MTS, the video encoder 200 and the video decoder 300 can determine one or more separable transforms to be used among a plurality of separable transforms. By including more choices of separable transforms, the decoding efficiency can be improved because the selected transform(s) may be more suitable for the content being decoded.

[0135] Figure 5 is an illustration of an exemplary low-frequency non-separable transform (LFNST) on the encoder side and the decoder side (e.g., the video encoder 200 and the video decoder 300), where the use of LFNST introduces a new stage between the separable transform and quantization in the codec. As Figure 5 shown, on the encoder side (e.g., the video encoder 200), the transform processing unit 206 may first apply a separable transform 500 to the transform block to obtain transform coefficients. The transform processing unit 206 may then apply the LFNST 502 to a part of the transform coefficients of the transform block (e.g., the LFNST region). As described above, the transform processing unit 206 may perform zeroing processing in combination with the LFNST application. The quantization unit 208 may then quantize the resulting transform coefficients before entropy coding.

[0136] On the decoder side (e.g., the video decoder 300), the inverse quantization unit 306 first inverse quantizes the entropy-decoded transform coefficients in the transform block (see Figure 4 ). Then, the inverse transform processing unit 308 of the video decoder 300 applies the inverse LFNST 504 to the LFSNT region of the transform block. Then, the inverse transform processing unit 308 applies the inverse separable transform 506 to the result of the inverse LFNST to generate a residual block.

[0137] In JEM-7.0, an exemplary LFNST is used (e.g., as shown in Figure 5 ), to further improve the decoding efficiency of MTS, where the implementation of LFNST is based on the exemplary Hypercube Givens Transform (HyGT) described in U.S. Patent Publication No. 10,448,053 filed on Feb. 14, 2017. U.S. Patent No. 10,491,922 filed on Sep. 20, 2016, U.S. Patent Publication No. 2017 / 0094314 published on Mar. 30, 2017, U.S. Patent No. 10,349,085 filed on Feb. 14, 2017, and U.S. Patent Application No. 16 / 354,007 filed on Mar. 25, 2019 describe other exemplary designs and further details. Recently, LFNST has been adopted in the VVC standard (see JVET-N0193, Reduced Secondary Transform (RST) (CE6-3.1), available online at http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N0193-v5.zip . Note that LFNST was previously known as the Non-Separable Secondary Transform (NSST) or the Secondary Transform.

[0138] Zeroing Process in Current VVC

[0139] In the LFNST design in VVC Draft 5, the encoder (e.g., video encoder 200) may be configured to perform a zeroing operation that leaves the K lowest-frequency transform coefficients unchanged (e.g., the values of the K lowest-frequency transform coefficients are not zeroed). The K lowest-frequency transform coefficients are transformed by an LFNST of size N (e.g., for an 8×8 LFNST region N = 64). The decoder (e.g., video decoder 300) reconstructs the separable coefficients (e.g., MTS coefficients) by using only those K coefficients (also referred to as the K LFNST coefficients). In VVC Draft 5, this zeroing process is specified only for LFNSTs of sizes 4x4 and 8x8, where the decoder implicitly infers (determines without being signaled or assumed) that the values of the remaining N-K higher-frequency transform coefficients are set to zero values, and the K LFNST coefficients are used for reconstruction.

[0140] Figure 6 is a representative illustration of the transform coefficients obtained after applying an LFNST of size N to a transformed block 602 with zeroing of size HxW, where Z of the N transform coefficients are zeroed and K coefficients are retained. As shown in Figure 6As shown, video encoder 200 applies a separable transform (e.g., using MTS technology) to transform block 602 to obtain MTS coefficients. Video encoder 200 then applies LFNST to the LFNST region 600 (having a size of h×w) of transform block 602. The dark region 601 of LFNST region 600 is the K coefficients (e.g., LFNST coefficients) that are retained. The white region of LFNST region 600 is the Z(N-K) coefficients that are set to zero (e.g., coefficients set to zero).

[0141] As described in U.S. Patent No. 10,491,922, filed on September 20, 2016, U.S. Patent Publication No. 2017 / 0094314, published on March 30, 2017, and U.S. Provisional Application No. 62 / 799,410, filed on January 31, 2019, LFNST can be performed as follows: First, convert a 2-D sub-block that is the LFNST region (e.g., Figure 6 the LFNST region 600 in ) to a 1-D list (or vector) of transform coefficients via a predefined scan / sorting and then apply a transform to a subset of the transform coefficients (e.g., transform coefficients that are not set to zero).

[0142] Figure 7 shows an example of separable transform coefficients (MTS) and LFNST coefficients obtained without any zeroing. As Figure 7 shown, video encoder 200 applies a separable transform (e.g., using MTS technology) to transform block 702 (having a size of HxW) to obtain MTS coefficients. Video encoder 200 then applies LFNST to the LFNST region 700 (having a size of h×w) of transform block 702. In Figure 7 the example of, all N coefficients of LFNST region 700 are retained (e.g., LFNST coefficients). That is, no zeroing is performed in Figure 7 the example of.

[0143] This disclosure describes various techniques that can address the problems of signaling overhead and complexity associated with prior LFNST techniques. The techniques of this disclosure can (i) reduce the signaling overhead of LFNST indices / flags and (ii) simplify the LFNST process by extending the zeroing of separable transform coefficients. Extending the zeroing region of separable transform coefficients allows a VVC-like codec (e.g., video decoder 300) to infer LFNST indices / flags based on existing syntax related to coefficient decoding (e.g., syntax for determining the last position of valid (e.g., non-zero) coefficients).

[0144] Although the signaling method described in the present disclosure is described with reference to LFNST, the techniques of the present disclosure are not limited to LFNST and can be applied to reduce the signaling of other transform-related syntax.

[0145] LFNST Signal Notification Technology

[0146] Video encoder 200 and video decoder 300 may be configured to use the following LFNST signaling techniques individually or in any combination. In the context of the present disclosure, signaling may refer to video encoder 200 encoding one or more syntax elements and / or flags in one or more syntax structures (e.g., headers or parameter sets). In a reciprocal manner, video decoder 300 may receive and decode such syntax elements and / or flags. In some examples, video decoder 300 may be configured to infer the values of certain syntax elements and / or flags without explicitly receiving them in the bitstream.

[0147] In some examples, video encoder 200 and video decoder 300 are configured to apply LFNST with canonical zeroing. In this context, canonical zeroing defines which regions of the transform block (e.g., regions inside and outside the LFNST region) are zeroed. Based on a predefined set of conditions (e.g., block size, block shape, and / or transform-related syntax such as MTS index / flag indicating separable transform), canonical zeroing is applied at both video encoder 200 and video decoder 300. When video encoder 200 and video decoder 300 are configured to apply LFNST with canonical zeroing, video decoder 300 may be configured to infer the LFNST index / flag directly based on the pattern of zero coefficients defined by the canon. Thus, video encoder 200 does not need to signal the LFNST index / flag.

[0148] For example, the pattern / shape of the zeroed region (e.g., see the white region of LFNST region 600 in Figure 6 ) may change according to a predefined set of rules (e.g., block size, block shape, and / or transform-related syntax such as MTS index / flag). Video decoder 300 may be configured to infer the value of the LFNST index / flag based on the observed pattern, and the LFNST index / flag may not be explicitly signaled by video encoder 200. In some examples, the LFNST flag may indicate whether LFNST is applied (e.g., LFNST flag = 1) or whether LFNST is not applied (LFNST flag = 0). In other examples, the LFNST index may indicate that LFNST is not applied (LFNST index = 0), or if LFNST is applied, it may indicate a specific type of LFNST to be applied (LFNST index > 0).

[0149] In one example, if the video decoder 300 determines that a non-zero coefficient is in a position that should be zeroed when using LFNST, the video decoder 300 may infer that LFNST is not applied (e.g., infer that the value of the LFNST index is zero). In this case, the video decoder 300 may infer that the value of the LFNST index / flag is 0, which corresponds to not applying LFNST. For example, if the position of the last non-zero coefficient is in the zeroing region of the transform block, the video decoder 300 may determine that the value of the LFNST index is zero. As will be explained below, the zeroing region may be a zeroing region inside the LFNST region of the transform block and / or a zeroing region outside the LFNST region of the transform block.

[0150] The video decoder 300 may be configured to determine the position of the last significant coefficient because the video encoder 200 may generate and signal in the encoded video bitstream one or more syntax elements indicating the position of the last significant coefficient. Since the video decoder 300 will first receive and decode the position of the last significant coefficient (e.g., before determining whether to apply LFNST), if the position of the last significant coefficient is in the zeroing region of the transform block, the video decoder 300 does not need to receive the LFNST index indicating that LFNST is not performed. Instead, the video decoder 300 may infer (e.g., determine without an explicit syntax element) that the value of the LFNST index is zero and that LFNST is not applied.

[0151] In VVC draft 5, canonical zeroing is used for the 4×4 and 8×8 LFNST regions of the transform block (e.g., as Figure 6 shown), where a subset of the coefficients within the LFNST region are canonically zeroed. As described in the co-pending U.S. Provisional Application No. 62 / 799,410, filed on January 31, 2019, separable transform coefficients outside the LFNST region (e.g., MTS coefficients outside the LFNST region) may also be zeroed (e.g., as Figure 8 shown). Figure 8 is an illustration of the LFNST coefficients obtained by applying the LFNST of size N and zeroing Z coefficients (e.g., the highest frequency coefficients) in the LFNST region 800 (of size hxw) of the transform block 802 (of size HxW) and also zeroing the MTS coefficients outside the LFNST region 800. The dark region 801 of the LFNST region 800 is the K coefficients (e.g., LFNST coefficients) that are retained.

[0152] In this case, the video encoder 200 and the video decoder 300 may also utilize the zeroing mode to infer and / or not signal the LFNST index / flag, as described below. In one example, if there is at least one non-zero coefficient in the zeroing region, the video decoder 300 may infer that LFNST is not applied and derive the corresponding LFNST index / flag value as (for example) 0. In Figure 8 , the zeroing region may be both within and outside the LFNST region 800 of the transform block 802.

[0153] In another example, the video decoder 300 may use existing side information to infer the value of the LFNST index / flag. For example, the video decoder 300 may use existing last significant coefficient position information (e.g., a syntax element indicating the position of the last significant coefficient) to infer the value of the LFNST index / flag. In VVC, the video encoder 200 may be configured to signal two syntax elements respectively indicating the last significant coefficient positions in the X and Y (horizontal and vertical) directions. The syntax element indicating the last significant coefficient position may indicate whether there is a non-zero (significant) coefficient in the zeroing region.

[0154] As a specific example, if the last significant coefficient position signaling (i.e., the (X, Y) coordinates in the transform block) points to a position in the zeroing region (e.g., inside or outside the LFNST region as in Figure 8 ), the video decoder 300 may infer that the value of the LFNST index / flag is (for example) 0 and not apply LFNST. In some examples, the last significant coefficient position may be defined in one dimension (e.g., using the index of a 1-D list of LFNST coefficients) rather than in 2-D coordinates (X, Y).

[0155] Thus, in view of the above examples, the video decoder 300 may be configured to determine the position of the last significant coefficient in the transform block of the video data. For example, the video decoder 300 may be configured to decode one or more syntax elements indicating the X position and the Y position of the last significant coefficient in the transform block. The video decoder 300 may then determine the value of the low-frequency non-separable transform (LFNST) index for the transform block based on the relative relationship between the position of the last significant coefficient and the zeroing region of the transform block.

[0156] According to Figure 8 's example, the zeroing region of the transform block includes a first region within the LFNST region 800 of the transform block 802 (e.g., the white region in the LFNST region 800) and a second region outside the LFNST region (800) of the transform block 802. The value of the LFNST index indicates whether LFNST is applied to the transform block and, if applied, indicates the type of LFNST applied.

[0157] In a particular example, when the position of the last valid coefficient in a transform block is within the zeroing region of the transform block, video decoder 300 may infer that the value of the LFNST index is zero, where a value of zero for the LFNST index indicates that LFNST is not applied to the transform block. That is, video decoder 300 may be configured to infer that the value of the LFNST index is zero without receiving a syntax element indicating the value of the LFNST index.

[0158] In another example, to determine the value of the LFNST index, video decoder 300 may be configured to receive a syntax element indicating the LFNST index and decode the syntax element to determine the value of the LFNST index when the position of the last valid coefficient in the transform block is not within the zeroing region of the transform block.

[0159] Video decoder 300 may then perform an inverse transform on the transform block according to the value of the LFNST index. In one example, to perform an inverse transform on the transform block, video decoder 300 may perform an inverse transform on the LFNST region of the transform block with the LFNST indicated by the LFNST index and perform an inverse transform on the transform block with one or more separable transforms after performing the inverse transform on the LFNST region of the transform block with the LFNST. In another example, video decoder 300 may not apply LFNST and instead may perform an inverse transform on the transform block only with one or more separable transforms. Whether or not LFNST is used, video decoder 300 can perform an inverse transform on the transform block to generate a residual block, determine a prediction block for the residual block (e.g., using prediction techniques such as inter prediction or intra prediction), and combine the prediction block and the residual block to generate a decoded block.

[0160] For cases where zeroing is not used for LFNST coefficients, video encoder 200 and video decoder 300 may still apply zeroing to separable transform coefficients outside the LFNST region (e.g., MTS coefficients outside the LFNST region), as Figure 9 shown. Figure 9 is an illustration of LFNST coefficients obtained by applying an LFNST of size N and zeroing only the MTS coefficients outside the LFNST region 900 (having a size of h×w) of the transform block 902 (having a size of HxW). Then, video encoder 200 and video decoder 300 may infer the value of the LFNST index / flag based on the position of non-zero (valid) coefficients by using one of the methods or a combination of the methods described above.

[0161] Figure 10 is a flowchart showing an exemplary method for encoding a current block. The current block may include a current CU. Although for video encoder 200 ( Figure 1 and3 ) has been described, but it should be understood that other devices may be configured to perform similar methods as Figure 10 .

[0162] In this example, video encoder 200 initially predicts a current block (350). For example, video encoder 200 may form a predicted block of the current block. Video encoder 200 may then calculate a residual block (352) of the current block. To calculate the residual block, video encoder 200 may calculate the difference between the original unencoded block and the predicted block of the current block. Video encoder 200 may then transform and quantize the coefficients of the residual block (354). Next, video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, video encoder 200 may perform entropy encoding on the coefficients (358). For example, video encoder 200 may use CAVLC or CABAC to encode the coefficients. Video encoder 200 may then output the entropy-coded data of the block (360).

[0163] Figure 11 is a flowchart showing an exemplary method for decoding a current block of video data. The current block may include a current CU. Although described with respect to video decoder 300 ( Figure 1 and 4 ), it should be understood that other devices may be configured to perform similar methods as Figure 11 .

[0164] Video decoder 300 may receive the entropy-coded data of the current block, such as the entropy-coded prediction information and the entropy-coded data of the coefficients of the residual block corresponding to the current block (370). Video decoder 300 may perform entropy decoding on the entropy-coded data to determine the prediction information of the current block and reproduce the coefficients of the residual block (372). Video decoder 300 may predict the current block (374), for example, using an intra or inter prediction mode indicated by the prediction information of the current block to calculate a predicted block of the current block. Video decoder 300 may then perform inverse scanning on the reproduced coefficients (376) to produce a block of quantized transform coefficients. Video decoder 300 may then perform inverse quantization and inverse transformation on the coefficients to produce a residual block (378). Video decoder 300 may finally decode the current block by combining the predicted block and the residual block (380).

[0165] Figure 12 is a flowchart showing an exemplary decoding method of the present disclosure. Figure 12 's technology further defines Figure 11 's process 378. Figure 12 's technology may be performed by one or more structural units of video decoder 300, including inverse transform processing unit 308.

[0166] In one example of the present disclosure, the video decoder 300 may be configured to determine the position of the last significant coefficient in a transform block of video data (1200). For example, the video decoder 300 may be configured to decode one or more syntax elements indicating the X position and the Y position of the last significant coefficient in the transform block. The video decoder 300 may then determine the value of a Low-Frequency Non-Separable Transform (LFNST) index for the transform block based on the relative relationship between the position of the last significant coefficient and the zeroing region of the transform block (1202).

[0167] According to Figure 8 the example, the zeroing region of the transform block includes a first region (e.g., the white region in the LFNST region 800) within the LFNST region 800 of the transform block 802 and a second region outside the LFNST region 800 of the transform block 802. The value of the LFNST index indicates whether the LFNST is applied to the transform block, and if so, indicates the type of LFNST applied.

[0168] In a particular example, when the position of the last significant coefficient in the transform block is within the zeroing region of the transform block, the video decoder 300 may infer that the value of the LFNST index is zero, where a value of zero for the LFNST index indicates that the LFNST is not applied to the transform block. That is, the video decoder 300 may be configured to infer that the value of the LFNST index is zero without receiving a syntax element indicating the value of the LFNST index.

[0169] In another example, to determine the value of the LFNST index, the video decoder 300 may be configured to receive a syntax element indicating the LFNST index and decode the syntax element to determine the value of the LFNST index when the position of the last significant coefficient in the transform block is not within the zeroing region of the transform block.

[0170] The video decoder 300 may then perform an inverse transform (1204) on the transform block according to the value of the LFNST index. In one example, to perform an inverse transform on the transform block, the video decoder 300 may perform an inverse transform on the LFNST region of the transform block using one of the multiple LFNSTs indicated by the LFNST index, and perform an inverse transform on the transform block using one or more separable transforms after performing the inverse transform on the LFNST region of the transform block using the LFNST. In another example, the video decoder 300 may not apply the LFNST, but instead perform an inverse transform on the transform block only using one or more separable transforms. Whether or not the LFNST is used, the video decoder 300 may perform an inverse transform on the transform block to generate a residual block, determine a prediction block of the residual block (e.g., using a prediction technique such as inter prediction or intra prediction), and combine the prediction block and the residual block to generate a decoded block.

[0171] Other illustrative examples of the present disclosure are described below.

[0172] Example 1 - A method for decoding video data, the method comprising: inferring a value of a low-frequency non-separable transform index or flag based on a pattern of zero coefficients defined in a video data block; and performing a transform on the video data block according to the low-frequency non-separable transform index or flag.

[0173] Example 2 - The method of Example 1, wherein the pattern of the zero coefficients defined in the video data block is a pattern of a zeroed region of the video data block.

[0174] Example 3 - The method of Example 2, wherein inferring the value of the low-frequency non-separable transform index or flag comprises: inferring the value of the low-frequency non-separable transform index or flag as zero in a case where non-zero coefficients are in the zeroed region of the video data block.

[0175] Example 4 - The method of Example 2, wherein inferring the value of the low-frequency non-separable transform index or flag comprises: inferring the value of the low-frequency non-separable transform index or flag as zero in a case where the last valid coefficient position information indicates that non-zero coefficients are in the zeroed region of the video data block.

[0176] Example 5 - The method of any one of Examples 1-4, wherein decoding comprises decoding.

[0177] Example 6 - The method of any one of Examples 1-4, wherein decoding comprises encoding.

[0178] Example 7 - An apparatus for decoding video data, the apparatus comprising one or more units for performing the method of any one of Examples 1 to 6.

[0179] Example 8 - The device of Example 7, wherein the one or more units include one or more processors implemented in a circuit.

[0180] Example 9 - The device of any one of Examples 7 and 8, further comprising: a memory for storing the video data.

[0181] Example 10 - The device of any one of Examples 7 - 9, further comprising: a display configured to display the decoded video data.

[0182] Example 11 - The device of any one of Examples 7 - 10, wherein the device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set - top box.

[0183] Example 12 - The device of any one of Examples 7 - 11, wherein the device includes a video decoder.

[0184] Example 13 - The device of any one of Examples 7 - 12, wherein the device includes a video encoder.

[0185] Example 14 - A computer - readable storage medium having instructions stored thereon that, when executed, cause one or more processors to perform the method of any one of Examples 1 - 6.

[0186] It should be recognized that, according to the examples, certain operations or events of any of the techniques described herein may be performed in a different order, may be added, combined, or completely omitted (e.g., not all described operations or events are necessary to implement the techniques). Additionally, in some examples, the operations or events may be performed concurrently rather than sequentially, such as through multithreading, interrupt handling, or multiple processors.

[0187] In one or more exemplary aspects, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer - readable medium and executed by a hardware - based processing unit. Computer - readable media includes: computer storage media, which corresponds to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another according to a communication protocol. In this manner, computer - readable media generally may correspond to (1) a non - transitory tangible computer - readable storage medium, or (2) a communication medium such as a signal or a carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer - readable medium.

[0188] By way of example and not limitation, such a computer-readable storage medium may include one or more of the following: RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection can be properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, wireless, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, wireless, and microwave are also included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disk and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks typically reproduce data optically using lasers. Combinations of the above are also included within the scope of computer-readable media.

[0189] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein, the terms "processor" and "processing circuitry" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Similarly, the techniques can be fully implemented in one or more circuits or logic elements.

[0190] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a group of ICs (e.g., a chipset). In the present disclosure, various components, modules, or units are described to emphasize the various functional aspects of the devices configured to perform the disclosed techniques, but need not be implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above in conjunction with appropriate software and / or firmware.

[0191] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprising: Determining a position of a last significant coefficient in a transform block of video data by decoding a syntax element indicating a position of the last significant coefficient along a predetermined scan order; Determining a value of a Low-Frequency Non-Separable Transform (LFNST) index based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, wherein the zeroing region of the transform block includes both a first region of the transform block within an LFNST region and a second region of the transform block outside the LFNST region, wherein determining the value of the LFNST index includes: inferring the value of the LFNST index to be zero when the position of the last significant coefficient in the transform block is within the zeroing region of the transform block, wherein a value of zero for the LFNST index indicates that the LFNST is not applied to the transform block; and Performing an inverse transform on the transform block according to the value of the LFNST index.

2. The method according to claim 1, wherein, The value of the LFNST index indicates whether the LFNST is applied to the transform block, and if so, indicates a type of the applied LFNST.

3. The method according to claim 1, wherein Inferring the value of the LFNST index to be zero includes: Inferring the value of the LFNST index to be zero when not receiving a syntax element indicating the value of the LFNST index.

4. The method according to claim 1, wherein, Performing an inverse transform on the transform block includes: Performing an inverse transform on the transform block using one or more separable transforms.

5. The method according to claim 1, wherein Determining the value of the LFNST index includes: Receiving a syntax element indicating the LFNST index when the position of the last significant coefficient in the transform block is not within the zeroing region of the transform block; and Decoding the syntax element to determine the value of the LFNST index.

6. The method according to claim 5, wherein Performing an inverse transform on the transform block includes: Performing an inverse transform on the LFNST region of the transform block using the LFNST indicated by the LFNST index; and After performing the inverse transform on the LFNST region of the transform block using the LFNST, performing an inverse transform on the transform block using one or more separable transforms.

7. The method according to claim 1, wherein Determining the position of the last significant coefficient in the transform block of video data includes: Decoding one or more syntax elements indicating an X position and a Y position of the last significant coefficient in the transform block.

8. The method according to claim 1, wherein Performing an inverse transform on the transform block includes performing an inverse transform on the transform block to generate a residual block, and the method further includes: Determining a prediction block for the residual block; and Combining the prediction block and the residual block to generate a decoded block.

9. The method according to claim 8, further comprising: Displaying a picture including the decoded block.

10. An apparatus configured to decode video data, the apparatus comprising: A memory configured to store a transform block of video data; And One or more processors in communication with the memory, the one or more processors being configured to: Determine the position of the last significant coefficient in the transform block of the video data by decoding a syntax element that indicates the position of the last significant coefficient along a predetermined scan order; Determine the value of a Low-Frequency Non-Separable Transform (LFNST) index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, wherein the zeroing region of the transform block includes both a first region of the transform block within an LFNST region and a second region of the transform block outside the LFNST region, and wherein, to determine the value of the LFNST index, the one or more processors are configured to: infer the value of the LFNST index as zero in a case where the position of the last significant coefficient in the transform block is within the zeroing region of the transform block, wherein a value of zero for the LFNST index indicates that the LFNST is not applied to the transform block; and Perform an inverse transform on the transform block according to the value of the LFNST index.

11. The device according to claim 10, wherein, The value of the LFNST index indicates whether the LFNST is applied to the transform block and, if so, indicates the type of LFNST applied.

12. The device according to claim 10, wherein, To infer the value of the LFNST index as zero, the one or more processors are configured to: Infer the value of the LFNST index as zero without receiving a syntax element that indicates the value of the LFNST index.

13. The device according to claim 10, wherein, To perform an inverse transform on the transform block, the one or more processors are configured to: Perform an inverse transform on the transform block using one or more separable transforms.

14. The apparatus according to claim 10, wherein, To determine the value of the LFNST index, the one or more processors are configured to: Receive a syntax element that indicates the LFNST index in a case where the position of the last significant coefficient in the transform block is not within the zeroing region of the transform block; And Decode the syntax element to determine the value of the LFNST index.

15. The apparatus according to claim 14, wherein, To perform an inverse transform on the transform block, the one or more processors are configured to: Perform an inverse transform on the LFNST region of the transform block using the LFNST indicated by the LFNST index; and After performing the inverse transform on the LFNST region of the transform block using the LFNST, perform an inverse transform on the transform block using one or more separable transforms.

16. The device according to claim 10, wherein, To determine the position of the last significant coefficient in the transform block of the video data, the one or more processors are configured to: Decode one or more syntax elements that indicate an X position and a Y position of the last significant coefficient in the transform block.

17. The device according to claim 10, wherein, To perform an inverse transform on the transform block, the one or more processors are configured to perform an inverse transform on the transform block to generate a residual block, and wherein the one or more processors are configured to: Determine a prediction block for the residual block; and Combine the prediction block and the residual block to generate a decoded block.

18. The apparatus according to claim 17, further comprising: A display configured to display a picture including the decoded block.

19. An apparatus configured to decode video data, the apparatus comprising: a unit configured to determine a position of a last significant coefficient in a transform block of video data by decoding a syntax element indicating the position of the last significant coefficient along a predetermined scan order; a unit configured to determine a value of a Low-Frequency Non-Separable Transform (LFNST) index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, wherein the zeroing region of the transform block includes both a first region of the transform block within an LFNST region and a second region of the transform block outside the LFNST region, and wherein the unit configured to determine the value of the LFNST index includes: a unit configured to infer the value of the LFNST index as zero when the position of the last significant coefficient in the transform block is within the zeroing region of the transform block, wherein a value of zero for the LFNST index indicates that the LFNST is not applied to the transform block; and a unit configured to perform an inverse transform on the transform block according to the value of the LFNST index.

20. The apparatus according to claim 19, wherein, The value of the LFNST index indicates whether the LFNST is applied to the transform block and, if so, indicates the type of LFNST applied.

21. The device according to claim 19, wherein The unit configured to infer the value of the LFNST index as zero includes: a unit configured to infer the value of the LFNST index as zero when not receiving a syntax element indicating the value of the LFNST index.

22. The apparatus according to claim 19, wherein, The unit configured to perform an inverse transform on the transform block includes: a unit configured to perform an inverse transform on the transform block using one or more separable transforms.

23. The device according to claim 19, wherein The unit configured to determine the value of the LFNST index includes: a unit configured to receive a syntax element indicating the LFNST index when the position of the last significant coefficient in the transform block is not within the zeroing region of the transform block; and a unit configured to decode the syntax element to determine the value of the LFNST index.

24. The device according to claim 23, wherein, The unit configured to perform an inverse transform on the transform block includes: a unit configured to perform an inverse transform on the LFNST region of the transform block using the LFNST indicated by the LFNST index; and a unit configured to perform an inverse transform on the transform block using one or more separable transforms after performing the inverse transform on the LFNST region of the transform block using the LFNST.

25. The apparatus according to claim 19, wherein The unit configured to determine the position of the last significant coefficient in the transform block of video data includes: a unit configured to decode one or more syntax elements indicating an X position and a Y position of the last significant coefficient in the transform block.

26. The apparatus according to claim 19, wherein, The unit configured to perform an inverse transform on the transform block includes a unit configured to perform an inverse transform on the transform block to generate a residual block, and the apparatus further includes: a unit configured to determine a prediction block for the residual block; and A unit for combining the prediction block and the residual block to generate a decoded block.

27. The apparatus according to claim 26, further comprising: A unit for displaying a picture including the decoded block.

28. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors configured to decode video data to perform the following operations: Determine the position of the last significant coefficient in a transform block of video data by decoding a syntax element indicating the position of the last significant coefficient along a predetermined scan order; Determine a value of a Low-Frequency Non-Separable Transform (LFNST) index for the transform block based on a relative relationship between the position of the last significant coefficient and a zeroing region of the transform block, wherein the zeroing region of the transform block includes both a first region of the transform block within the LFNST region and a second region of the transform block outside the LFNST region, and wherein, to determine the value of the LFNST index, the instructions further cause the one or more processors: in the case where the position of the last significant coefficient in the transform block is within the zeroing region of the transform block, infer the value of the LFNST index to be zero, wherein a value of zero for the LFNST index indicates that the LFNST is not applied to the transform block; and Perform an inverse transform on the transform block according to the value of the LFNST index.

29. The non-transitory computer-readable storage medium according to claim 28, wherein, The value of the LFNST index indicates whether the LFNST is applied to the transform block, and if so, indicates the type of LFNST applied.

30. The non-transitory computer-readable storage medium according to claim 28, wherein, To infer the value of the LFNST index to be zero, the instructions further cause the one or more processors: Infer the value of the LFNST index to be zero when not receiving a syntax element indicating the value of the LFNST index.

31. The non-transitory computer-readable storage medium according to claim 28, wherein, To perform an inverse transform on the transform block, the instructions further cause the one or more processors: Perform an inverse transform on the transform block using one or more separable transforms.

32. The non-transitory computer-readable storage medium according to claim 28, wherein, To determine the value of the LFNST index, the instructions further cause the one or more processors: Receive a syntax element indicating the LFNST index in the case where the position of the last significant coefficient in the transform block is not within the zeroing region of the transform block; and Decode the syntax element to determine the value of the LFNST index.

33. The non-transitory computer-readable storage medium according to claim 32, wherein, To perform an inverse transform on the transform block, the instructions further cause the one or more processors: Perform an inverse transform on the LFNST region of the transform block using the LFNST indicated by the LFNST index; and After performing the inverse transform on the LFNST region of the transform block using the LFNST, perform an inverse transform on the transform block using one or more separable transforms.

34. The non-transitory computer-readable storage medium according to claim 28, wherein, To determine the position of the last significant coefficient in the transform block of video data, the instructions further cause the one or more processors: Decode one or more syntax elements indicating the X and Y positions of the last significant coefficient in the transform block.

35. The non-transitory computer-readable storage medium according to claim 28, wherein, To perform an inverse transform on the transform block, the instruction further causes the one or more processors to perform an inverse transform on the transform block to generate a residual block, and wherein the instruction further causes the one or more processors to: Determine a prediction block for the residual block; and Combine the prediction block with the residual block to generate a decoded block.

36. The non-transitory computer-readable storage medium according to claim 35, wherein the instruction further causes the one or more processors to: Display a picture including the decoded block.

Citation Information

Patent Citations

  • Enhanced multiple transforms for prediction residual

    US10306229B2

  • Efficient parameter storage for compact multi-pass transforms

    US10349085B2

  • Multi-pass non-separable transforms for video coding

    US10448053B2

  • Non-separable secondary transform for video coding

    US10491922B2

  • Non-separable secondary transform for video coding with reorganizing

    US20170094314A1