Transforming video data using indivisible primary transform
By directly processing residual blocks of video data using indivisible main transform (NSPT) in video decoding technology, the problem of inefficient processing of large block shapes in the prior art is solved, and more efficient video data compression and decoding are achieved.
Patent Information
- Application Number
- CN202380070973.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-12
- Filing Date
- 2023-10-13
- Publication Date
- 2025-05-16
AI Technical Summary
The existing video decoding technology has problems of inefficiency when processing video data, especially when processing large block shapes, traditional segmentable transformations and low-frequency indivisible transformations are difficult to improve the decoding efficiency.
The indivisible main transformation (NSPT) is used as the transformation method of video data, and it is directly applied to the residual data, avoiding the segmentable transformation step before NSPT, thereby improving the efficiency of large-block shape processing.
Through NSPT transformation, the compression efficiency and decoding efficiency of video data are significantly improved. Especially when processing larger block shapes, the NSPT transformation core can be used more effectively to improve system performance.
Smart Images

Figure CN120019648A_ABST
Abstract
Description
Related Applications This application claims the benefit of U.S. Patent Application No. 18 / 485,707, filed on October 12, 2023, U.S. Provisional Application No. 63 / 379,422, filed on October 13, 2022, and U.S. Provisional Application No. 63 / 385,678, filed on December 1, 2022, each of which is incorporated herein by reference in its entirety. U.S. Patent Application No. 18 / 485,707, filed on October 12, 2023, claims the benefit of U.S. Provisional Application No. 63 / 379,422, filed on October 13, 2022, and U.S. Provisional Application No. 63 / 385,678, filed on December 1, 2022. Technical Field The present disclosure relates to video coding, including video encoding and video decoding. Background Art Digital video capabilities may be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio telephones (so-called "smart phones"), video teleconferencing devices, video streaming devices, etc. Digital video devices implement video coding techniques (such as those described in the standards defined by MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 (Part 10, Advanced Video Coding (AVC)), ITU-TH.265 / High Efficiency Video Coding (HEVC), ITU-TH.266 / Versatile Video Coding (VVC) and extensions of such standards, as well as proprietary video codecs / formats (such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media)). By implementing such video coding techniques, video devices may more efficiently send, receive, encode, decode and / or store digital video information. Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction relative to reference samples in neighboring blocks in the same picture or temporal prediction relative to reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames. Summary of the invention In general, this disclosure describes techniques related to transforming video data during video coding. That is, transforms are used to convert representations of residual video data (representing the difference between an uncoded block and a prediction block) between a spatial or pixel domain and a frequency domain. For example, during video encoding, a residual block may be transformed to a transform block in the frequency domain, while during video decoding, a transform block may be transformed from the frequency domain to a reconstructed residual block in the spatial or pixel domain. This disclosure describes techniques related to using an inseparable primary transform when transforming or inverse transforming video data. In one example, a method for decoding video data includes: inversely transforming a transform coefficient block of a video data block using an inverse non-separable primary transform (NSPT) to reconstruct a residual block of the video data block without using an inverse separable transform; and decoding the video data block using the residual block. In another example, a device for decoding video data includes: a memory configured to store video data; and a processing system including one or more processors implemented in a circuit system, the processing system being configured to: inversely transform a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and decode the block using the residual block. In another example, an apparatus for decoding video data includes: a unit for inversely transforming a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and a unit for decoding the video data block using the residual block. The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a block diagram illustrating an example video encoding and decoding system that may perform the techniques of this disclosure.
[0011] Figure 2 is a block diagram illustrating an example video encoder that may perform the techniques of this disclosure.
[0012] Figure 3 is a block diagram illustrating an example video decoder that may perform the techniques of this disclosure.
[0013] Figure 4A and 4B is a conceptual diagram illustrating the use of a separable and low-frequency non-separable transform (LFNST) when encoding and decoding video data.
[0014] Figure 5 is a conceptual diagram illustrating a reverse non-dividable primary transform (NSPT) process.
[0015] Figures 6A to 6C is a conceptual diagram illustrating an example scanning pattern for reorganizing inverse quantized coefficients before applying NSPT.
[0016] Figure 7 is a conceptual diagram showing a reordered one-dimensional array from a two-dimensional matrix.
[0017] Figure 8 is a conceptual diagram illustrating deinterleaving a set of decoded coefficients into four one-dimensional arrays.
[0018] Fig. 9 is a flowchart illustrating an example method for encoding a current block according to techniques of this disclosure.
[0019] Fig.10 is a flow chart illustrating an example method for decoding a current block according to techniques of this disclosure.
[0020] Fig.11 is a flow chart illustrating an example method of encoding video data using a non-splittable primary transform (NSPT) in accordance with the techniques of this disclosure.
[0021] Fig.12 is a flow chart illustrating an example method of decoding video data using a non-splittable primary transform (NSPT) in accordance with the techniques of this disclosure. DETAILED DESCRIPTION
[0022] Video encoding typically includes forming a prediction block for a current block of video data, then calculating the difference between the current block and the prediction block to form a residual block. The video encoder may then apply one or more transforms to the residual block to produce a transform block in the transform domain. A "transform block" may be a block of transform coefficients. A video decoder may perform the inverse process, where the video decoder may decode the transform block, inverse transform the transform block to reproduce the residual block, and then add the residual block to the prediction block to reproduce the original uncoded block of video data.
[0023] In video coding standards prior to ITU-T H.265 / High Efficiency Video Coding (HEVC), only fixed separable transforms were used, such as DCT-2, which can be used vertically and horizontally. In addition to DCT-2 for 4x4 blocks as a fixed separable transform, HEVC adds DST-7 as a possible transform. Separable transforms usually refer to performing a transform in multiple steps. For example, in the first step, a vertical transform is applied, and a horizontal transform is applied to the result, or vice versa.
[0024] In addition to the separable transform, a low frequency non-separable transform (LFNST) can be applied to further improve coding efficiency. A non-separable transform generally refers to applying the transform once (e.g., instead of splitting into horizontal and vertical transforms, the values are transformed in one step). The LFNST is a non-separable secondary transform (NSST) that is applied to the main transform coefficients resulting from applying the separable transform as the main transform.
[0025] In some examples, three LFNST classes (LFNST4, LFNST8, LFNST16) may be used depending on the shape of the video data block. Each of these classes may include, for example, 35 sets and 3 candidates. In one example, LFNST4 is applied to blocks of size 4xN / Nx4 (N≥4), LFNST8 is applied to blocks of size 8xN / Nx8 (N≥8), and LFNST16 is applied to blocks of size MxN (M, N≥16).
[0026] In some examples, the set is selected based on the intra mode of the current block. The intra mode can be mapped to one of 35 sets according to the following lookup table: const uint8_t g_nsptLut
[97] = 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 86 87 88 89 90 91 92 93 94 95 96+0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,33,32,31,30,29,28,27,26,2 5,24,23,22,21,20,19,18,17,16,15,14,13,12,11,10,9,8,7,6,5,4,3,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2};
[0027] In some examples, for intra modes > 34, the corresponding main transform coefficient blocks are transposed before applying LFNST.
[0028] The signs of the transform coefficients can be predicted as part of the sign prediction process. In some examples, the prediction region is adaptively selected to be a 32x32 region, and the number of predicted symbols is configurable for non-LFNST transform blocks. In the case of LFNST, the prediction region can be limited to the top left 4×4 region, and up to 4 coefficients can be predicted.
[0029] The present disclosure describes techniques including the use of a non-separable primary transform (NSPT) as a supplement to existing transform kernels. NSPT is an additional class of transforms that is applied directly to residual data and therefore does not require a prior separable transform. Using NSPT can improve the coding efficiency that can be achieved using a larger NSPT transform kernel (i.e., larger than a separable transform kernel), as opposed to using both separable and LFNST.
[0030] According to the technology of the present disclosure, the affected block shapes can be processed using NSPT. Each block shape can be processed by a dedicated NSPT core, which can vary in the number of sets and candidates. If the block shape indicates processing by NSPT, the splittable main transform plus LFNST is not applied. Therefore, the compression efficiency can be improved.
[0031] Figure 1 1 is a block diagram illustrating an example video encoding and decoding system 100 that can perform the techniques of the present disclosure. In general, the techniques of the present disclosure relate to decoding (encoding and / or decoding) video data. Generally, video data includes any data used to process video. Thus, video data can include original uncoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).
[0032] like Figure 1 As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may include any of a wide variety of devices, including desktop computers, notebook computers (i.e., laptop computers), mobile devices, tablet computers, set-top boxes, telephone handsets such as smart phones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and, therefore, may be referred to as wireless communication devices.
[0033] exist Figure 1 In the example of , source device 102 includes video source 104, memory 106, video encoder 200 and output interface 108. Destination device 116 includes input interface 122, video decoder 300, memory 120 and display device 118. According to the present disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply a technique for transforming video data using an indivisible main transform. Therefore, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may be docked with an external display device instead of including an integrated display device.
[0034] like Figure 1The system 100 shown is only an example. In general, any digital video encoding and / or decoding device can perform techniques for transforming video data using NSPT. The source device 102 and the destination device 116 are only examples of such decoding devices, wherein the source device 102 generates decoded video data for transmission to the destination device 116. The present disclosure refers to a "decoding" device as a device that performs decoding (e.g., encoding and / or decoding) of data. Therefore, the video encoder 200 and the video decoder 300 represent examples of decoding devices (specifically, video encoders and video decoders), respectively. In some examples, the source device 102 and the destination device 116 can operate in a substantially symmetrical manner, so that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Therefore, the system 100 can support one-way or two-way video transmission between the source device 102 and the destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0035] Typically, the video source 104 represents the source of video data (i.e., the original, undecoded video data), and provides a series of pictures (also referred to as "frames") of the order of the video data to the video encoder 200, which encodes the data for the pictures. The video source 104 of the source device 102 may include a video capture device, such as a camera, a video archive unit containing previously captured original video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 104 may generate computer graphics-based data as the source video, or a combination of real-time video, archived video, and computer-generated video. In each case, the video encoder 200 may encode captured, pre-captured, or computer-generated video data. The video encoder 200 may rearrange the pictures from the received order (sometimes referred to as "display order") to a decoding order for decoding. The video encoder 200 may generate a bitstream comprising encoded video data. Source device 102 may then output the encoded video data onto computer-readable medium 110 via output interface 108 to be received and / or retrieved by, for example, input interface 122 of destination device 116 .
[0036] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general purpose memories. In some examples, the memories 106, 120 can store original video data, for example, original video from the video source 104 and original decoded video data from the video decoder 300. In addition or alternatively, the memories 106, 120 can store software instructions that can be executed by, for example, the video encoder 200 and the video decoder 300, respectively. Although the memory 106 and the memory 120 are shown as separate from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 can also include internal memories for functionally similar or equivalent purposes. In addition, the memories 106, 120 can store, for example, encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, portions of the memories 106, 120 can be allocated as one or more video buffers, for example, to store original decoded and / or encoded video data.
[0037] The computer-readable medium 110 may represent any type of medium or device capable of delivering the encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium that enables the source device 102 to directly send the encoded video data to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 may demodulate the transmission signal including the encoded video data according to a communication standard such as a wireless communication protocol, and the input interface 122 may demodulate the received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, for example, a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a portion of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful for facilitating communication from the source device 102 to the destination device 116.
[0038] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0039] In some examples, source device 102 may output the encoded video data to file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading.
[0040] The file server 114 may be any type of server device capable of storing encoded video data and sending the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a server configured to provide a file transfer protocol service (such as the File Transfer Protocol (FTP) or the File Delivery over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or an enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Real-time Streaming (HLS), Real-time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.
[0041] Destination device 116 may access the encoded video data from file server 114 through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on file server 114. Input interface 122 may be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.
[0042] Output interface 108 and input interface 122 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components that operate according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to other wireless standards (such as IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee TM ), Bluetooth TM Standards, etc.) to transmit data (such as encoded video data). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include a SoC device for performing the functions assigned to video encoder 200 and / or output interface 108, and destination device 116 may include a SoC device for performing the functions assigned to video decoder 300 and / or input interface 122.
[0043] The techniques of the present disclosure may be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0044] The input interface 122 of the destination device 116 receives the encoded video bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information such as the following syntax elements defined by the video encoder 200 (which are also used by the video decoder 300): the syntax elements have values that describe the characteristics and / or processing of video blocks or other decoding units (e.g., slices, pictures, groups of pictures, sequences, etc.). The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0045] Despite Figure 1Not shown, but in some examples, the video encoder 200 and the video decoder 300 can each be integrated with an audio encoder and / or an audio decoder, and can include appropriate MUX-DEMUX units or other hardware and / or software to process a multiplexed stream including both audio and video in a common data stream.
[0046] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium, and use one or more processors to execute the instructions in hardware to perform the technology of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, and any of the encoders or decoders can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. The device including the video encoder 200 and / or the video decoder 300 can include an integrated circuit, a microprocessor, and / or a wireless communication device (such as a cellular phone).
[0047] The video encoder 200 and the video decoder 300 may operate according to a video coding standard, such as the ITU-T H.265 (also known as the High Efficiency Video Coding (HEVC) standard) or an extension thereof, such as a multi-view and / or scalable video coding extension. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards, such as the ITU-T H.266 standard, also known as Versatile Video Coding (VVC). In other examples, the video encoder 200 and the video decoder 300 may operate according to a proprietary video codec / format, such as AOMedia Video 1 (AV1), an extension of AV1, and / or a subsequent version of AV1 (e.g., AV2). In other examples, the video encoder 200 and the video decoder 300 may operate according to other proprietary formats or industry standards. However, the techniques of the present disclosure are not limited to any particular coding standard or format. In general, the video encoder 200 and the video decoder 300 may be configured to perform the techniques of the present disclosure in conjunction with any video coding technique that uses NSPT to transform video data.
[0048] Typically, the video encoder 200 and the video decoder 300 can perform block-based decoding of the picture. The term "block" generally refers to a structure that includes data to be processed (e.g., to be encoded, decoded, or otherwise used in the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, the video encoder 200 and the video decoder 300 can decode video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, instead of decoding red, green, and blue (RGB) data for samples of a picture, the video encoder 200 and the video decoder 300 can decode luminance and chrominance components, wherein the chrominance components may include both red hue and blue hue chrominance components. In some examples, the video encoder 200 converts the received RGB formatted data to a YUV representation before encoding, and the video decoder 300 converts the YUV representation to an RGB format. Alternatively, a pre-processing and post-processing unit (not shown) may perform these conversions.
[0049] In general, the present disclosure may relate to the decoding (e.g., encoding and decoding) of a picture to include the process of encoding or decoding data for the picture. Similarly, the present disclosure may relate to the decoding of a block of a picture to include the process of encoding or decoding data for the block (e.g., prediction and / or residual decoding). The encoded video bitstream typically includes a series of values for syntax elements that represent decoding decisions (e.g., decoding modes) and partitioning of the picture into blocks. Therefore, references to decoding a picture or a block should generally be understood as decoding the values of the syntax elements used to form the picture or the block.
[0050] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as the video encoder 200) partitions a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. A node without child nodes may be referred to as a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video decoder may further partition the PU and TU. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of a TU. In HEVC, a PU represents inter-frame prediction data, and a TU represents residual data. An intra-predicted CU includes intra-frame prediction information, such as an intra-frame mode indication.
[0051] As another example, the video encoder 200 and the video decoder 300 may be configured to operate according to VVC. According to VVC, a video decoder (such as the video encoder 200) partitions a picture into a plurality of coding tree units (CTUs). The video encoder 200 may partition the CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CU, PU, and TU of HEVC. The QTBT structure includes two levels: a first level partitioned according to quadtree partitioning, and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).
[0052] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also known as ternary tree (TT)) partitioning. A ternary tree or ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, the ternary tree or ternary tree partitioning divides the block into three sub-blocks without dividing the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.
[0053] When operating according to the AV1 codec, the video encoder 200 and the video decoder 300 can be configured to decode the video data in blocks. In AV1, the largest decoding block that can be processed is called a super block. In AV1, a super block can be 128x128 luminance samples or 64x64 luminance samples. However, in subsequent video decoding formats (e.g., AV2), super blocks can be defined by different (e.g., larger) luminance sample sizes. In some examples, the super block is the top layer of the block quadtree. The video encoder 200 can further divide the super block into smaller decoding blocks. The video encoder 200 can use square or non-square partitioning to divide the super block and other decoding blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN and NxN / 4 blocks. The video encoder 200 and the video decoder 300 can perform separate prediction and transformation processes for each decoding block.
[0054] AV1 also defines tiles of video data. A tile is a rectangular array of super blocks that can be decoded independently of other tiles. That is, the video encoder 200 and the video decoder 300 can encode and decode the decoding blocks within the tile respectively without using video data from other tiles. However, the video encoder 200 and the video decoder 300 can perform filtering across tile boundaries. The size of the tile can be uniform or uneven. Tile-based decoding can achieve parallel processing and / or multithreading for encoders and decoders.
[0055] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma component and the chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).
[0056] The video encoder 200 and the video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, super block partitioning, or other partitioning structures.
[0057] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples for a picture with three sample arrays, or a CTB of samples for a monochrome picture or a picture coded using three separate color planes and syntax structures for coding the samples. A CTB can be an NxN block of samples (for some value of N) such that the division of components into CTBs is a partitioning. A component can be an array or a single sample from one of the three arrays (one luma and two chroma) of a picture in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array of a picture in a monochrome format. In some examples, a coding block is an MxN block of samples (for some values of M and N) such that the division of a CTB into coding blocks is a partitioning.
[0058] Blocks (e.g., CTUs or CUs) may be grouped in a picture in various ways. As an example, a brick may refer to a rectangular area of a CTU row within a particular tile in a picture. A tile may be a rectangular area of a CTU within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular area of a CTU having a height equal to the height of the picture and a width specified by a syntax element (e.g., such as in a picture parameter set). A tile row refers to a rectangular area of a CTU having a height specified by a syntax element (e.g., such as in a picture parameter set) and a width equal to the width of the picture.
[0059] In some examples, a tile may be partitioned into multiple bricks, each of which may include one or more CTU rows within the tile. Tiles that are not partitioned into multiple bricks may also be referred to as bricks. However, bricks that are true subsets of tiles may not be referred to as tiles. Bricks in a picture may also be arranged in slices. A slice may be an integer number of bricks of a picture that may be uniquely contained in a single network abstraction layer (NAL) unit. In some examples, a slice includes multiple complete tiles or a continuous sequence of complete bricks of only one tile.
[0060] This disclosure may use "NxN" and "N by N" interchangeably to refer to the sample size of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y=16) and 16 samples in the horizontal direction (x=16). Similarly, an NxNCU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. The samples in a CU may be arranged in rows and columns. In addition, a CU does not necessarily need to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may include NxM samples, where M is not necessarily equal to N.
[0061] The video encoder 200 encodes video data representing prediction and / or residual information and other information for a CU. The prediction information indicates how the CU will be predicted in order to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between the samples of the CU before encoding and the prediction block.
[0062] In order to predict a CU, the video encoder 200 can generally form a prediction block for the CU by inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting a CU based on data of a previously decoded picture, while intra-frame prediction generally refers to predicting a CU based on previously decoded data of the same picture. In order to perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate a prediction block. The video encoder 200 can generally perform motion search to identify a reference block that closely matches the CU, for example, in terms of the difference between the CU and the reference block. The video encoder 200 can use the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean square difference (MSD), or other such difference calculations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional prediction or bidirectional prediction to predict the current CU.
[0063] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).
[0064] To perform intra prediction, the video encoder 200 can select an intra prediction mode to generate a prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as planar modes and DC modes. Typically, the video encoder 200 selects an intra prediction mode that describes neighboring samples of the current block according to which the samples of the current block (e.g., a block of a CU) are to be predicted. Assuming that the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be above, above left, or to the left of the current block in the same picture as the current block.
[0065] The video encoder 200 encodes data representing a prediction mode for the current block. For example, for an inter-frame prediction mode, the video encoder 200 may encode data representing which of the various available inter-frame prediction modes is used and motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. The video encoder 200 may use a similar mode to encode motion vectors for affine motion compensation mode.
[0066] AV1 includes two general techniques for encoding and decoding coded blocks of video data. The two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when an intra prediction mode is used to predict a block of a current frame of video data, the video encoder 200 and the video decoder 300 do not use video data from other frames of the video data. For most intra prediction modes, the video encoder 200 encodes a block of the current frame based on the difference between the sample values in the current block and the prediction values generated from the reference samples in the same frame. The video encoder 200 determines the prediction values generated from the reference samples based on the intra prediction mode.
[0067] After a prediction such as intra prediction or inter prediction of a block, the video encoder 200 may calculate residual data for the block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and a prediction block for the block, which is formed using a corresponding prediction mode. The video encoder 200 may apply one or more transforms to the residual block to produce transformed data in a transform domain rather than in a sample domain. For example, the video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. In addition, the video encoder 200 may apply a secondary transform after the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.
[0068] According to the technology of the present disclosure, the video encoder 200 may determine to apply a non-separable primary transform (NSPT) to a residual block without applying a separable transform to the residual block before applying the NSPT. Because the NSPT is applied without a separable transform, the video encoder 200 may apply the NSPT to the entire residual block. The video encoder 200 may determine to apply the NSPT according to the prediction mode used to form the corresponding prediction block and / or according to the size of the residual block. For example, the video encoder 200 may determine that the residual block has a size of one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8. The video encoder 200 may include different NSPTs for residual blocks of different sizes and / or for different prediction modes (e.g., various intra-frame prediction modes). Therefore, a combination of a specific intra-frame prediction mode and the size of the residual block may be mapped to a specific NSPT. For blocks having sizes different than the size mapped to the NSPT and / or other prediction modes (e.g., inter-prediction or affine), the video encoder 200 may apply a conventional separable transform, which may then be followed by a low-frequency inseparable transform, which the video encoder 200 may apply to only a portion of the resulting transform block.
[0069] As described above, after any transform to produce transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all transform coefficients. For example, the video encoder 200 may round down an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift on the value to be quantized.
[0070] After quantization, the video encoder 200 may scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place the transform coefficients of higher energy (and therefore lower frequency) in front of the vector and the transform coefficients of lower energy (and therefore higher frequency) in the back of the vector. In some examples, the video encoder 200 may scan the quantized transform coefficients using a predefined scan order to generate a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 may perform adaptive scanning. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 may entropy encode the one-dimensional vector, for example, according to context adaptive binary arithmetic coding (CABAC). The video encoder 200 may also entropy encode the values of the syntax elements used to describe the metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.
[0071] To perform CABAC, the video encoder 200 may assign context within a context model to a symbol to be transmitted. The context may relate to, for example, whether neighboring values of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.
[0072] The video encoder 200 may also generate syntax data (such as block-based syntax data, picture-based syntax data, and sequence-based syntax data) for the video decoder 300, for example, in a picture header, a block header, a slice header, or other syntax data (such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS). Likewise, the video decoder 300 may decode such syntax data to determine how to decode the corresponding video data.
[0073] In this way, the video encoder 200 can generate a bitstream, which includes the encoded video data, for example, a syntax element describing the partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0074] In general, the video decoder 300 performs a process that is the reverse of the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may use CABAC to decode the values of the syntax elements for the bitstream in a manner substantially similar to, but reverse to, the CABAC encoding process of the video encoder 200. The syntax elements may define partitioning information for partitioning a picture into CTUs, and partitioning each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define CUs of the CTUs. The syntax elements may also define prediction and residual information for a block (e.g., CU) of video data.
[0075] The residual information may be represented by, for example, quantized transform coefficients. The video decoder 300 may inverse quantize and inverse transform the quantized transform coefficients of the block to reproduce a residual block for the block. The video decoder 300 uses the signaled prediction mode (intra-frame prediction or inter-frame prediction) and related prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for the block. The video decoder 300 may then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The video decoder 300 may perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the blocks.
[0076] According to the techniques of the present disclosure, the video decoder 300 may entropy decode a set of quantized transform coefficients for a current block together with data indicating the size of the block and the prediction mode of the block. The video decoder 300 may inverse quantize the quantized transform coefficients and inverse scan the resulting transform candidates to reconstruct the transform block. If the size and / or prediction mode is mapped to a non-dividable primary transform (NSPT), the video decoder 300 may apply an inverse NSPT corresponding to the NSPT to which the size and / or prediction mode is mapped to the transform block to reconstruct a residual block for the block.
[0077] Alternatively, if the size and / or prediction mode or other data indicating potential use of the NSPT does not map to the NSPT, the video decoder 300 may determine an inverse separable transform to be applied to the transform block. In some examples, the video decoder 300 may also determine a low frequency non-separable transform (LFNST) to be applied to the transform block. The transform block may also be referred to as a transform coefficient block.
[0078] In general, the present disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" may generally refer to the transmission of values for syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 may signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As described above, source device 102 may transmit the bitstream to destination device 116 in substantially real time or not in real time (such as may occur when storing syntax elements to storage device 112 for later retrieval by destination device 116).
[0079] Figure 2 is a block diagram illustrating an example video encoder 200 that may perform the techniques of this disclosure. Figure 2 It is provided for the purpose of explanation and should not be considered to limit the techniques generally exemplified and described in this disclosure. For the purpose of explanation, this disclosure describes a video encoder 200 according to VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure can be performed by video encoding devices configured for other video coding standards and video coding formats, such as AV1 and subsequent versions of the AV1 video coding format.
[0080] exist Figure 2 In the example of , the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, an entropy coding unit 220, and transform data 232. Any one or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 can be implemented in one or more processors or in processing circuits. For example, the units of the video encoder 200 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. In addition, the video encoder 200 may include additional or alternative processors or processing circuits to perform these and other functions.
[0081] The video data memory 230 may store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive video data from, for example, the video source 104 ( Figure 1) receives the video data stored in the video data memory 230. DPB 218 can act as a reference picture memory, which stores reference video data for use when subsequent video data is predicted by the video encoder 200. Video data memory 230 and DPB 218 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or a separate memory device. In various examples, video data memory 230 can be on chip with other components of the video encoder 200 (as shown), or off chip relative to those components.
[0082] In the present disclosure, references to the video data memory 230 should not be interpreted as limited to memory internal to the video encoder 200 (unless specifically described as such), or to memory external to the video encoder 200 (unless specifically described as such). Rather, references to the video data memory 230 should be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data for a current block to be encoded). Figure 1 The memory 106 may also provide temporary storage for outputs from the various units of the video encoder 200 .
[0083] Shows Figure 2 The various units of the video encoder 200 are provided to help understand the operations performed by the video encoder 200. These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are pre-set with respect to the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions with the operations that can be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits may execute software instructions (e.g., to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0084] The video encoder 200 may include an arithmetic logic unit (ALU), an elementary function unit (EFU), a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example in which the operation of the video encoder 200 is performed using software executed by a programmable circuit, the memory 106 ( Figure 1 ) may store instructions (eg, object code) for software that the video encoder 200 receives and executes, or another memory (not shown) within the video encoder 200 may store such instructions.
[0085] The video data memory 230 is configured to store the received video data. The video encoder 200 may retrieve the picture of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be the original video data to be encoded.
[0086] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.
[0087] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include partitioning of a CTU into CUs, a prediction mode for a CU, a transform type for residual data of a CU, a quantization parameter for residual data of a CU, etc. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a better rate-distortion value than other tested combinations.
[0088] The video encoder 200 may partition the picture retrieved from the video data memory 230 into a series of CTUs, and encapsulate one or more CTUs in a slice. The mode selection unit 202 may partition the CTUs of the picture according to a tree structure (such as the above-mentioned MTT structure, QTBT structure, super block structure, or quadtree structure). As described above, the video encoder 200 may form one or more CUs by partitioning the CTUs according to the tree structure. Such CUs may also be generally referred to as "video blocks" or "blocks".
[0089] Typically, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for the current block (e.g., the current CU, or the overlapping portion of the PU and TU in HEVC). In order to inter-predict the current block, the motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 218). Specifically, the motion estimation unit 222 may calculate a value representing the degree of similarity of the potential reference block to the current block, for example, based on the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 may typically perform these calculations using the sample-by-sample difference between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block with the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.
[0090] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, for unidirectional inter prediction, the motion estimation unit 222 may provide a single motion vector, while for bidirectional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then use the motion vectors to generate a prediction block. For example, the motion compensation unit 224 may use the motion vectors to retrieve data for the reference block. As another example, if the motion vectors have fractional sample precision, the motion compensation unit 224 may interpolate the values for the prediction block according to one or more interpolation filters. In addition, for bidirectional inter prediction, the motion compensation unit 224 may retrieve data for two reference blocks identified by the corresponding motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.
[0091] When operating according to the AV1 video coding format, the motion estimation unit 222 and the motion compensation unit 224 may be configured to encode coding blocks of video data (e.g., both luma coding blocks and chroma coding blocks) using translational motion compensation, affine motion compensation, overlapped block motion compensation (OBMC), and / or composite intra prediction.
[0092] As another example, for intra prediction or intra prediction decoding, the intra prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, for directional mode, the intra prediction unit 226 can generally mathematically combine the values of adjacent samples and fill these calculated values across the current block in a defined direction to produce a prediction block. As another example, for DC mode, the intra prediction unit 226 can calculate the average of adjacent samples of the current block and generate a prediction block to include the resulting average for each sample of the prediction block.
[0093] When operating according to the AV1 video coding format, the intra prediction unit 226 may be configured to encode a coding block of video data (e.g., both a luma coding block and a chroma coding block) using directional intra prediction, non-directional intra prediction, recursive filter intra prediction, chroma prediction according to luma (CFL) prediction, intra block copy (IBC), and / or palette mode. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes.
[0094] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the original uncoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines a residual block for the current block. In some examples, the residual generation unit 204 may calculate the difference between the sample values in the residual block to generate the residual block using residual differential pulse coding modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.
[0095] In an example where the mode selection unit 202 partitions the CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As noted above, the size of the CU may refer to the size of the luma coding block of the CU, and the size of the PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 may support PU sizes of 2Nx2N or NxN for intra-frame prediction, and 2Nx2N, 2NxN, Nx2N, NxN or similar symmetrical PU sizes for inter-frame prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.
[0096] In an example where mode selection unit 202 does not further partition a CU into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As described above, the size of a CU may refer to the size of the luma coding block of the CU. Video encoder 200 and video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0097] For other video coding techniques (such as intra-block copy mode coding, affine mode coding, and linear model (LM) mode coding, to name a few examples), the mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the coding technique. In some examples (such as palette mode coding), the mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how to reconstruct the block based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.
[0098] As described above, the residual generation unit 204 receives video data for a current block and a corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0099] The transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 can apply various transforms to the residual block to form a transform coefficient block. For example, the transform processing unit 206 can apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transform processing unit 206 can perform a variety of transforms on the residual block, for example, a primary transform and a secondary transform (such as a rotation transform). In some examples, the transform processing unit 206 does not apply a transform to the residual block.
[0100] According to the techniques of this disclosure, transform processing unit 206 may determine a transform to apply to the residual block using transform data 232. Transform data 232 represents one or more computer-readable media devices, such as memory, storing data for various transforms, such as a separable transform, a low-frequency non-separable transform (LFNST), and a non-separable primary transform (NSPT). The NSPT may be a non-separable transform that is applied directly to the residual block without intervening a separable transform. That is, transform processing unit 206 may apply the NSPT to the residual block without applying any separable transform to the block.
[0101] In addition, the transform data 232 may store data that maps various characteristics of the block to various transforms. For example, the transform data 232 may store data that maps block size and / or prediction mode data to corresponding transforms. In the case of NSPT, the data may map blocks having a size of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8 to the NSPT. The data may also map blocks corresponding to prediction blocks formed using various intra-frame prediction modes to the NSPT. Therefore, in some examples, the transform processing unit 206 may apply NSPT to residual blocks having a size of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8 and formed using prediction blocks generated according to the intra-frame prediction mode. For other modes and block sizes, the transform processing unit 206 may select a separable transform based on the mapping data, and in some examples, LFNST may be applied after the separable transform.
[0102] When operating in accordance with AV1, the transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). The transform processing unit 206 may apply various transforms to the residual block to form a transform coefficient block. For example, the transform processing unit 206 may apply a horizontal / vertical transform combination, which may include a discrete cosine transform (DCT), an asymmetric discrete sine transform (ADST), a flipped ADST (e.g., inverse ADST), and an identity transform (IDTX). When an identity transform is used, a transform is skipped in one of the vertical or horizontal directions. In some examples, the transform process may be skipped.
[0103] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. The quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause a loss of information, and therefore, the quantized transform coefficients may have a lower precision than the original transform coefficients produced by the transform processing unit 206.
[0104] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. As noted above, in some examples, the transform processing unit 206 may have applied NSPT to the residual block to produce the transform block. In such cases, for example, based on the size and prediction mode of the block corresponding to the transform block, the inverse transform processing unit 212 may determine the inverse NSPT to be applied to the transform block to reconstruct the corresponding residual block. That is, the inverse transform processing unit 212 may apply the inverse NSPT to directly reconstruct the residual block without also applying a separable transform to reconstruct the residual block. However, in the case where the transform processing unit 206 applies a separable transform (potentially followed by LFNST) to the residual block, the inverse transform processing unit 212 may retrieve the corresponding inverse LFNST (if applicable) and the corresponding inverse separable transform from the transform data 232, and apply the inverse LFNST and the inverse separable transform to the transform block to reconstruct the residual block.
[0105] The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add samples of the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit 202 to generate a reconstructed block.
[0106] Filter unit 216 may perform one or more filter operations on the reconstructed blocks. For example, filter unit 216 may perform a deblocking operation to reduce blocking artifacts along CU edges. In some examples, the operations of filter unit 216 may be skipped.
[0107] When operating according to AV1, the filter unit 216 may perform one or more filter operations on the reconstructed blocks. For example, the filter unit 216 may perform a deblocking operation to reduce blocking artifacts along the edges of the CU. In other examples, the filter unit 216 may apply a constrained directional enhancement filter (CDEF), which may be applied after deblocking and may include applying an inseparable, non-linear, low-pass directional filter based on an estimated edge direction. The filter unit 216 may also include a loop recovery filter applied after CDEF, and may include a separable symmetric normalized Wiener filter or a dual self-guided filter.
[0108] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not performed, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is performed, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve the reference picture formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction on the blocks of the subsequently encoded pictures. In addition, the intra-frame prediction unit 226 can use the reconstructed blocks of the current picture in the DPB 218 to perform intra-frame prediction on other blocks in the current picture.
[0109] In general, the entropy coding unit 220 may entropy encode syntax elements received from other functional components of the video encoder 200. For example, the entropy coding unit 220 may entropy encode a quantized transform coefficient block from the quantization unit 208. As another example, the entropy coding unit 220 may entropy encode a prediction syntax element (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from the mode selection unit 202. The entropy coding unit 220 may perform one or more entropy coding operations on syntax elements as another example of video data to generate entropy-encoded data. For example, the entropy coding unit 220 may perform a context-adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable-to-variable (V2V) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an exponential Golomb coding operation, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 may operate in a bypass mode in which the syntax elements are not entropy encoded.
[0110] The video encoder 200 may output a bitstream including entropy-encoded syntax elements required for reconstructing a block of a slice or picture. Specifically, the entropy encoding unit 220 may output a bitstream.
[0111] According to AV1, the entropy coding unit 220 may be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 include an alphabet of N elements, and the context (e.g., a probability model) includes a set of N probabilities. The entropy coding unit 220 may store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). The entropy coding unit 220 may perform recursive scaling using an update factor based on the alphabet size to update the context.
[0112] The above operations are described with respect to blocks. Such descriptions should be understood as operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and chroma coding blocks are the luma components and chroma components of a CU. In some examples, the luma coding blocks and chroma coding blocks are the luma components and chroma components of a PU.
[0113] In some examples, operations performed with respect to luma coding blocks need not be repeated for chroma coding blocks. As an example, operations for identifying motion vectors (MVs) and reference pictures for luma coding blocks need not be repeated to identify MVs and reference pictures for chroma blocks. Specifically, the MVs for luma coding blocks may be scaled to determine MVs for chroma blocks, and the reference pictures may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.
[0114] In this manner, the video encoder 200 represents an example of an apparatus for decoding video data, the apparatus comprising: a memory configured to store the video data; and a processing system comprising one or more processors (e.g., an inverse transform processing unit 212 and a reconstruction unit 214) implemented in a circuit system, the processing system being configured to: inversely transform a transform coefficient block of a video data block using an inverse non-separable primary transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and decode the block using the residual block.
[0115] Figure 3 is a block diagram illustrating an example video decoder 300 that may perform the techniques of this disclosure. Figure 3 It is provided for the purpose of explanation and does not limit the techniques generally illustrated and described in this disclosure. For the purpose of explanation, this disclosure describes the video decoder 300 according to the techniques of VVC (ITU-TH.266, under development) and HEVC (ITU-TH.265). However, the techniques of this disclosure can be performed by video decoding devices configured for other video decoding standards.
[0116] exist Figure 3In the example of , the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, transform data 322, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 134 can be implemented in one or more processors or in processing circuits. For example, the units of the video decoder 300 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor, ASIC, or FPGA. In addition, the video decoder 300 may include additional or alternative processors or processing circuits to perform these and other functions.
[0117] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include an addition unit that performs predictions according to other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0118] When operating according to AV1, the compensation unit 316 may be configured to decode the coding blocks of the video data (e.g., both the luma coding blocks and the chroma coding blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite intra-frame inter-frame prediction, as described above. The intra-frame prediction unit 318 may be configured to decode the coding blocks of the video data (e.g., both the luma coding blocks and the chroma coding blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL prediction, intra-block copy (IBC), and / or palette mode, as described above.
[0119] CPB memory 320 may store video data, such as an encoded video bitstream, to be decoded by components of video decoder 300. For example, the video bitstream may be obtained from computer readable medium 110 ( Figure 1) obtains video data stored in CPB memory 320. CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from an encoded video bitstream. In addition, CPB memory 320 may store video data other than syntax elements of decoded pictures, such as temporary data representing outputs from various units of video decoder 300. DPB 314 typically stores decoded pictures, and video decoder 300 may output decoded pictures and / or use decoded pictures as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory 320 and DPB 314 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, CPB memory 320 can be on-chip with other components of video decoder 300 , or off-chip relative to those components.
[0120] Additionally or alternatively, in some examples, video decoder 300 may retrieve the video from memory 120 ( Figure 1 ) to retrieve the decoded video data. That is, memory 120 may utilize CPB memory 320 to store data as discussed above. Likewise, when some or all of the functions of video decoder 300 are implemented in software to be executed by processing circuitry of video decoder 300, memory 120 may store instructions to be executed by video decoder 300.
[0121] Shows Figure 3 The various units shown in FIG. 3 are used to help understand the operations performed by the video decoder 300. These units may be implemented as fixed function circuits, programmable circuits, or a combination thereof. Figure 2 , fixed-function circuits refer to circuits that provide specific functions and are pre-set with respect to the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions with respect to the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by fixed-function circuits are generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0122] The video decoder 300 may include an ALU, an EFU, a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuit, an on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0123] The entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate decoded video data based on the syntax elements extracted from the bitstream.
[0124] Generally, the video decoder 300 reconstructs a picture block by block. The video decoder 300 may perform a reconstruction operation on each block individually (wherein a block currently being reconstructed (ie, decoded) may be referred to as a "current block").
[0125] The entropy decoding unit 302 may entropy decode syntax elements defining the quantized transform coefficients of the quantized transform coefficient block and transform information such as a quantization parameter (QP) and / or a transform mode indication. The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine a degree of quantization and, likewise, determine a degree of inverse quantization for the inverse quantization unit 306 to apply. The inverse quantization unit 306 may, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.
[0126] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the transform coefficient block.
[0127] According to the techniques of this disclosure, transform data 322 includes coefficients for a variety of inverse transforms, such as an inverse separable transform, an inverse low-frequency non-separable transform (LFNST), and an inverse non-separable main transform (NSPT). Transform data 322 also includes data that maps characteristics of a block to the various inverse transforms, such as size and prediction mode information. In some examples, the data may map sizes 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8 and intra-prediction modes to an inverse NSPT, and other sizes and / or prediction modes to an inverse separable transform (and possibly an inverse LFNST). Thus, when the transform block corresponds to an intra-predicted block having a size of one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8, the inverse transform processing unit 308 may apply an inverse NSPT to the transform block (i.e., a transform coefficient block) without using an inverse partitionable transform; and the inverse transform processing unit 308 may apply an inverse partitionable transform to other transform blocks.
[0128] In addition, prediction processing unit 304 generates a prediction block based on the prediction information syntax element entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter-predicted, motion compensation unit 316 can generate a prediction block. In this case, the prediction information syntax element can indicate a reference picture in DPB 314 from which to retrieve the reference block, and a motion vector that identifies the location of the reference block in the reference picture relative to the location of the current block in the current picture. Motion compensation unit 316 can generally generate a prediction block in the same manner as described for motion compensation unit 224 ( Figure 2 )The inter-frame prediction process is performed in a manner substantially similar to that described in detail above.
[0129] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 may generally generate a prediction block in the same manner as described with respect to the intra-prediction unit 226 ( Figure 2 The intra prediction process is performed in a substantially similar manner as described in the foregoing description. The intra prediction unit 318 may retrieve data of neighboring samples of the current block from the DPB 314.
[0130] The reconstruction unit 310 may reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 may reconstruct the current block by adding samples of the residual block to corresponding samples of the prediction block.
[0131] The filter unit 312 may perform one or more filter operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 may not necessarily be performed in all examples.
[0132] The video decoder 300 may store the reconstructed block in the DPB 314. For example, in examples where the operation of the filter unit 312 is not performed, the reconstruction unit 310 may store the reconstructed block in the DPB 314. In examples where the operation of the filter unit 312 is performed, the filter unit 312 may store the filtered reconstructed block in the DPB 314. As discussed above, the DPB 314 may provide reference information (such as samples of the current picture for intra-frame prediction and previously decoded pictures for subsequent motion compensation) to the prediction processing unit 304. In addition, the video decoder 300 may output a decoded picture (e.g., a decoded video) from the DPB 314 for use in, for example, Figure 1 Subsequent presentation on a display device such as display device 118.
[0133] In this way, the video decoder 300 represents an example of a device for decoding video data, the device comprising: a memory configured to store the video data; and a processing system comprising one or more processors implemented in a circuit system, the processing system being configured to: inversely transform a transform coefficient block of a block of the video data using an inverse non-separable primary transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the block of the video data; and decode the block using the residual block.
[0134] Figure 4A and 4B is a conceptual diagram illustrating the use of splittable and low frequency non-splittable transform (LFNST) when encoding and decoding video data. Specifically, Figure 4A The application of a separable transform 130 is depicted, followed by a LFNST 132 and then quantization 134 . Figure 4B Depicted is inverse quantization 136, followed by inverse LFNST 138, and then inverse separable transform 140. In accordance with the techniques of this disclosure, video encoder 200 may apply separable transform 130 and LFNST 132 to blocks having sizes and / or prediction modes that do not map to NSPTs. Similarly, video decoder 300 may apply inverse LFNST 138 and inverse separable transform 140 to blocks having sizes and / or prediction modes that do not map to NSPTs.
[0135] Figure 5is a conceptual diagram illustrating an inverse non-separable primary transform (NSPT) process. According to the techniques of the present disclosure, after obtaining a list of two-dimensional (2-D) dequantized coefficients (based on a coefficient decoding step), video decoder 300 may apply an inverse non-separable primary transform (NSPT) to obtain a list of two-dimensional (2-D) dequantized coefficients according to the coefficient decoding step. Figure 5 Reconstruct the residual values in a 2-D block / array. That is, Figure 2 The inverse transform processing unit 212 of the video encoder 200 and Figure 3 The inverse transform processing unit 308 of the video decoder 300 may apply the inverse NSPT to reconstruct the residual block of video data.
[0136] like Figure 5 As shown in , a 2D block 150 of M transform coefficients may initially be reproduced, for example, by inverse quantizing the transform coefficients. The values of the 2D block 150 of M transform coefficients may be reorganized (152). Then, an inverse NSPT may be applied to the reorganized block of transform coefficients (154). This produces a 2D block 156 having N residual values (i.e., a reconstructed residual block).
[0137] The NSPT decoding process may be specified based on a reorganization pattern / scan that may be used to define how the NSPT input coefficients and output residual values are organized or grouped, and a transform matrix that defines an integral transform to be used on all or a subset of the dequantized coefficients.
[0138] In NSPT, the 2-D blocks / arrays of decoded coefficients need to be organized so as to align the entries and inputs of the matrix. This can be achieved by constructing a one-dimensional (1-D) list of coefficients M from the 2-D array of coefficients, and then applying the inverse NSPT (of size M×N) on the 1-D list to reconstruct the residual block.
[0139] Figures 6A to 6C is a conceptual diagram showing an example scanning pattern for reorganizing inverse quantized coefficients before applying NSPT. In particular, Fig. 6A depicts a sub-block diagonal scan 160, Figure 6B A horizontal scan 162 is depicted, and Figure 6C A vertical scan 164 is depicted.
[0140] Can be used Figures 6A-6C The input reorganization may be performed by any of the various scans of . The reorganization may also depend on the block size and / or the prediction mode (e.g., intra mode). If the codec normalizes a subset of the input coefficients to zero, then the input coefficient list for the NSPT may not contain those zeroed coefficients. The reorganization step may only accept coefficients that can be non-zero (i.e., coefficients that are not normalized to zero). The reorganization may require deinterleaving the coefficients into multiple 1D arrays, as described below with respect to Figure 8 shown.
[0141] Figure 7 is a conceptual diagram showing a reordered one-dimensional array 172 constructed from a two-dimensional matrix 170. In this example, Fig. 6A The sub-block diagonal scan 160 reorders the original two-dimensional matrix 170 into a one-dimensional array 172.
[0142] Figure 8 is a conceptual diagram illustrating deinterleaving a set of decoded coefficients 180 into four one-dimensional arrays 182A-182D.
[0143] The NSPT can be a matrix of size MxN, where M is an integer value representing the number of basis vectors and also represents the number of rows, and N is an integer value representing the number of NSPT coefficients reconstructed after applying the transform (also called the number of support samples for the transform).
[0144] The video encoder 200 and the video decoder 300 may reorganize a one-dimensional (1-D) list of N output NSPT coefficients based on an array (defining a pattern / scan), where each value in the array may correspond to a position / location in a 2-D block. The values in the array (for reorganization) may represent an index of a 2-D block in any predefined order. As an example, an index value may correspond to a position in a 2-D block. Given an index value v, the corresponding position in the 2-D block may be calculated as: row index (r) = floor(v / w) and column index c = mod(v, w), where mod(x, y) represents a modulo operation that returns the remainder of x divided by y, and where w represents the width of the NSPT sub-block.
[0145] Based on these formulas, the following reorganized array can correspond to the raster order of a 4x4 block: const int raster_order
[16] = { / / 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15}; Thus, the i-th element in the 1-D coefficient list is mapped to the row and column position in the 2D block. For example, for i=0, 1, 2, ... 15: r[i]=floor(raster_order[i] / w) c[i]=mod(raster_order[i],w)
[0146] The NSPT matrices may be grouped into sets S0, S1, S2, ... Each set may include one or several candidates (matrices) C0, C1, C2, ... dedicated to a certain block size WxH. In various examples, the number of S and C may vary. In one example, each intra mode is associated with a certain set in S. Thus, assuming 97 intra modes are available, there may be 97 possible sets S: S0, S1, ... S 96 Therefore, the transform may be selected based on the candidate index, for example, C=4.
[0147] As another example, intra-modes (0, ... 96) may be mapped to corresponding indices (e.g., 0, ... 35) according to a lookup table, such as: const uint8_t g_nsptLut
[97] = 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 86 87 88 89 90 91 92 93 94 95 96+0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,33,32,31,30,29,28,27,26,2 5,24,23,22,21,20,19,18,17,16,15,14,13,12,11,10,9,8,7,6,5,4,3,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2}.
[0148] Therefore, in this example, there are 35 possible sets S: S0, S1, ... S 34 This mapping can be achieved by symmetry considerations or other means. In the case of symmetry, the reorganization of the residual blocks can be handled in transposed order. The NSPT matrix can be selected based on the candidate index, for example, C=3.
[0149] As another example, only certain intra modes may be associated with NSPT. For example, intra mode 0 (PLANAR), mode 1 (DC), mode 18 (horizontal), mode 50 (vertical), and mode 34 (diagonal). Thus, there may be five sets S: S0, S1, S2, S3, S4. Only one candidate may be selected, so that no candidate index is needed, and therefore, C=1.
[0150] S and C may vary depending on the block size WxH. For certain block sizes, the NSPT matrix may not exist, and thus S=0, C=0.
[0151] In some examples, NSPT matrices exist for block sizes 4x4, 4x8, 8x4, 8x8, i.e., NSPT4x4, NSPT4x8, NSPT8x4, NSPT8x8. Each NSPT class may include a set of 35 sets S, depending on the intra mode and C=3 candidates.
[0152] Additionally, in some examples, NSPT matrices exist for block sizes 16x16 and higher. The following example addresses the possibility of utilizing NSPT to process larger blocks: For block size 16x16, a dedicated NSPT16x16 class is used, including 35 sets S and 3 candidates C. A block of size 16x16 can be divided into 4 subgroups of size 4x4. Each 4x4 subgroup is processed by a dedicated NSPT4x4 class that includes 35 sets S and 3 candidates C. After the forward transform (performed by the video encoder 200), the coefficients of each of the 4 subgroups are interleaved according to their energy. The video decoder 300 can recover the interleaving (e.g., according to Figure 8 ), resulting in 4 1D arrays of dequantized coefficients, each of which is inverse transformed using a dedicated NSPT 4x4 class, set and candidate. o Additionally, 32x32 blocks can be divided into 4 sub-groups of size 16x16 (processed by a dedicated NSPT16x16 transform) or 16 sub-groups of size 4x4 (processed by a dedicated NSPT4x4 transform). A block of size WxH can be divided into L subgroups of arbitrary size. For example, a 32x32 block can be divided into a 24x32 (processed by NSPT24x32) and an 8x32 subgroup (processed by NSPT8x32). The partitioning can depend on the intra mode and the block size. The coefficients of the 32x32 block can be processed using a dedicated NSPT16x16. After inverse transform and reordering, the 16x16 residual block is restored. The residual block can be upsampled to size 32x32.
[0153] In some examples, there are NSPT & LFNST matrices for block sizes 4x4, 4x8, 8x4, 8x8. Depending on the intra mode (35 sets S) that exploit symmetry, each block size can correspond to one of the three candidates (C0, C1, C2) in the set. The 3 candidates can be NSPT or LFNST kernels, for example, candidate 1 is NSPT, candidates 2 and 3 are LFNST, and so on.
[0154] For non-square blocks, the following special case can be applied in order to exploit the symmetry: If the residual block is transposed, the dimensions are swapped (e.g., 4x8->8x4). Thus, the transposed 4x8 block can use the transformation matrix associated with the non-transposed 8x4 block, and vice versa.
[0155] For example, it may be necessary to use NSPT to process a 4x8 block containing dequantized coefficients. First, the coefficients can be reordered into a 1D array. Next, the intra mode of the block can be used to decide whether to transpose the encoder side residual before the forward transform. Conceptually, the vertical mode can be mapped to the horizontal mode. If the intra mode is in the range (34, ..., 67), the block can be transposed. Two situations may occur: if the intra mode is a smaller 34, the indicated set and candidates of NSPT4x8 can be used for the inverse transform reordered 1D array. The residual values can then be reordered into a 2D residual block without transposition. On the other hand, if the intra mode is in the range of (34, ..., 67), the indicated set and candidates of NSPT8x4 can be used for the inverse transform reordered 1D array. The residual values can be reordered into a 2D residual block and the block can be transposed.
[0156] The signs of the coefficients obtained by applying NSPT can be predicted by sign prediction. In one example, the sign prediction follows the LFNST design, that is, only the signs of up to 4 coefficients can be predicted, limited to the upper left 4×4 area. In another example, up to 4 coefficient symbols can be predicted in an area adaptively increased to 32x32.
[0157] For hardware implementation, the worst case number of multiplications (per transform coefficient) can be an important complexity criterion. A simple way to reduce the worst case number of multiplications is to normalize some of the transform coefficients to zero. However, the zeroing process may generally be undesirable unless a certain worst case number of multiplications is required to be met. For NSPT, the degree of zeroing may depend on the size of the block and the current complexity limit, which may change in the future. This defines the size of the NSPT matrix, MxN.
[0158] In some examples, for a 4x4 residual block, the NSPT4x4 matrix has size 16x16, and all 16 NSPT coefficients are retained.
[0159] In some examples, for 4x8 and 8x4 residual blocks, the NSPT4x8 and NSPT8x4 matrices have size 32x20. 20 of the 32 NSPT coefficients are retained. The remaining 12 coefficients are set equal to zero.
[0160] In some examples, for an 8x8 residual block, the NSPT8x8 matrix is 64x32. 32 of the 64 NSPT coefficients are retained. The remaining 32 are set to zero.
[0161] The video encoder 200 and the video decoder 300 may be configured to apply NSPT to block sizes larger than 8x8, such as blocks of size 4x16, 16x4, 8x16, 16x8, and 16x16. These techniques may also be extended to even larger sizes, such as 32xN and Nx32, but transform coefficient storage may become an issue. To address storage issues and computational complexity, the following techniques may be employed.
[0162] In some examples, the video encoder 200 and the video decoder 300 may be configured to apply NSPT to blocks of size 4x16 and / or 16x4, and the storage requirements for transforms for these block sizes may be equivalent to the 8x8 case. A zeroing mechanism as discussed above may be employed, where only 20, 24, or 32 of the resulting coefficients are stored, and the remainder of the coefficients may be normalized to zero. The video encoder 200 and the video decoder 300 may determine the number of coefficients to maintain based on a complexity and performance tradeoff that is maintained below a certain worst-case computational complexity (e.g., multiplication and addition operations).
[0163] In some examples, the video encoder 200 and the video decoder 300 may be configured to apply NSPT to blocks of size 8x16, 16x8, or 16x16. The above-described zeroing mechanism may be employed, where only 32 or 40 coefficients are retained for blocks of size 8x16 or 16x8, and 32, 40, or 44 resulting coefficients may be retained for blocks of size 16x16. The video encoder 200 and the video decoder 300 may set the remainder of the coefficients to zero.
[0164] The storage requirements for larger transforms can be reduced by using a smaller number of transform sets. This can be achieved by further clustering adjacent directional intra prediction modes to reduce the number of mapped intra modes to reduce the number of mapped intra modes from 35. Larger transform block sizes will have a smaller number of mapped intra modes. For example, a 16x16 block may have only 4 mapping modes, while an 8x16 or 16x8 block may have 11 mapping modes.
[0165] In some examples, for a particular block size, NSPT may be applied to certain intra modes and LFNST may be applied to other intra modes.
[0166] Alternatively, for a particular block size and intra prediction mode, certain signaled transform selection indices (eg, low frequency non-divisible transform index syntax element lfnst_idx) may correspond to whether NSPT or LFNST will be used. In some examples, other transform indices may correspond to LFNST.
[0167] In some examples, each mapped intra prediction mode uses three different inseparable kernels. One of the inseparable kernels may correspond to NSPT, while the other two inseparable kernels may correspond to LFNST kernels. In some examples, the video encoder 200 and the video decoder 300 may be configured to apply a combination of hybrid LFNST and NSPT kernels based on the intra prediction mode and the signaled index value.
[0168] Fig. 9 2 is a flowchart illustrating an example method for encoding a current block according to the techniques of the present disclosure. The current block may include a current CU. Although with respect to the video encoder 200 ( Figure 1 and 2 ), but it should be understood that other devices may be configured to perform the same Fig. 9 A similar approach to the one used in this paper.
[0169] In this example, the video encoder 200 initially predicts a current block (350). For example, the video encoder 200 may form a prediction block for the current block. The video encoder 200 may then calculate a residual block for the current block (352). To calculate the residual block, the video encoder 200 may calculate the difference between the original uncoded block and the prediction block for the current block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). For example, the video encoder 200 may transform the residual block using an inseparable primary transform according to any of the various techniques of the present disclosure, either alone or in any combination. Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may encode the transform coefficients using CAVLC or CABAC. The video encoder 200 may then output entropy encoded data for the block (360).
[0170] The video encoder 200 may also decode the current block after encoding the current block to use the decoded version of the current block as reference data for subsequently coded data (e.g., in an inter-frame or intra-frame prediction mode). Thus, the video encoder 200 may inverse quantize and inverse transform the coefficients to reproduce a residual block (362). The video encoder 200 may inverse transform the coefficients using the NSPT according to any of the various techniques of the present disclosure. The video encoder 200 may combine the residual block with the prediction block to form a decoded block (364). The video encoder 200 may then store the decoded block in the DPB 218 (366).
[0171] Fig.10 1 is a flowchart illustrating an example method for decoding a current block of video data according to the techniques of the present disclosure. The current block may include a current CU. Although with respect to the video decoder 300 ( Figure 1 and 3 ), but it should be understood that other devices may be configured to perform the same Fig.10 A similar approach to the one used in this paper.
[0172] The video decoder 300 may receive entropy encoded data for a current block (e.g., entropy encoded prediction information and entropy encoded data for transform coefficients of a residual block corresponding to the current block) (370). The video decoder 300 may entropy decode the entropy encoded data to determine the prediction information for the current block and reproduce the transform coefficients of the residual block (372). The video decoder 300 may predict the current block (374), for example, using an intra-frame or inter-frame prediction mode as indicated by the prediction information for the current block to calculate the prediction block for the current block. The video decoder 300 may then inverse scan the reproduced transform coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (378). The video decoder 300 may inverse transform the coefficients using NSPT according to any of the various techniques of the present disclosure. Finally, the video decoder 300 may decode the current block by combining the prediction block and the residual block ( 380 ).
[0173] Fig.11 is a flowchart illustrating an example method for encoding video data using a non-separable primary transform (NSPT) according to the techniques of the present disclosure. Figure 1 and Figure 2 Video Encoder 200 Explanation Fig.11 Other video encoding devices may be configured to perform this or similar methods. Fig.11 The method can usually correspond to Fig. 9 Step 354 of the method.
[0174] Initially, video encoder 200 determines a prediction mode for a block of video data (400). Video encoder 200 also determines the size of the block (402). Video encoder 200 may then determine whether the size and prediction mode map to a non-separable primary transform (NSPT) (404). For example, sizes 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8 and intra prediction modes may map to an NSPT, while other combinations of size and prediction mode may map to a separable transform (and, in some cases, a low-frequency non-separable transform).
[0175] Thus, if the size and prediction mode are mapped to NSPT ("yes" branch of 404), video encoder 200 may apply NSPT to the residual block to form a transform block (i.e., a transform coefficient block) (406). On the other hand, if the size and prediction mode are not mapped to NSPT ("no" branch of 404), video encoder 200 may apply a separable transform to the residual block (408), and in some cases, apply LFNST to at least a portion of the resulting transform block (410).
[0176] In either case, video encoder 200 may then quantize the resulting transform coefficients (412) and entropy encode the quantized transform coefficients, the prediction mode, and the size data (414).
[0177] In this way, Fig.11 The method represents an example of a method of encoding video data, including transforming a residual block of video data using a non-separable primary transform (NSPT) without using a separable transform to construct block transform coefficients; and encoding the transform coefficient block.
[0178] Fig.12 is a flowchart illustrating an example method for decoding video data using a non-separable primary transform (NSPT) according to the techniques of the present disclosure. Figure 1 and 3 The video decoder 300 explains Fig.12 Other video decoding devices may be configured to perform this or similar methods. For example, the decoding loop portion of the video encoder 200 may also perform Fig.12 method. Fig.12 The method can usually correspond to Fig. 9 Step 364 of the method or Fig.10 Step 378 of the method.
[0179] Initially, video decoder 300 entropy decodes quantized transform coefficients, size data, and a prediction mode for a current block of video data (420). Video decoder 300 may inverse quantize the quantized transform coefficients (422) to reconstruct a block of transform coefficients. Video decoder 300 may also determine a prediction mode for the block (424) and determine a size of the block using the entropy decoded data (426).
[0180] Video decoder 300 may then determine whether the size and prediction mode are mapped to an NSPT (428). If the size and prediction mode are mapped to an NSPT (the "yes" branch of 428), video decoder 300 may inverse transform the transform coefficient block (430) using the corresponding NSPT to reconstruct a residual block for the block, i.e., the inverse separable transform is not applied when reconstructing the residual block. On the other hand, if the size and prediction mode are not mapped to an NSPT (the "no" branch of 428), video decoder 300 may apply an inverse LFNST to the transform block (432) and reconstruct the residual block using the inverse separable transform (434).
[0181] In either case, video decoder 300 may form a prediction block using the indicated prediction mode ( 436 ) and combine the prediction block with the residual block to reconstruct the current block ( 438 ).
[0182] In this way, Fig.12 The method represents an example of a method for decoding video data, including inverse transforming a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and decoding the video data block using the residual block.
[0183] Various examples of the techniques of the present disclosure are summarized in the following clauses:
[0184] Clause 1: A method of decoding video data, the method comprising: inversely transforming a transform coefficient block of a video data block using an inverse non-splittable primary transform (NSPT) to reconstruct a residual block of the video data block; and decoding the video data block using the residual block.
[0185] Clause 2: The method of clause 1, wherein inverse transforming the block of transform coefficients comprises: reorganizing the block of transform coefficients to form a reorganized block of transform coefficients; and inverse transforming the reorganized block of transform coefficients.
[0186] Clause 3: A method according to any one of clauses 1 and 2, wherein inverse transforming the transform coefficient block comprises: constructing a one-dimensional list of coefficients from the transform coefficient block; and applying the inverse NSPT to the one-dimensional list of coefficients to reconstruct the residual block.
[0187] Clause 4: The method of clause 3, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a sub-block diagonal scan to the transform coefficient block.
[0188] Clause 5: The method of clause 3, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a horizontal scan to the block of transform coefficients.
[0189] Clause 6: The method of clause 3, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a vertical scan to the transform coefficient block.
[0190] Clause 7: A method according to any one of clauses 1 to 6, wherein the inverse NSPT is defined as a matrix of size MxN, where M is an integer value representing the number of basis vectors and also the number of rows in the matrix, and where N is an integer value representing the number of support samples of the inverse NSPT.
[0191] Clause 8: The method of clause 7, wherein the matrix comprises 8-bit precision values.
[0192] Clause 9: The method of any one of clauses 1 to 8, further comprising: selecting the inverse NSPT from a set of possible inverse NSPTs.
[0193] Clause 10: The method of clause 9, further comprising: selecting the possible inverse NSPT set from a plurality of possible inverse NSPT sets.
[0194] Clause 11: The method of clause 10, wherein selecting the set of possible inverse NSPTs comprises selecting the set of possible inverse NSPTs according to an intra prediction mode for the block of video data.
[0195] Clause 12: The method of any of clauses 9 to 11, wherein selecting the inverse NSPT comprises selecting the inverse NSPT based on a size of the block of video data.
[0196] Clause 13: The method of any of clauses 1-12, further comprising: performing sign prediction to predict one or more signs of one or more of the transform coefficients.
[0197] Clause 14: A method according to any of clauses 1 to 13, wherein decoding the block of video data comprises: forming a prediction block for the block of video data; and combining the prediction block with the residual block to form a decoded block of the block of video data.
[0198] Clause 15: A method according to any one of clauses 1-14, wherein the transform coefficient block is one of a 4x16 or 16x4 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 20, 24 or 32 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0199] Clause 16: A method according to clause 15, wherein when there are 20 non-zero-valued transform coefficients, the remaining transform coefficients are 44 zero-valued transform coefficients, when there are 24 non-zero-valued transform coefficients, the remaining transform coefficients are 40 zero-valued transform coefficients, or when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 32 zero-valued transform coefficients.
[0200] Clause 17: A method according to any one of clauses 1-14, wherein the transform coefficient block is one of an 8x16 block or a 16x8 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32 or 40 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0201] Clause 18: The method of clause 17, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 96 zero-valued transform coefficients, or when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 88 zero-valued transform coefficients.
[0202] Clause 19: A method according to any one of clauses 1-14, wherein the transform coefficient block is a 16x16 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32, 40 or 44 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0203] Clause 20: A method according to clause 19, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 224 zero-valued transform coefficients, when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 216 zero-valued transform coefficients, or when there are 44 non-zero-valued transform coefficients, the remaining transform coefficients are 212 zero-valued transform coefficients.
[0204] Clause 21: The method of clause 1, wherein inverse transforming the block of transform coefficients comprises: reorganizing the block of transform coefficients to form a reorganized block of transform coefficients; and inverse transforming the reorganized block of transform coefficients.
[0205] Clause 22: The method of clause 1, wherein inverse transforming the transform coefficient block comprises: constructing a one-dimensional list of coefficients from the transform coefficient block; and applying the inverse NSPT to the one-dimensional list of coefficients to reconstruct the residual block.
[0206] Clause 23: The method of clause 16, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a sub-block diagonal scan to the transform coefficient block.
[0207] Clause 24: The method of clause 16, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a horizontal scan to the block of transform coefficients.
[0208] Clause 25: The method of clause 16, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a vertical scan to the block of transform coefficients.
[0209] Clause 26: A method according to clause 1, wherein the inverse NSPT is defined as a matrix of size MxN, where M is an integer value representing the number of basis vectors and also the number of rows in the matrix, and where N is an integer value representing the number of support samples of the inverse NSPT.
[0210] Clause 27: The method of clause 20, wherein the matrix comprises 8-bit precision values.
[0211] Clause 28: The method of clause 1, further comprising: selecting the inverse NSPT from a set of possible inverse NSPTs.
[0212] Clause 29: The method according to clause 22, further comprising: selecting the possible inverse NSPT set from a plurality of possible inverse NSPT sets.
[0213] Clause 30: The method of clause 23, wherein selecting the set of possible inverse NSPTs comprises selecting the set of possible inverse NSPTs according to an intra prediction mode for the block of video data.
[0214] Clause 31: The method of clause 22, wherein selecting the set of possible inverse NSPTs or at least one of the inverse NSPTs comprises selecting the set of possible inverse NSPTs or at least one of the inverse NSPTs based on a size of the block of video data.
[0215] Clause 32: The method of clause 1, further comprising: performing sign prediction to predict one or more signs of one or more of the transform coefficients.
[0216] Clause 33: The method of clause 1, wherein decoding the block of video data comprises: forming a prediction block for the block of video data; and combining the prediction block with the residual block to form a decoded block of the block of video data.
[0217] Clause 34: A method according to clause 1, wherein the transform coefficient block is one of a 4x16 or 16x4 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 20, 24 or 32 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0218] Clause 35: A method according to clause 34, wherein, when there are 20 non-zero-valued transform coefficients, the remaining transform coefficients are 44 zero-valued transform coefficients, when there are 24 non-zero-valued transform coefficients, the remaining transform coefficients are 40 zero-valued transform coefficients, or when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 32 zero-valued transform coefficients.
[0219] Clause 36: A method according to clause 1, wherein the transform coefficient block is one of an 8x16 block or a 16x8 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32 or 40 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0220] Clause 37: A method according to clause 36, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 96 zero-valued transform coefficients, or when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 88 zero-valued transform coefficients.
[0221] Clause 38: A method according to clause 1, wherein the transform coefficient block is a 16x16 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32, 40 or 44 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0222] Clause 39: A method according to clause 38, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 224 zero-valued transform coefficients, when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 216 zero-valued transform coefficients, or when there are 44 non-zero-valued transform coefficients, the remaining transform coefficients are 212 zero-valued transform coefficients.
[0223] Clause 40: The method of clause 1, further comprising: encoding the block of video data prior to decoding the block of video data.
[0224] Clause 41: The method of any of clauses 1 to 39, further comprising: encoding the block of video data prior to decoding the block of video data.
[0225] Clause 42: An apparatus for decoding video data, the apparatus comprising one or more means for performing the method according to any of clauses 1 to 41.
[0226] Clause 43: The apparatus of clause 42, wherein the one or more units comprise one or more processors implemented in circuitry.
[0227] Clause 44: The apparatus of any of clauses 42 and 43, further comprising: a display configured to display the decoded video data.
[0228] Clause 45: The device of any of clauses 42-44, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0229] Clause 46: The apparatus of clauses 42-45, further comprising: a memory configured to store video data.
[0230] Clause 47: A computer-readable storage medium having stored thereon instructions which, when executed, cause a processor of a device for decoding video data to perform the method of any one of clauses 1 to 41.
[0231] Item 48: A device for decoding video data, the device comprising: a unit for inversely transforming a transform coefficient block of a video data block using an inverse non-splittable primary transform (NSPT) to reconstruct a residual block of the video data block; and a unit for decoding the video data block using the residual block.
[0232] Item 49: A method for decoding video data, the method comprising: inversely transforming a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and decoding the video data block using the residual block.
[0233] Clause 50: The method of clause 49, further comprising: forming a prediction block for the block of video data using an intra prediction mode; and determining an inverse NSPT based on the intra prediction mode.
[0234] Clause 51: The method according to Clause 49 further includes: determining the size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and selecting the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the block.
[0235] Clause 52: The method according to Clause 49 further includes: retrieving coefficients of the inverse NSPT from a memory storing coefficients for multiple inverse NSPTs, wherein the multiple inverse NSPTs include 4x4 inverse NSPT, 4x8 inverse NSPT, 8x4 inverse NSPT, 4x16 inverse NSPT, 16x4 inverse NSPT, 8x8 inverse NSPT, 8x16 inverse NSPT and 16x8 inverse NSPT.
[0236] Clause 53: A method according to Clause 49, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes the first video data block, and the residual block includes a first residual block, the method further comprising: determining that a second transform coefficient block of a second video data block has a second size different from the first size; based on the second size being different from the first size, inversely transforming the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and decoding the second video data block using the second residual block.
[0237] Clause 54: The method of clause 49, wherein inverse transforming the block of transform coefficients comprises: reorganizing the block of transform coefficients to form a reorganized block of transform coefficients; and inverse transforming the reorganized block of transform coefficients.
[0238] Clause 55: The method of clause 49, wherein inverse transforming the transform coefficient block comprises: constructing a one-dimensional list of coefficients from the transform coefficient block; and applying the inverse NSPT to the one-dimensional list of coefficients to reconstruct the residual block.
[0239] Clause 56: The method of clause 55, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a sub-block diagonal scan to the transform coefficient block.
[0240] Clause 57: The method of clause 55, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a horizontal scan to the block of transform coefficients.
[0241] Clause 58: The method of clause 55, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a vertical scan to the transform coefficient block.
[0242] Clause 59: A method according to clause 49, wherein the inverse NSPT is defined as a matrix of size MxN, where M is an integer value representing the number of basis vectors and also the number of rows in the matrix, and where N is an integer value representing the number of support samples of the inverse NSPT.
[0243] Clause 60: The method of clause 59, wherein the matrix comprises 8-bit precision values.
[0244] Clause 61: The method of clause 49, further comprising: selecting the inverse NSPT from a set of possible inverse NSPTs.
[0245] Clause 62: The method according to clause 61, further comprising: selecting the possible inverse NSPT set from a plurality of possible inverse NSPT sets.
[0246] Clause 63: The method of clause 62, wherein selecting the set of possible inverse NSPTs comprises selecting the set of possible inverse NSPTs according to an intra-prediction mode for the block of video data.
[0247] Clause 64: The method of clause 61, wherein selecting the set of possible inverse NSPTs or at least one of the inverse NSPTs comprises selecting the set of possible inverse NSPTs or the at least one of the inverse NSPTs based on a size of the block of video data.
[0248] Clause 65: The method of clause 49, further comprising: performing sign prediction to predict one or more signs of one or more of the transform coefficients.
[0249] Clause 66: The method of clause 49, wherein decoding the block of video data comprises: forming a prediction block for the block of video data; and combining the prediction block with the residual block to form a decoded block of the block of video data.
[0250] Clause 67: A method according to clause 49, wherein the transform coefficient block is one of a 4x16 block or a 16x4 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 20, 24 or 32 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0251] Clause 68: A method according to clause 67, wherein, when there are 20 non-zero-valued transform coefficients, the remaining transform coefficients are 44 zero-valued transform coefficients, when there are 24 non-zero-valued transform coefficients, the remaining transform coefficients are 40 zero-valued transform coefficients, or when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 32 zero-valued transform coefficients.
[0252] Clause 69: A method according to clause 49, wherein the transform coefficient block is one of an 8x16 block or a 16x8 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32 or 40 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0253] Clause 70: A method according to clause 69, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 96 zero-valued transform coefficients, or when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 88 zero-valued transform coefficients.
[0254] Clause 71: A method according to clause 49, wherein the transform coefficient block is a 16x16 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32, 40 or 44 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0255] Clause 72: A method according to clause 71, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 224 zero-valued transform coefficients, when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 216 zero-valued transform coefficients, or when there are 44 non-zero-valued transform coefficients, the remaining transform coefficients are 212 zero-valued transform coefficients.
[0256] Clause 73: The method of clause 49, further comprising: encoding the block of video data prior to decoding the block of video data.
[0257] Item 74: A device for decoding video data, the device comprising: a memory configured to store the video data; and a processing system comprising one or more processors implemented in a circuit system, the processing system being configured to: inversely transform a transform coefficient block of a block of video data using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the block of video data; and decode the block using the residual block.
[0258] Clause 75: The apparatus of clause 74, wherein the processing system is further configured to: form a prediction block for the block of video data using an intra prediction mode; and determine the inverse NSPT based on the intra prediction mode.
[0259] Clause 76: An apparatus according to clause 74, wherein the processing system is further configured to: determine the size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and select the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the video data block.
[0260] Clause 77: An apparatus according to clause 74, wherein the memory is further configured to: store coefficients for multiple inverse NSPTs, the multiple inverse NSPTs including 4x4 inverse NSPT, 4x8 inverse NSPT, 8x4 inverse NSPT, 4x16 inverse NSPT, 16x4 inverse NSPT, 8x8 inverse NSPT, 8x16 inverse NSPT and 16x8 inverse NSPT, and wherein the processing system is configured to retrieve the coefficients for one inverse NSPT among the multiple inverse NSPTs from the memory.
[0261] Clause 78: An apparatus according to clause 74, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes the first video data block, and the residual block includes a first residual block, and wherein the processing system is further configured to: determine that a second transform coefficient block of a second video data block has a second size different from the first size; based on the second size being different from the first size, inverse transform the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and decode the second video data block using the second residual block.
[0262] Clause 79: The apparatus of clause 74, further comprising: a display configured to display the decoded video data.
[0263] Clause 80: The device of clause 74, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0264] Item 81: A device for decoding video data, the device comprising: a unit for inversely transforming a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and a unit for decoding the video data block using the residual block.
[0265] Clause 82: The apparatus of clause 81, further comprising: means for forming a prediction block for the block of video data using an intra prediction mode; and means for determining an inverse NSPT based on the intra prediction mode.
[0266] Clause 83: The apparatus of clause 81, further comprising: a unit for determining a size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and a unit for selecting the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the block.
[0267] Clause 84: The apparatus according to clause 81 further includes: a unit for retrieving coefficients of the inverse NSPT from a memory storing coefficients for a plurality of inverse NSPTs, wherein the plurality of inverse NSPTs include 4x4 inverse NSPT, 4x8 inverse NSPT, 8x4 inverse NSPT, 4x16 inverse NSPT, 16x4 inverse NSPT, 8x8 inverse NSPT, 8x16 inverse NSPT and 16x8 inverse NSPT.
[0268] Item 85: An apparatus according to Item 81, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes the first video data block, and the residual block includes a first residual block, and the apparatus further comprises: a unit for determining that a second transform coefficient block of a second video data block has a second size different from the first size; a unit for inversely transforming the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform based on the second size being different from the first size to reconstruct a second residual block of the second video data block; and a unit for decoding the second video data block using the second residual block.
[0269] Clause 86: The apparatus of clause 81, wherein the means for inverse transforming the transform coefficient block comprises: means for reorganizing the transform coefficient block to form a reorganized transform coefficient block; and means for inverse transforming the reorganized transform coefficient block.
[0270] Item 87: A method for decoding video data, the method comprising: inversely transforming a transform coefficient block of a video data block using an inverse non-separable primary transform (NSPT) to reconstruct a residual block of the video data block without using an inverse separable transform; and decoding the video data block using the residual block.
[0271] Clause 88: The method of clause 87, further comprising: forming a prediction block for the block of video data using an intra-frame prediction mode; and determining the inverse NSPT according to the intra-frame prediction mode.
[0272] Clause 89: The method according to any one of clauses 87 and 88, further comprising: determining the size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and selecting the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the block.
[0273] Clause 90: The method according to any one of clauses 87-89 further includes: retrieving coefficients of the inverse NSPT from a memory storing coefficients for multiple inverse NSPTs, wherein the multiple inverse NSPTs include 4x4 inverse NSPT, 4x8 inverse NSPT, 8x4 inverse NSPT, 4x16 inverse NSPT, 16x4 inverse NSPT, 8x8 inverse NSPT, 8x16 inverse NSPT and 16x8 inverse NSPT.
[0274] Clause 91: A method according to any one of clauses 87-90, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes a first video data block, and the residual block includes a first residual block, the method further comprising: determining that a second transform coefficient block of a second video data block has a second size different from the first size; based on the second size being different from the first size, inversely transforming the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and decoding the second video data block using the second residual block.
[0275] Clause 92: A method according to any of clauses 87-91, wherein inverse transforming the transform coefficient block comprises: reorganizing the transform coefficient block to form a reorganized transform coefficient block; and inverse transforming the reorganized transform coefficient block.
[0276] Clause 93: A method according to any one of clauses 87-92, wherein inverse transforming the transform coefficient block comprises: constructing a one-dimensional list of coefficients from the transform coefficient block; and applying the inverse NSPT to the one-dimensional list of coefficients to reconstruct the residual block.
[0277] Clause 94: The method of clause 93, wherein constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a sub-block diagonal scan to the transform coefficient block.
[0278] Clause 95: The method of any of clauses 93 and 94, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a horizontal scan to the block of transform coefficients.
[0279] Clause 96: A method according to any of clauses 93 to 95, wherein constructing the one-dimensional list of coefficients from the block of transform coefficients comprises applying a vertical scan to the block of transform coefficients.
[0280] Clause 97: A method according to any one of clauses 87-96, wherein the inverse NSPT is defined as a matrix of size MxN, where M is an integer value representing the number of basis vectors and also the number of rows in the matrix, and where N is an integer value representing the number of support samples of the inverse NSPT.
[0281] Clause 98: The method of clause 97, wherein the matrix comprises 8-bit precision values.
[0282] Clause 99: The method of any of clauses 87-98, further comprising: selecting the inverse NSPT from a set of possible inverse NSPTs.
[0283] Clause 100: The method according to clause 99, further comprising: selecting the possible inverse NSPT set from a plurality of possible inverse NSPT sets.
[0284] Clause 101: The method of clause 100, wherein selecting the set of possible inverse NSPTs comprises selecting the set of possible inverse NSPTs according to an intra prediction mode for the block of video data.
[0285] Clause 102: The method of any of clauses 100 and 101, wherein selecting the set of possible inverse NSPTs or at least one of the inverse NSPTs comprises selecting the set of possible inverse NSPTs or the at least one of the inverse NSPTs based on a size of the block of video data.
[0286] Clause 103: The method of any of clauses 87-102, further comprising: performing sign prediction to predict one or more signs of one or more of the transform coefficients.
[0287] Clause 104: The method of any of clauses 87-103, wherein decoding the block of video data comprises: forming a prediction block for the block of video data; and combining the prediction block with the residual block to form a decoded block of the block of video data.
[0288] Clause 105: A method according to any one of clauses 87-104, wherein the transform coefficient block is one of a 4x16 block or a 16x4 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 20, 24 or 32 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0289] Clause 106: A method according to clause 105, wherein, when there are 20 non-zero-valued transform coefficients, the remaining transform coefficients are 44 zero-valued transform coefficients, when there are 24 non-zero-valued transform coefficients, the remaining transform coefficients are 40 zero-valued transform coefficients, or when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 32 zero-valued transform coefficients.
[0290] Clause 107: A method according to any one of clauses 87-106, wherein the transform coefficient block is one of an 8x16 block or a 16x8 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32 or 40 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0291] Clause 108: A method according to clause 107, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 96 zero-valued transform coefficients, or when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 88 zero-valued transform coefficients.
[0292] Clause 109: A method according to any one of clauses 87-108, wherein the transform coefficient block is a 16x16 block, and wherein inverse transforming the transform coefficient block includes inverse transforming 32, 40 or 44 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
[0293] Clause 110: A method according to clause 109, wherein when there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 224 zero-valued transform coefficients, when there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 216 zero-valued transform coefficients, or when there are 44 non-zero-valued transform coefficients, the remaining transform coefficients are 212 zero-valued transform coefficients.
[0294] Clause 111: The method of any of clauses 87-110, further comprising: encoding the block of video data prior to decoding the block of video data.
[0295] Item 112: A device for decoding video data, the device comprising: a memory configured to store the video data; and a processing system comprising one or more processors implemented in a circuit system, the processing system being configured to: inversely transform a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and decode the block using the residual block.
[0296] Clause 113: The apparatus of clause 112, wherein the processing system is further configured to: form a prediction block for a block of video data using an intra prediction mode; and determine the inverse NSPT based on the intra prediction mode.
[0297] Clause 114: An apparatus according to any one of clauses 112 and 113, wherein the processing system is further configured to: determine the size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and select the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the video data block.
[0298] Clause 115: A device according to any one of clauses 112-114, wherein the memory is further configured to: store coefficients for multiple inverse NSPTs, the multiple inverse NSPTs including 4x4 inverse NSPT, 4x8 inverse NSPT, 8x4 inverse NSPT, 4x16 inverse NSPT, 16x4 inverse NSPT, 8x8 inverse NSPT, 8x16 inverse NSPT and 16x8 inverse NSPT, and wherein the processing system is configured to retrieve the coefficients for one inverse NSPT among the multiple inverse NSPTs from the memory.
[0299] Clause 116: An apparatus according to any one of clauses 112-115, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes first video data, and the residual block includes a first residual block, and wherein the processing system is further configured to: determine that a second transform coefficient block of a second video data block has a second size different from the first size; based on the second size being different from the first size, inversely transform the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and decode the second video data block using the second residual block.
[0300] Clause 117: The apparatus of any of clauses 112-116, further comprising: a display configured to display the decoded video data.
[0301] Clause 118: The device of any of clauses 112-117, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0302] Item 119: A device for decoding video data, the device comprising: a unit for inversely transforming a transform coefficient block of a video data block using an inverse non-separable main transform (NSPT) without using an inverse separable transform to reconstruct a residual block of the video data block; and a unit for decoding the video data block using the residual block.
[0303] Clause 120: The apparatus of clause 119, further comprising: means for forming a prediction block for the block of video data using an intra prediction mode; and means for determining an inverse NSPT based on the intra prediction mode.
[0304] Clause 121: The apparatus of any one of clauses 119 and 120, further comprising: a unit for determining the size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and a unit for selecting the inverse NSPT so that the inverse NSPT has a size corresponding to the size of the block.
[0305] Clause 122: The apparatus according to any one of clauses 119-121 further comprises: a unit for retrieving coefficients of the inverse NSPT from a memory storing coefficients for a plurality of inverse NSPTs, wherein the plurality of inverse NSPTs comprises a 4x4 inverse NSPT, a 4x8 inverse NSPT, an 8x4 inverse NSPT, a 4x16 inverse NSPT, a 16x4 inverse NSPT, an 8x8 inverse NSPT, an 8x16 inverse NSPT and a 16x8 inverse NSPT.
[0306] Clause 123: A device according to any one of clauses 119-122, wherein the transform coefficient block includes a first transform coefficient block having a first size, the video data block includes the first video data block, and the residual block includes a first residual block, and the device further includes: a unit for determining that a second transform coefficient block of a second video data block has a second size different from the first size; a unit for inversely transforming the second transform coefficient block using the inverse separable transform and the inverse low-frequency non-separable transform (LFNST) transform based on the second size being different from the first size to reconstruct a second residual block of the second video data block; and a unit for decoding the second video data block using the second residual block.
[0307] Clause 124: An apparatus according to any one of clauses 119-123, wherein the unit for inverse transforming the transform coefficient block comprises: a unit for reorganizing the transform coefficient block to form a reorganized transform coefficient block; and a unit for inverse transforming the reorganized transform coefficient block.
[0308] It is to be appreciated that, according to examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary to implement the techniques). In addition, in some examples, actions or events may be performed concurrently rather than sequentially, such as through multithreading, interrupt handling, or multiple processors.
[0309] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium, the communication medium including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-temporary tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.
[0310] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, optical fiber cable, twisted pair, digital subscriber line (DSL) or wireless technology (e.g., infrared, radio and microwave) is used to transmit instructions from a website, server or other remote source, then the coaxial cable, optical fiber cable, twisted pair, DSL or wireless technology (e.g., infrared, radio and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other temporary media, but are instead directed to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser optical disks, optical disks, digital versatile disks (DVDs), floppy disks and blue-ray disks, wherein disks usually copy data magnetically, while optical disks use lasers to optically copy data. Combinations of the above should also be included within the scope of computer-readable media.
[0311] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. In addition, the techniques may be implemented entirely in one or more circuits or logic elements.
[0312] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of a device configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Specifically, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.
[0313] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for decoding video data, the method comprising: Inversely transforming a transform coefficient block of a video data block using an inverse non-separable primary transform (NSPT) to reconstruct a residual block of the video data block without using an inverse separable transform; as well as The video data block is decoded using the residual block.
2. The method according to claim 1, further comprising: forming a prediction block for the block of video data using an intra prediction mode; as well as The inverse NSPT is determined according to the intra prediction mode.
3. The method according to claim 1, further comprising: Determining a size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8; and The inverse NSPT is selected such that the inverse NSPT has a size corresponding to a size of a block.
4. The method according to claim 1, further comprising: Coefficients of the inverse NSPT are retrieved from a memory storing coefficients for a plurality of inverse NSPTs, the plurality of inverse NSPTs including a 4x4 inverse NSPT, a 4x8 inverse NSPT, an 8x4 inverse NSPT, a 4x16 inverse NSPT, a 16x4 inverse NSPT, an 8x8 inverse NSPT, an 8x16 inverse NSPT, and a 16x8 inverse NSPT.
5. The method according to claim 1, wherein: The transform coefficient block comprises a first transform coefficient block having a first size, the video data block comprises a first video data block, and the residual block comprises a first residual block, the method further comprising: determining that a second transform coefficient block for a second block of video data has a second size different than the first size; Based on the second size being different from the first size, inversely transforming the second transform coefficient block using the inverse separable transform and an inverse low frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and The second video data block is decoded using the second residual block.
6. The method according to claim 1, wherein: Performing an inverse transform on the transform coefficient block comprises: reorganizing the transform coefficient block to form a reorganized transform coefficient block; and The reorganized transform coefficient block is inversely transformed.
7. The method according to claim 1, wherein: Performing an inverse transform on the transform coefficient block comprises: constructing a one-dimensional list of coefficients from the block of transform coefficients; and The inverse NSPT is applied to the one-dimensional list of coefficients to reconstruct the residual block.
8. The method according to claim 7, wherein: Constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a sub-block diagonal scan to the transform coefficient block.
9. The method according to claim 7, wherein: Constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a horizontal scan to the transform coefficient block.
10. The method according to claim 7, wherein: Constructing the one-dimensional list of coefficients from the transform coefficient block comprises applying a vertical scan to the transform coefficient block.
11. The method according to claim 1, wherein: The inverse NSPT is defined as a matrix of size MxN, where M is an integer value representing the number of basis vectors and also the number of rows in the matrix, and where N is an integer value representing the number of support samples of the inverse NSPT.
12. The method according to claim 11, wherein: The matrix includes 8-bit precision values.
13. The method according to claim 1, further comprising: The inverse NSPT is selected from a set of possible inverse NSPTs.
14. The method according to claim 13, further comprising: The possible inverse NSPT set is selected from a plurality of possible inverse NSPT sets.
15. The method according to claim 14, wherein: Selecting the set of possible inverse NSPTs includes selecting the set of possible inverse NSPTs according to an intra-prediction mode for the block of video data.
16. The method according to claim 13, wherein: Selecting the inverse NSPT includes selecting the inverse NSPT according to a size of the video data block.
17. The method according to claim 1, further comprising: Sign prediction is performed to predict one or more signs of one or more of the transform coefficients.
18. The method according to claim 1, wherein: Decoding the video data block comprises: forming a prediction block for the block of video data; and The prediction block is combined with the residual block to form a decoded block of the block of video data.
19. The method according to claim 1, wherein: The transform coefficient block is one of a 4x16 block or a 16x4 block, and wherein inverse transforming the transform coefficient block comprises inverse transforming 20, 24 or 32 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
20. The method according to claim 19, wherein: When there are 20 non-zero-valued transform coefficients, the remaining transform coefficients are 44 zero-valued transform coefficients, When there are 24 non-zero-valued transform coefficients, the remaining transform coefficients are 40 zero-valued transform coefficients, or When there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 32 zero-valued transform coefficients.
21. The method according to claim 1, wherein: The transform coefficient block is one of an 8x16 block or a 16x8 block, and wherein inverse transforming the transform coefficient block comprises inverse transforming 32 or 40 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
22. The method according to claim 21, wherein: When there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 96 zero-valued transform coefficients, or When there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 88 zero-valued transform coefficients.
23. The method according to claim 1, wherein: The transform coefficient block is a 16x16 block, and wherein inverse transforming the transform coefficient block comprises inverse transforming 32, 40 or 44 non-zero-valued transform coefficients and zero-valued transform coefficients for the remaining transform coefficients.
24. The method according to claim 23, wherein: When there are 32 non-zero-valued transform coefficients, the remaining transform coefficients are 224 zero-valued transform coefficients, When there are 40 non-zero-valued transform coefficients, the remaining transform coefficients are 216 zero-valued transform coefficients, or When there are 44 non-zero-valued transform coefficients, the remaining transform coefficients are 212 zero-valued transform coefficients.
25. The method of claim 1, further comprising: The block of video data is encoded before decoding the block of video data.
26. A device for decoding video data, the device comprising: a memory configured to store video data; as well as A processing system comprising one or more processors implemented in a circuit system, the processing system being configured to: Inversely transforming a transform coefficient block of a video data block using an inverse non-separable primary transform (NSPT) to reconstruct a residual block of the video data block without using an inverse separable transform; as well as The block is decoded using the residual block.
27. The apparatus of claim 26, wherein: The processing system is also configured to: forming a prediction block for the block of video data using an intra prediction mode; and The inverse NSPT is determined according to the intra prediction mode.
28. The apparatus of claim 26, wherein: The processing system is also configured to: Determining a size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, or 16x8; and The inverse NSPT is selected such that the inverse NSPT has a size corresponding to a size of the block of video data.
29. The apparatus of claim 26, wherein: The memory is also configured to store coefficients for a plurality of inverse NSPTs, the plurality of inverse NSPTs comprising a 4x4 inverse NSPT, a 4x8 inverse NSPT, an 8x4 inverse NSPT, a 4x16 inverse NSPT, a 16x4 inverse NSPT, an 8x8 inverse NSPT, an 8x16 inverse NSPT, and a 16x8 inverse NSPT, and wherein the processing system is configured to retrieve the coefficients for one of the plurality of inverse NSPTs from the memory.
30. The apparatus of claim 26, wherein: The transform coefficient block comprises a first transform coefficient block having a first size, the video data block comprises a first video data block, and the residual block comprises a first residual block, and wherein the processing system is further configured to: determining that a second transform coefficient block for a second block of video data has a second size different than the first size; Based on the second size being different from the first size, inversely transforming the second transform coefficient block using the inverse separable transform and an inverse low frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second video data block; and The second video data block is decoded using the second residual block.
31. The apparatus of claim 26, further comprising: A display configured to display the decoded video data.
32. The apparatus of claim 26, wherein: The device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
33. An apparatus for decoding video data, the apparatus comprising: means for inversely transforming a block of transform coefficients of a block of video data using an inverse non-splittable primary transform (NSPT) to reconstruct a residual block of the block of video data without using an inverse separable transform; as well as Means for decoding the block of video data using the residual block.
34. The apparatus of claim 33, further comprising: means for forming a prediction block for the block of video data using an intra prediction mode; as well as Means for determining the inverse NSPT based on an intra prediction mode.
35. The apparatus of claim 33, further comprising: means for determining a size of the transform coefficient block to be one of 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16 or 16x8; and means for selecting the inverse NSPT such that the inverse NSPT has a size corresponding to a size of a block.
36. The apparatus of claim 33, further comprising: A unit for retrieving coefficients of an inverse NSPT from a memory storing coefficients for a plurality of inverse NSPTs, the plurality of inverse NSPTs comprising a 4x4 inverse NSPT, a 4x8 inverse NSPT, an 8x4 inverse NSPT, a 4x16 inverse NSPT, a 16x4 inverse NSPT, an 8x8 inverse NSPT, an 8x16 inverse NSPT, and a 16x8 inverse NSPT.
37. The apparatus of claim 33, wherein: The transform coefficient block comprises a first transform coefficient block having a first size, the video data block comprises a first video data block, and the residual block comprises a first residual block, the apparatus further comprising: for determining that a second transform coefficient block for a second block of video data has units of a second size different from the first size; means for inversely transforming the second transform coefficient block using the inverse separable transform and an inverse low frequency non-separable transform (LFNST) transform to reconstruct a second residual block of the second block of video data based on the second size being different from the first size; and Means for decoding the second block of video data using the second residual block.
38. The apparatus of claim 33, wherein: The unit for inverse transforming the transform coefficient block comprises: means for reorganizing the transform coefficient block to form a reorganized transform coefficient block; and Means for inverse transforming the reorganized block of transform coefficients.