Method and device for decoding video data
By determining the maximum possible value of the quadratic element of the quadratic transformation of the video data block and using a unified binarization or reverse binarization scheme, the problem of low processing efficiency of the quadratic element of the video data block in the prior art is solved, and more efficient video data block decoding is achieved.
Patent Information
- Application Number
- CN202110709660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-05-02
- Filing Date
- 2017-05-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2037-05-03
AI Technical Summary
When existing video decoding technology processes the secondary transformation syntax elements of video data blocks, it is difficult to implement a unified binary or reverse binary scheme, resulting in low bitstream efficiency.
By determining the maximum possible value of the quadratic syntax element of the video data block and using a common binarization or reverse binarization scheme, unified processing of the quadratic syntax element is achieved.
The decoding efficiency of video data blocks is improved, the bit rate of the bit stream is reduced, and the same binarization or reverse binarization scheme can be applied regardless of the maximum possible value.
Smart Images

Figure CN113453019B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of May 3, 2017, application number "201780026951.8", and invention name "Method and device for decoding video data".
[0002] This application claims the benefit of each of the following U.S. provisional applications:
[0003] U.S. Provisional Application No. 62 / 331,290, filed May 3, 2016;
[0004] U.S. Provisional Application No. 62 / 332,425, filed May 5, 2016;
[0005] U.S. Provisional Application No. 62 / 337,310, filed May 16, 2016;
[0006] U.S. Provisional Application No. 62 / 340,949, filed May 24, 2016; and
[0007] U.S. Provisional Application No. 62 / 365,853, filed on July 22, 2016,
[0008] The entire contents of each of the U.S. Provisional Applications are hereby incorporated by reference. Technical Field
[0009] The present invention relates to video decoding. Background Art
[0010] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, electronic book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio telephones (so-called "smart phones"), video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) standards, and extensions of these standards. Video devices may transmit, receive, encode, decode, and / or store digital video information more efficiently by implementing these video coding techniques.
[0011] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent to video sequences. For block-based video coding, a video slice (e.g., a video picture or portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in neighboring blocks in the same picture, or temporal prediction with respect to reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames.
[0012] Spatial or temporal prediction produces a predictive block for the block to be coded. The residual data represents the pixel difference between the original block to be coded and the predictive block. The inter-coded block is encoded according to the motion vector pointing to the reference sample block forming the predictive block and the residual data indicating the difference between the coded block and the predictive block. The intra-coded block is encoded according to the intra-coding mode and the residual data. For further compression, the residual data can be transformed from the pixel domain to the transform domain, thereby producing residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, and entropy coding can be applied to achieve even more compression. Summary of the invention
[0013] In general, this disclosure describes techniques related to entropy coding (encoding or decoding) secondary transform syntax elements for blocks of video data. The secondary transform syntax elements may include, for example, non-separable secondary transform (NSST) syntax elements, rotation transform syntax elements, and the like. In general, entropy coding of these syntax elements may include binarization or inverse binarization. The binarization or inverse binarization scheme may be unified such that the same binarization or inverse binarization scheme is applied regardless of the maximum possible value of the secondary transform syntax elements. The techniques of this disclosure may further include coding (encoding or decoding) a signaling unit syntax element, wherein a signaling unit may include two or more adjacent blocks. The signaling unit syntax element may precede each of the blocks, or be placed immediately before the block to which the signaling unit syntax element is applied (in coding order).
[0014] In one example, a method of decoding video data includes determining a maximum possible value of a secondary transform syntax element for a block of video data; entropy decoding the value of the secondary transform syntax element for the block to form a binarized value representing a secondary transform for the block; de-binarizing the value of the secondary transform syntax element using a common de-binarization scheme regardless of the maximum possible value to determine the secondary transform for the block; and inverse transforming transform coefficients of the block using the determined secondary transform.
[0015] In another example, a device for decoding video data includes: a memory configured to store video data; and one or more processors implemented in a circuit system and configured to: determine a maximum possible value of a secondary transform syntax element for a block of video data; entropy decode the value of the secondary transform syntax element for the block to form a binarized value representing a secondary transform for the block; de-binarize the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value to determine the secondary transform for the block; and inverse transform transform coefficients of the block using the determined secondary transform.
[0016] In another example, a device for decoding video data includes: means for determining a maximum possible value of a secondary transform syntax element for a block of video data; means for entropy decoding the value of the secondary transform syntax element for the block to form a binarized value representing a secondary transform for the block; means for de-binarizing the value of the secondary transform syntax element using a common de-binarization scheme regardless of the maximum possible value to determine the secondary transform for the block; and means for inverse transforming transform coefficients of the block using the determined secondary transform.
[0017] In another example, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) has stored thereon instructions that, when executed, cause one or more processors to: determine a maximum possible value for a secondary transform syntax element for a block of video data; entropy decode the value of the secondary transform syntax element for the block to form a binarized value representing a secondary transform for the block; de-binarize the value of the secondary transform syntax element using a common de-binarization scheme regardless of the maximum possible value to determine the secondary transform for the block; and inverse transform transform coefficients of the block using the determined secondary transform.
[0018] In another example, a method of encoding video data includes: transforming intermediate transform coefficients of a block of video data using a secondary transform; determining a maximum possible value of a secondary transform syntax element for the block, the value of the secondary transform syntax element representing the secondary transform; binarizing the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value; and entropy encoding the binarized value of the secondary transform syntax element for the block to form a binarized value representing the secondary transform for the block.
[0019] In another example, a device for encoding video data includes a memory configured to store video data; and one or more processors implemented in a circuit system and configured to: transform intermediate transform coefficients of a block of video data using a secondary transform; determine a maximum possible value of a secondary transform syntax element for the block, the value of the secondary transform syntax element representing the secondary transform; binarize the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value; and entropy encode the binarized value of the secondary transform syntax element for the block to form a binarized value representing the secondary transform for the block.
[0020] The details of one or more examples are set forth below in the accompanying drawings and the detailed description. Other features, objects, and advantages will be apparent from the detailed description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 2 is a block diagram illustrating an example video encoding and decoding system that may utilize techniques for binarizing secondary transform indices.
[0022] Figure 2 is a block diagram illustrating an example of a video encoder that may implement techniques for binarizing secondary transform indices.
[0023] Figure 3 is a block diagram of an example entropy encoding unit that may be configured to perform CABAC according to the techniques of this disclosure.
[0024] Figure 4 is a block diagram illustrating an example of a video decoder that may implement techniques for binarizing secondary transform indices.
[0025] Figure 5 is a block diagram of an example entropy encoding unit that may be configured to perform CABAC according to the techniques of this disclosure.
[0026] Figure 6 A flowchart illustrating an example method of encoding video data in accordance with the techniques of this disclosure is shown.
[0027] Figure 72 is a flow chart illustrating an example of a method of decoding video data in accordance with the techniques of this disclosure. DETAILED DESCRIPTION
[0028] Video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC (Advanced Video Coding)), ITU-T H.265 (also known as HEVC or "High Efficiency Video Coding"), including extensions such as scalable video coding (SVC), multi-view video coding (MVC), and screen content coding (SCC). The techniques of this disclosure may be applied to these or future video coding standards, such as the Joint Video Exploration Team (JVET) test model (which may also be referred to as the Joint Exploration Model-JEM), which is undergoing development activities other than HEVC. Video coding standards also include proprietary video codecs, such as Google VP8, VP9, VP10, and video codecs developed by other organizations (e.g., the Alliance for Open Media).
[0029] In the JVET test model, there is an intra prediction method called position-dependent intra prediction combination (PDPC). The JVET test model also includes a non-separable secondary transform (NSST) tool. Both the PDPC tool and the NSST tool use syntax elements (e.g., indices) to indicate whether the corresponding tool is applied and which transformation to use. For example, an index of 0 may mean that the tool is not used.
[0030] The maximum number of NSST indices for a block of video data may depend on the intra-prediction mode or partition size of the block. In one example, if the intra-prediction mode is PLANAR or DC and the partition size is 2N×2N, the maximum number of NSST indices is 3, otherwise the maximum number of NSST indices is 4. Under the JVET test model, two types of binarization are used to represent the NSST indices. In the JVET test model, if the maximum value is 3, truncated unary binarization is used, otherwise fixed binary binarization is applied. In the JVET test model, if the PDPC index is not equal to 0, NSST is not applied and the NSST index is not signaled.
[0031] This disclosure describes various techniques that can be applied alone or in any combination to improve, for example, the coding of NSST syntax elements (e.g., NSST index and / or NSST flag). For example, these techniques can improve the operation of a video encoder / video decoder and thereby improve bitstream efficiency in that these techniques can reduce the bitrate of the bitstream relative to the current JVET test model.
[0032] Figure 1 1 is a block diagram illustrating an example video encoding and decoding system 10 that may utilize techniques for binarizing secondary transform indices. Figure 1 As shown, system 10 includes a source device 12 that provides encoded video data to be decoded at a later time by a destination device 14. In particular, source device 12 provides the video data to destination device 14 via a computer-readable medium 16. Source device 12 and destination device 14 may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart” phones, so-called “smart” tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 12 and destination device 14 may be equipped for wireless communication.
[0033] Destination device 14 may receive the encoded video data to be decoded via computer-readable medium 16. Computer-readable medium 16 may include any type of media or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, computer-readable medium 16 may include a communication medium to enable source device 12 to transmit the encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.
[0034] In some examples, the encoded data may be output from output interface 22 to a storage device. Similarly, the encoded data may be accessed from the storage device by the input interface. The storage device may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device may correspond to a file server or another intermediate storage device, which may store the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and transmitting that encoded video data to the destination device 14. Example file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection, including an Internet connection. This connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0035] The techniques of this disclosure are not necessarily limited to wireless applications or settings. The techniques may be applied to video coding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (e.g., HTTP Dynamic Adaptive Streaming (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system 10 may be configured to support one-way or two-way video transmissions to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0036] exist Figure 1 In an example of , source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to this disclosure, video encoder 20 of source device 12 may be configured to apply techniques for binarizing secondary transform indices. In other examples, source device and destination device may include other components or arrangements. For example, source device 12 may receive video data from an external video source 18, such as an external camera. Likewise, destination device 14 may interface with an external display device rather than including an integrated display device.
[0037] Figure 1The illustrated system 10 is merely one example. The techniques for binarizing the secondary transform index may be performed by any digital video encoding and / or decoding device. Although the techniques of the present invention are typically performed by a video encoding device, the techniques may also be performed by a video encoder / decoder (commonly referred to as a "codec (CODEC)"). In addition, the techniques of the present invention may also be performed by a video preprocessor. Source device 12 and destination device 14 are merely examples of these decoding devices that generate decoded video data for transmission to destination device 14 by source device 12. In some examples, devices 12, 14 may operate in a substantially symmetrical manner such that each of devices 12, 14 includes video encoding and decoding components. Thus, system 10 may support one-way or two-way video transmission between video devices 12, 14, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0038] The video source 18 of the source device 12 may include a video capture device, such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source 18 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source 18 is a video camera, the source device 12 and the destination device 14 may form a so-called camera phone or video phone. However, as mentioned above, the techniques described in this disclosure may be applicable to video coding in general, and may be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video information may then be output by the output interface 22 to the computer-readable medium 16.
[0039] Computer-readable medium 16 may include: temporary media, such as wireless broadcast or wired network transmission; or storage media (i.e., non-temporary storage media), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from source device 12 and provide the encoded video data to destination device 14, for example, via a network transmission. Similarly, a computing device of a media production facility (such as a disc stamping facility) may receive encoded video data from source device 12 and produce a disc containing the encoded video data. Therefore, in various examples, computer-readable medium 16 may be understood to include one or more computer-readable media in various forms.
[0040] Input interface 28 of destination device 14 receives information from computer-readable medium 16. The information of computer-readable medium 16 may include syntax information defined by video encoder 20, which is also used by video decoder 30, including syntax elements that describe characteristics and / or processing of blocks and other coded units. Display device 32 displays the decoded video data to a user, and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0041] Video encoder 20 and video decoder 30 may operate according to a video coding standard, such as the High Efficiency Video Coding (HEVC) standard, also referred to as ITU-T H.265. Alternatively, video encoder 20 and video decoder 30 may operate according to other proprietary or industry standards, such as the ITU-T H.264 standard (alternatively referred to as MPEG-4), Part 10, Advanced Video Coding (AVC), or extensions of these standards. However, the techniques of this disclosure are not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although Figure 1 20 and video decoder 30 may each be integrated with an audio encoder and decoder, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or in separate data streams. Where applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).
[0042] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, a device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in a respective device.
[0043] In general, according to ITU-T H.265, a video picture may be divided into a series of coding tree units (CTUs) (or largest coding units (LCUs)) that may include both luma samples and chroma samples. Alternatively, a CTU may include monochrome data (i.e., only luma samples). Syntax data within the bitstream may define the size of a CTU, which is the largest coding unit in terms of the number of pixels. A slice includes a number of consecutive CTUs in coding order. A video picture may be partitioned into one or more slices. Each CTU may be split into coding units (CUs) according to a quadtree. In general, a quadtree data structure includes one node per CU, with the root node corresponding to the CTU. If a CU is split into four sub-CUs, the node corresponding to the CU includes four leaf nodes, each of which corresponds to one of the sub-CUs.
[0044] Each node of the quadtree data structure may provide syntax data for the corresponding CU. For example, a node in the quadtree may include a split flag indicating whether the CU corresponding to the node is split into sub-CUs. Syntax elements for a CU may be defined recursively and may depend on whether the CU is split into sub-CUs. If a CU is not further split, it is referred to as a leaf-CU. In the present invention, the four sub-CUs of a leaf-CU will be referred to as leaf-CUs even if there is no explicit splitting of the original leaf-CU. For example, if a CU in size 16×16 is not further split, the four 8×8 sub-CUs will also be referred to as leaf-CUs, but a 16×16 CU is never split.
[0045] A CU has a similar purpose to a macroblock of the H.264 standard, except that a CU does not have a size distinction. For example, a CTU can be split into four child nodes (also referred to as sub-CUs), and each child node can be a parent node and split into another four child nodes. The final unsplit child nodes, referred to as leaf nodes of a quadtree, include coding nodes, also referred to as leaf CUs. Syntax data associated with the coded bitstream may define the maximum number of times a CTU can be split (referred to as the maximum CU depth), and may also define the minimum size of a coding node. Thus, the bitstream may also define a minimum coding unit (SCU). This disclosure uses the term "block" to refer to any of a CU, a prediction unit (PU), or a transform unit (TU) in the context of HEVC, or to similar data structures in the context of other standards (e.g., macroblocks and subblocks thereof in H.264 / AVC).
[0046] A CU includes a coding node and prediction units (PUs) and transform units (TUs) associated with the coding node. The size of the CU corresponds to the size of the coding node and is generally square in shape. The size of the CU may range from 8×8 pixels to the size of a CTU having a maximum size, for example, 64×64 pixels or larger. Each CU may contain one or more PUs and one or more TUs. Syntax data associated with a CU may describe, for example, partitioning the CU into one or more PUs. The partitioning mode may differ between whether the CU is skipped or direct mode encoded, intra-prediction mode encoded, or inter-prediction mode encoded. The PU may be partitioned into a non-square shape. Syntax data associated with a CU may also describe, for example, partitioning the CU into one or more TUs according to a quadtree. A TU may be square or non-square (e.g., rectangular) in shape.
[0047] The HEVC standard allows for transforms according to TUs, which may be different for different CUs. TUs are typically sized based on the size of PUs within a given CU defined for a partitioned CTU, but this may not always be the case. TUs are typically the same size or smaller than a PU. In some examples, a quadtree structure called a "residual quadtree" (RQT) may be used to subdivide the residual samples corresponding to a CU into smaller units. The leaf nodes of the RQT may be referred to as transform units (TUs). Pixel difference values associated with a TU may be transformed to produce transform coefficients that may be quantized.
[0048] A leaf-CU may include one or more prediction units (PUs). In general, a PU represents a spatial region corresponding to all or part of a corresponding CU, and may include data for retrieving and / or generating reference samples for the PU. In addition, a PU includes data related to prediction. For example, when a PU is intra-mode encoded, data for the PU may be included in a residual quad tree (RQT), and the residual RQT may include data describing an intra-prediction mode for a TU corresponding to the PU. The RQT may also be referred to as a transform tree. In some instances, an intra-prediction mode may be signaled in a leaf-CU syntax instead of an RQT. As another example, when a PU is inter-mode encoded, the PU may include data defining motion information (e.g., one or more motion vectors) for the PU. The data defining a motion vector for the PU may describe, for example, a horizontal component of the motion vector, a vertical component of the motion vector, a resolution of the motion vector (e.g., quarter-pixel precision or eighth-pixel precision), a reference picture to which the motion vector points, and / or a reference picture list (e.g., list 0, list 1, or list C) for the motion vector.
[0049] A leaf-CU having one or more PUs may also include one or more transform units (TUs). A transform unit may be specified using an RQT (also referred to as a TU quadtree structure), as discussed above. For example, a split flag may indicate whether a leaf-CU is split into four transform units. Each transform unit may then be further split into additional sub-TUs. When a TU is not further split, it may be referred to as a leaf-TU. Typically, for intra-coding, all leaf-TUs belonging to a leaf-CU share the same intra-prediction mode. That is, the same intra-prediction mode is typically applied to calculate predicted values for all TUs of a leaf-CU. For intra-coding, a video encoder may use an intra-prediction mode to calculate a residual value for each leaf-TU as the difference between the portion of the CU corresponding to the TU and the original block. A TU is not necessarily limited to the size of a PU. Therefore, a TU may be larger or smaller than a PU. For intra-coding, a PU may be co-located with a corresponding leaf-TU for the same CU. In some examples, the maximum size of a leaf-TU may correspond to the size of the corresponding leaf-CU.
[0050] In addition, the TUs of a leaf-CU may also be associated with corresponding quadtree data structures, referred to as residual quadtrees (RQTs). That is, a leaf-CU may include a quadtree indicating how the leaf-CU is partitioned into TUs. The root node of a TU quadtree typically corresponds to a leaf-CU, while the root node of a CU quadtree typically corresponds to a CTU (or LCU). A TU of an RQT that is not split is referred to as a leaf-TU. In general, unless otherwise noted, this disclosure uses the terms CU and TU to refer to a leaf-CU and a leaf-TU, respectively.
[0051] A video sequence typically includes a series of video frames or pictures, which starts at a random access point (RAP) picture. A video sequence may include syntax data in a sequence parameter set (SPS) that represents characteristics of the video sequence. Each slice of a picture may include slice syntax data that describes a coding mode for the corresponding slice. Video encoder 20 typically operates on video blocks within individual video slices in order to encode the video data. A video block may correspond to a coding node within a CU. A video block may have a fixed or varying size, and its size may differ according to a specified coding standard.
[0052] As an example, prediction may be performed for PUs of various sizes. Assuming a particular CU is 2N×2N in size, intra prediction may be performed for PU sizes of 2N×2N or N×N, and inter prediction may be performed for symmetric PU sizes of 2N×2N, 2N×N, N×2N, or N×N. Asymmetric partitioning for inter prediction may also be performed for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N. In an asymmetric partitioning, one direction of the CU is not partitioned, while the other direction is partitioned into 25% and 75%. The portion of the CU corresponding to the 25% partition is indicated by an indication of "n" followed by "Up," "Down," "Left," or "Right." Thus, for example, "2N×nU" refers to a 2N×2N CU that is partitioned horizontally such that a 2N×0.5N PU is at the top and a 2N×1.5N PU is at the bottom.
[0053] In this disclosure, "NxN" and "N by N" may be used interchangeably to refer to the pixel dimensions of a video block in terms of the vertical and horizontal dimensions, e.g., 16x16 pixels or 16 by 16 pixels. In general, a 16x16 block will have 16 pixels in the vertical direction (y=16) and 16 pixels in the horizontal direction (x=16). Likewise, an NxN block typically has N pixels in the vertical direction and N pixels in the horizontal direction, where N represents a non-negative integer value. The pixels in a block may be arranged in rows and columns. Furthermore, a block need not necessarily have the same number of pixels in the horizontal direction as in the vertical direction. For example, a block may include NxM pixels, where M is not necessarily equal to N.
[0054] After intra-predictive or inter-predictive coding using a PU of a CU, video encoder 20 may calculate residual data for a TU of the CU. A PU may include syntax data describing a method or mode of generating predictive pixel data in a spatial domain (also referred to as a pixel domain), and a TU may include coefficients in a transform domain after applying a transform (e.g., a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform) to the residual video data. The residual data may correspond to pixel differences between pixels of an unencoded picture and prediction values corresponding to the PU. Video encoder 20 may form a TU to include quantized transform coefficients representing the residual data for the CU. That is, video encoder 20 may calculate the residual data (in the form of a residual block), transform the residual block to generate a block of transform coefficients, and then quantize the transform coefficients to form quantized transform coefficients. Video encoder 20 may form a TU including quantized transform coefficients, as well as other syntax information (e.g., split information for the TU).
[0055] As mentioned above, after any transforms used to produce transform coefficients, video encoder 20 may perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing transform coefficients to possibly reduce the amount of data used to represent the coefficients, providing further compression. The quantization process may reduce the bit depth associated with some or all of the coefficients. For example, an n-bit value may be truncated to an m-bit value during quantization, where n is greater than m.
[0056] After quantization, the video encoder may scan the transform coefficients, thereby generating a one-dimensional vector from the two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and therefore lower frequency) coefficients at the front of the array and lower energy (and therefore higher frequency) coefficients at the back of the array. In some examples, video encoder 20 may utilize a predefined scan order to scan the quantized transform coefficients to generate a serialized vector that can be entropy encoded. In other examples, video encoder 20 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 20 may entropy encode the one-dimensional vector, for example, according to context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method. Video encoder 20 may also entropy encode syntax elements associated with the encoded video data for use by video decoder 30 to decode the video data.
[0057] To perform CABAC, video encoder 20 may assign context within a context model to a symbol to be transmitted. The context may be related to, for example, whether neighboring values of the symbol are non-zero. To perform CAVLC, video encoder 20 may select a variable length code for the symbol to be transmitted. Codewords in VLC may be constructed so that relatively shorter codes correspond to more likely symbols, while longer codes correspond to less likely symbols. In this way, using VLC may achieve bit savings compared to, for example, using codewords of equal length for each symbol to be transmitted. Probability determinations may be based on the context assigned to the symbol.
[0058] In general, video decoder 30 performs a process substantially similar to, but inverse of, the process performed by video encoder 20 to decode the encoded data. For example, video decoder 30 inverse quantizes and inverse transforms the coefficients of the received TU to reproduce the residual block. Video decoder 30 uses the signaled prediction mode (intra-prediction or inter-prediction) to form a predicted block. Then, video decoder 30 combines the predicted block with the residual block (pixel by pixel) to reproduce the original block. Additional processing may be performed, such as performing a deblocking process to reduce visual artifacts along block boundaries. In addition, video decoder 30 may decode syntax elements using CABAC in a manner substantially similar to, but inverse of, the CABAC encoding process of video encoder 20.
[0059] According to the techniques of this disclosure, a video coder (e.g., video encoder 20 or video decoder 30) may unify the binarization of NSST syntax elements. For example, the video coder may be configured to use only one binarization (e.g., truncated or truncated unary binarization). For a block for which an NSST syntax element is coded, a maximum value of the NSST syntax element may be defined (and thus determined by the video coder) according to an intra-mode and optionally according to a block size condition. For example, the video coder may apply a truncated unary binarization for the NSST index, where the maximum value is equal to 3 if the current intra-mode is non-angular (e.g., PLANAR or DC for chroma components, or optionally LM mode), and the maximum value is equal to 4 otherwise. In addition, the video coder may apply a block size condition. For example, the video coder may determine that if the current block is square or the width×height is less than a certain threshold (e.g., 64), the maximum value is equal to 3.
[0060] In one example, the video coder may context entropy code each bin or only certain predetermined bins (e.g., an ordinal first number of bins) from the binarization codeword. The video coder may entropy code bins other than the predetermined bins without applying context modeling (e.g., in bypass mode). If NSST is applied to luma and chroma separately, context modeling may be used for luma and chroma separately. Alternatively, bins from the binarization codeword may share context for luma and chroma, for example, the context for the first bin indicating whether the NSST index is 0 (meaning that NSST is not applied) may be shared between the luma component and the chroma component, and other bins may have separate contexts for luma and chroma.
[0061] In another example, the context modeling used for the NSST index may depend on the maximum value that the NSST index may have. For example, if the maximum value may be 3 or 4, one context set may be used to signal the NSST index for the maximum value of 3, and another context set may be used to signal the NSST index for the maximum value of 4. Similar context sets may be defined for other maximum values that the NSST index may have, and more than two maximum values may be used.
[0062] Optionally, the context for the first binary (indicating that the NSST index is equal to or not equal to 0) may be shared across all context sets, or may be shared across context sets corresponding to the same color component (e.g., for luma, chroma, or both chroma components, or all color components).
[0063] In the current JVET test model, if the PDPC index is not equal to 0, then NSST is not applied and the NSST index is not signaled. This process of avoiding NSST and not signaling the NSST index can reduce decoding complexity. However, the present disclosure recognizes that the process currently implemented in the JVET test model does not necessarily achieve the best decoding results and may not achieve the desired trade-off between decoder complexity and bit rate.
[0064] According to the techniques of this disclosure, when the NSST index of a block has a non-zero value (i.e., in other words, the NSST method is applied to the current block), a video coder (e.g., video encoder 20 or video decoder 30) need not apply and / or code (e.g., signal) a position-dependent intra-prediction combination (PDPC) syntax element for the block. This may result in similar coder complexity, but the resulting compression efficiency may be higher because the NSST method generally has better efficiency than PDPC. In this case, the PDPC index may be signaled in the bitstream at a position after the NSST index.
[0065] Additionally or alternatively, the NSST index context may be based on the PDPC index. For example, if the PDPC index is 0, one context may be used to entropy code the NSST index, and if the PDPC index is not 0, another context may be used to entropy code the NSST index. In another example, each PDPC index may have its own context used to entropy code the NSST index. Additionally or alternatively, the context of the NSST index may depend jointly on the PDPC index of the current block and other elements, such as prediction mode, block size, etc. Similarly, the context of the PDPC index may depend jointly on the NSST index of the current block and other elements, such as prediction mode, block size, etc.
[0066] Alternatively, if the NSST index is coded in the bitstream before the PDPC index, the same method may be applied. In this case, in the above method, the NSST and PDPC are swapped in the description. For example, if the NSST index is 0, one context may be used to entropy code the PDPC index, and if the NSST index is not 0, another context may be used to entropy code the PDPC index. In another example, each NSST index may have its own context used to entropy code the PDPC index. Additionally or alternatively, the context of the PDPC index may depend jointly on the NSST index of the current block and other elements, such as prediction mode, block size, etc. Similarly, the context of the NSST index may depend jointly on the PDPC index of the current block and other elements, such as prediction mode, block size, etc.
[0067] The PDPC techniques mentioned herein may be extended to any other techniques related to intra / inter prediction techniques, and / or the NSST techniques mentioned herein may be extended to any techniques related to transform techniques. The signaling of syntax elements (index / flag / mode) of prediction techniques may interact with the signaling of syntax elements (index / flag / mode) of transform techniques. The interaction may be, but is not limited to, the context of the prediction technique syntax being dependent on the context of the transform technique syntax, or vice versa.
[0068] Additionally, the video coder may be configured to apply the techniques discussed above to other coding modes, including, but not limited to, PDPC or motion parameter inheritance (MPI) modes.
[0069] NSST indices may be signaled and shared for multiple components. For example, one NSST index may be signaled and shared for a luma (Y) component, a blue hue chroma (Cb) component, and a red hue chroma (Cr) component. Alternatively, one NSST index may be signaled and shared for Cb and Cr components (a separate NSST index may be signaled for the Y component). In some examples, when one NSST index is shared for multiple components, the NSST index signaling depends on some conditions, and when these conditions are met for each of the included components, or when these conditions are met for some (but not all) of the included components, or when these conditions are met for any of the included components, the NSST index is not signaled but is derived as a default value (e.g., 0).
[0070] These conditions may include, but are not limited to: the number of non-zero coefficients (or the sum of the absolute values of the non-zero coefficients) when the block is not coded by certain coding modes, and these certain coding modes include, but are not limited to, transform skip mode and / or LM mode and / or cross-component prediction mode.
[0071] The blocks in the above examples may be blocks for each component considered independently, or they may be related blocks for some color components (e.g., related blocks for Cb and Cr), or they may be blocks for all available components (e.g., blocks for Y, Cb, and Cr). In one example, the conditions may be applied jointly to those blocks together.
[0072] For example, when the condition is applied to multiple components (e.g., Cb and Cr), the condition may include, but is not limited to, that the sum of the number of non-zero coefficients (or the sum of the absolute values of the non-zero coefficients) of each included component block is not coded by certain coding modes, and these certain coding modes include, but are not limited to, transform skip mode and / or LM mode and / or cross-component prediction mode, etc.
[0073] In some examples, when multiple NSST indices are signaled, and each NSST index is signaled for one or more components, the multiple NSST indices may be jointly binarized into one syntax element, and one binarization and / or context modeling may be applied to this jointly coded one syntax element. For example, a flag may first be coded to indicate whether there is at least one non-zero NSST index (meaning that NSST is applied to at least one component). After the flag, the multiple NSST indices are binarized into one syntax element and coded. In this example, some signaling redundancy may be removed. For example, if the flag indicates that there is at least one non-zero NSST index, then the last signaled NSST index may be inferred to be non-zero if all previous indices have values equal to 0.
[0074] In the above example, a joint NSST index signaling technique may be applied to signal the NSST index for a group of blocks. A flag may be signaled for the group to indicate that there is at least one block using a non-zero NSST index (in which case the flag is equal to 1), or that all blocks have a zero NSST index (in which case the flag is equal to 0). Signaling redundancy may also be removed for the last NSST index in the group, taking into account that the last NSST index cannot be equal to 0. In another example, if only two NSST indices (0 or 1) are possible, the last index may not be signaled if all previous indices are equal to 0, and the last NSST index may be inferred to be equal to 1. In another example, if more than two NSST index values are possible, the last index may be decremented by 1 if all previous indices are equal to 0.
[0075] The above techniques may be used in any combination.
[0076] The NSST index is used as an example. The same techniques may be applied to any transform or secondary transform index, flag, or syntax element signaling. For example, these techniques may be applied to signaling a rotational transform (ROT) index.
[0077] Likewise, the PDPC index is also used as an example. The same techniques may be applied to any intra or inter prediction index, flag, or syntax element signaling. For example, these techniques may be applied to signaling a motion parameter inheritance (MPI) index.
[0078] In some examples, video encoder 20 and / or video decoder 30 may perform transform-related syntax coding (e.g., encoding / signaling or decoding / interpretation) at special structural units, which may be referred to as signaling units (SUs). In general, a signaling unit includes a plurality of blocks. For example, a signaling unit may correspond to a single quadtree-binary tree (QTBT) of a QTBT architecture. Alternatively, a signaling unit may correspond to a group of blocks, each of which corresponds to a different respective QTBT.
[0079] In the QTBT architecture, the signaling unit can be partitioned according to a polytree, which includes a first portion partitioned according to a quadtree (where each node is partitioned into zero or four child nodes), and each leaf node of the quadtree can be further partitioned using a binary tree partition (where each node is partitioned into zero or two child nodes). Each node partitioned into zero child nodes is considered a leaf node of the corresponding tree.
[0080] As discussed above, various syntax elements (e.g., NSST index, PDPC index, prediction mode, block size, etc.) may be signaled jointly for a group of blocks. Such joint signaling may generally be described as "signaling data at a signaling unit level," where a signaling unit includes a plurality of blocks to which data is signaled at the signaling unit level, and such data applies to each block included in the signaling unit.
[0081] Problems may arise when the signaling unit forms part of a non-I slice, such as a P slice or a B slice. In these or other non-I slices, the slice may include some blocks predicted using intra mode and other blocks predicted using inter mode. However, some tools may apply to only one of intra or inter modes, but not both. Therefore, signaling some syntax at the signaling unit level for mixed blocks (intra and inter) may be inefficient, especially when the tool does not apply to a certain prediction mode.
[0082] Therefore, this disclosure also describes a variety of techniques that may be used alone or in combination with each other and / or with the techniques discussed above. Certain techniques of this disclosure may be applied to resolve a mix of inter-predicted blocks and intra-predicted blocks in non-I slices, yet still have signaling for signaling unit blocks. A video coder may use blocks arranged in a signaling unit in such a way that the signaling unit contains only blocks affected by signaling performed at the signaling unit level.
[0083] For example, the transform may be of two types: a first (or primary) transform and a secondary transform. According to the JVET model, the first transform may be a discrete cosine transform (DCT) or an enhanced multiple transform (EMT), and the secondary transform may be, for example, NSST and ROT. It should be understood that DCT, EMT, NSST, and ROT are merely examples, and the techniques of this disclosure are not limited to these transforms, but other transforms may also be used (in addition or in the alternative).
[0084] Assuming for purposes of example that an EMT flag or EMT index is signaled at the signaling unit level, then those syntax elements have values identifying which particular transform is used for the blocks included in the signaling unit. Blocks may be predicted by intra, inter, or skip modes. The signaled EMT flag or EMT index may be effective for intra-predicted blocks, but may be less effective or inefficient for inter-predicted blocks. In this case, the signaling unit may further include any or both of the following types of blocks: 1) intra-predicted blocks and skip-predicted blocks; and / or 2) inter-predicted blocks and skip-predicted blocks.
[0085] According to this example, the transform-related syntax signaled at the signaling unit level will be valid for intra-coded blocks, but the skip mode is based on the assumption that the residual is 0 and no transform is required, so the signaled transform will not affect the skipped predicted blocks and there will be no inter-coded blocks in this signaling unit block. Similarly, according to the signaling unit composition, the transform-related syntax signaled at the signaling unit level for inter-predicted blocks will be valid for inter-predicted blocks, but it will not affect the skip mode, and there will be no intra-coded blocks in this signaling unit block.
[0086] By arranging signaling units according to the techniques of this disclosure, certain syntax elements may become redundant. In the above example, it is clear that if the signaling unit type (#1 or #2) is signaled at the signaling unit level in addition to the transform syntax element, then the prediction mode is not needed. In this case, the prediction mode does not need to be signaled for each block included in the signaling unit, and the prediction mode may be inferred from the signaling unit type. In one example, the signaling unit type may be signaled as that syntax element with a context specific to a separate syntax element, or the prediction mode syntax element may be reused and signaled to indicate the signaling unit type.
[0087] As another example, the signaling unit may include blocks arranged according to either or both of the following arrangements: 1) an intra-predicted block, a skipped predicted block, and an inter-predicted block with a residual equal to 0 (a zero block); and / or 2) an inter-predicted block, a skipped predicted block, and an intra-predicted block with a zero residual.
[0088] In the first example discussed above, the coded block flag (CBF) syntax element (indicating whether a block includes a zero residue, i.e., whether the block includes one or more non-zero residual values, i.e., whether the block is “coded”) does not need to be signaled per inter-predicted blocks for signaling unit type 1, and does not need to be signaled for intra-predicted blocks for signaling unit type 2, since only zero blocks are possible.
[0089] In yet another example, the signaling unit may be constructed as follows: (1) intra-predicted blocks, skip-predicted blocks, and inter-coded blocks with a residual equal to 0 (zero blocks), and blocks coded using transform skip; and / or (2) inter-predicted blocks, skip-predicted blocks, and intra-predicted blocks with a zero residual, and blocks coded using transform skip.
[0090] Similarly, as in the above example, the CBF syntax elements need not be signaled in blocks contained in a signaling unit.
[0091] In the above example, the signaling unit blocks are classified into two types: "intra-related" type and "inter-related" type. However, it may still be possible that a mixture of intra blocks and inter blocks may share similar tool decisions, for example, the transform type may be the same for both types of predicted blocks. Then, the signaling unit types may be further expanded to the following three types: (1) intra-predicted blocks, and inter-predicted blocks with zero residual (skipped blocks, inter-blocks with zero residual, or transformed skipped inter-blocks); (2) inter-predicted blocks, and intra-blocks with zero residual, or transformed skipped intra-blocks; and (3) inter- and intra-mixing is allowed without restriction.
[0092] In this example, some redundant syntax elements may not need to be signaled per block (i.e., within each block included in the signaling unit) for signaling unit types 1 and 2, such as prediction mode or CBF syntax. Instead, video encoder 20 may encode those syntax elements once at the signaling unit level and video decoder 30 may decode those syntax elements once at the signaling unit level, and the coded values may be applied to each block included in the signaling unit.
[0093] In the above examples, EMT or first transform is used as an example. In a similar manner, secondary transforms (such as NSST or ROT) can be signaled at the signaling unit level, and redundant syntax elements (such as prediction mode or CBF syntax) can be signaled at the signaling unit level, and those elements do not need to be signaled at the block level.
[0094] Video encoder 20 and video decoder 30 may use context modeling to context code (e.g., using CABAC) transform decision related syntax elements. Transform related syntax elements may be context coded, such as flags or indices from a transform set, such as, but not limited to, an EMT flag, an NSST flag, an EMT index, an NSST index, etc. The context may be defined according to the number of non-zero transform coefficients in a block, the absolute sum of the non-zero transform coefficients, and / or the location of the non-zero transform coefficients within a TU (e.g., whether there is only one non-zero DC coefficient).
[0095] In addition, the number of non-zero coefficients may be categorized into some subgroups; for example, the number of non-zero coefficients within a certain range is one subgroup, another range of values is another subgroup, etc. Context may be defined in terms of subgroups.
[0096] In addition, the context may be defined based on the position of the last non-zero coefficient in the block, the context may be defined based on the first non-zero coefficient in the block, and / or the context may be defined based on the value or the sign (negative or positive) of the last and / or first coefficient in the block.
[0097] The number of non-zero coefficient signaling is described below. Currently, in HEVC or JVET, the position of the last non-zero coefficient and a significance map (e.g., 0—coefficient is zero, 1—coefficient is non-zero, or vice versa) are signaled for transform coefficients to indicate which coefficients are non-zero until the last non-zero coefficient.
[0098] However, if the block has only a few coefficients, the current signaling of JVET and HEVC may not be effective. For example, if the transform block has only one non-zero coefficient and that coefficient is not in the beginning of the block, the last position already indicates the position of that coefficient; however, the significance map is still signaled, which contains all zeros.
[0099] This disclosure also describes techniques related to signaling an additional syntax element having a value indicating the number of nonzero coefficients in a transform block. Video encoder 20 may signal the value of this syntax element, and video decoder 30 may decode the value of this syntax element to determine the number of nonzero transform coefficients in the transform block. This syntax element value may be signaled using any binarization (e.g., unary code, truncated unary code, Golomb code, exponential Golomb code, Rice code, fixed length binary code, truncated binary code, etc.). For truncated binarization, the maximum element may be the number of possible coefficients up to the last position coefficient.
[0100] In one example, this new syntax element may be signaled after the last non-zero coefficient number for the transform block. In another example, this new syntax element may be signaled before the last non-zero coefficient. In the latter case, the flag may indicate whether the block has only one DC coefficient.
[0101] Since the last non-zero coefficient and the number of non-zero coefficients are signaled, the techniques of this disclosure may result in a reduction in the size of a coded significance map that forms part of a bitstream. For example, when signaling a significance map, the number of non-zero coefficients that have been signaled may be counted; when a number of non-zero coefficients equal to the signaled number of non-zero coefficients minus 1 has been signaled, there is no need to continue signaling the significance map for the block, since the only possible next non-zero coefficient is the last coefficient in the block.
[0102] In one example, the syntax element mentioned above may be a flag (a coefficient flag) indicating whether the transform block has only one non-zero coefficient. This flag may be signaled after the position of the last non-zero coefficient and may also depend on that position. For example, if the last non-zero coefficient is the first coefficient in the block (DC), then it is already known that only one coefficient is possible and a coefficient flag is not needed. Similarly, a flag may be signaled only for the case when the position of the last non-zero coefficient is greater than a certain threshold. For example, if the last non-zero coefficient is numerically set to a certain distance from the beginning of the block, then a coefficient flag is signaled.
[0103] The context model selection for a coefficient flag may depend on the position of the last nonzero coefficient in the block, the distance of that last position from the beginning of the block, the last nonzero coefficient value, and / or the sign of that value, alone or in any combination.
[0104] A coefficient flag may be signaled after the position of the last non-zero coefficient, in another alternative after the position of the last non-zero coefficient and its value, in yet another alternative after the position of the last non-zero coefficient, its value and sign. This may depend on which context model is applied (see above).
[0105] In yet another example, a coefficient flag may be signaled before the last non-zero coefficient digital position and may indicate whether the block has only one DC (first transform coefficient) coefficient. In this example, the last non-zero coefficient digital position may depend on that flag and be signaled when the flag has a value representing "disabled", which means that there are more than one non-zero coefficient or one coefficient is not a DC coefficient. In addition, the last position signaling may be modified by subtracting 1 from the position coordinates, since if a coefficient flag is disabled, then the last position equal to the DC coefficient cannot be signaled; otherwise, that flag would be enabled.
[0106] When this one coefficient flag is signaled and has a value indicating "enabled" (i.e., the block has only one non-zero coefficient), a significance map may not be needed, and only the position of the last coefficient and its value and sign may be signaled. Thus, video encoder 20 may signal only the position of the last coefficient, and video decoder 30 may receive only data indicating the position of the last coefficient and determine that subsequent data of the bitstream applies to a different set of syntax elements (e.g., syntax elements for the same block, but not associated with transform coefficient data, or syntax elements for a subsequent block).
[0107] A coefficient flag may be conditionally signaled based on which transform type (e.g., DCT or EMT) is used, and may depend on an EMT flag or an EMT index. Additionally, a coefficient flag signaling may depend on whether a secondary transform (e.g., NSST or ROT) is used in a block; secondary transform syntax, such as an NSST flag, NSST index, ROT flag, or ROT index; etc. For example, if a secondary transform is used, then the flag may not be signaled.
[0108] The more detailed example described for one non-zero coefficient flag is applicable to the case when more than one non-zero coefficient value is signaled in a block.
[0109] Video encoder 20 and video decoder 30 may switch between different transform types based on non-zero coefficients. Two different types of transforms may be used, for example, one type being a separable transform and the other type being a non-separable transform. For the use of each type of transform, some restrictions may be added, namely, only non-zero coefficients may be present for certain positions inside the transform unit. In this way, the selected type of transform is not signaled, but the video decoder 30 may derive the selected type of transform based on the positions of the non-zero coefficients inside the transform unit after decoding the coefficients. By deriving the transform type rather than receiving explicit signaling, the size of the encoded video bitstream may be reduced, which may thereby improve bitstream efficiency without introducing too much complexity into the video decoder 30 and without losing the quality of the resulting decoded video data. In addition, providing multiple types of transforms in this manner may result in even further improvements in bitstream efficiency, in that the resulting transform type may, on average, better compress the residual data.
[0110] In one example, if there is at least one non-zero coefficient after the Nth coefficient in the scan order (where N may be predefined or derived based on some conditions), then a separable transform is applied; otherwise (all non-zero coefficients are present only in the first N coefficients in the scan order), a non-separable transform is applied.
[0111] In another example, the type of transform is still signaled by a flag / index, but the context model used for entropy coding (entropy encoding or entropy decoding) coefficients at different positions may depend on the value of the signaled flag / index.
[0112] In another example, a flag or index to indicate the transform selection mentioned above is signaled after the Nth coefficient or all coefficients. The flag or index may be context coded, where the context depends on the position of the last non-zero coefficient. For example, the context may depend on whether the last non-zero coefficient occurs before or after the Nth coefficient. If the last non-zero coefficient stops at the Nth coefficient itself, then before or after the earlier mentioned Nth coefficient, the context model may be associated with either group, or a separate context may be assigned.
[0113] Video encoder 20 may encode / signal syntax elements for a signaling unit, and video decoder 30 may decode and interpret the values of the syntax elements for the signaling unit. As described earlier, syntax elements may be signaled at the signaling unit level. However, some syntax elements may not be applicable to every block included in a signaling unit.
[0114] For example, a secondary transform (e.g., NSST) may be applied only to intra-predicted blocks, which have non-zero coefficients. It may be the case that there are no blocks in the signaling unit to which the secondary transform will be applied. For these cases, signaling NSST information (e.g., NSST index or NSST flag) for this signaling unit is not needed and may just be a waste of bits. In another example, a first transform (e.g., EMT) is applied to non-zero residual blocks. It may also be the case that all blocks included in the signaling unit have zero residuals, and signaling EMT information (e.g., EMT flag or EMT index) is not needed for this signaling unit and may just be a waste of bits.
[0115] In some examples, video encoder 20 may defer signaling of a signaling unit syntax until the first block contained in a signaling unit to which such signaling applies. In other words, signaling unit syntax is not signaled for blocks at the beginning of a signaling unit in scan order for which such signaling does not apply. Likewise, video decoder 30 will only apply the value of a signaling unit syntax element to blocks following the signaling unit syntax element in a signaling unit.
[0116] For example, video encoder 20 may not signal some type of information that applies to all blocks within a signaling unit until there are blocks in the signaling unit to which the information applies. Similarly, video decoder 30 may not analyze some type of information that applies to all blocks within a signaling unit until there are blocks in the signaling unit to which the information applies. The information may be information that identifies a specific coding tool, syntax element, etc.
[0117] As an example, video encoder 20 may signal and video decoder 30 may receive NSST information (indices, flags, etc.) in the first intra-block with non-zero residual in a signaling unit. In another example, video encoder 20 may signal and video decoder 30 may receive EMT information (indices, flags, etc.) at the first non-zero block in a signaling unit. These blocks may not necessarily be at the beginning of the corresponding signaling unit. In some examples, once a syntax element (e.g., information for a coding tool or other type of syntax element) is signaled for the first block using the syntax element, that information may be uniform for all blocks using the syntax element after that first block in block scan order. However, this should not be considered a requirement in all cases.
[0118] By deferring the signaling of these syntax elements, bits associated with the syntax elements may be saved if there are no blocks in the signaling unit that require these syntax elements or if there are no blocks in the signaling unit to which such signaling is applicable, compared to a signaling and reception technique that always signals syntax elements at the signaling unit level regardless of whether the signaling unit contains any blocks to which the signaling unit syntax elements will be applicable.
[0119] Video encoder 20 may utilize similar techniques to defer signaling of other syntax elements (not necessarily transform related) at the signaling unit level, depending on the signaled information and the block types to which such information included in the signaling unit applies. The above examples of deferring signaling and analyzing information of a signaling unit should not be considered limiting.
[0120] Various syntax elements may be considered specific to a signaling unit. Some syntax elements may be introduced only for a signaling unit and may not be present for other blocks. For example, these syntax elements may be control flags and coding mode related parameters. In one example, the signaling unit syntax elements include any or all of the first transform (e.g., EMT) and / or second transform syntax elements (e.g., NSST or ROT flags and / or indices) as mentioned earlier, and these syntax elements need not be present for blocks larger than a signaling unit or not included in a signaling unit.
[0121] Alternatively or additionally, existing syntax elements for blocks signaled for a signaling unit may have different range values or different semantics / interpretation than the same syntax elements signaled for blocks larger than or not included in the signaling unit. In one example, the non-zero coefficient thresholds for identifying when to signal first transform and second transform syntax elements may be different for a signaling unit than for other blocks. These thresholds may be greater or less than corresponding thresholds for other blocks.
[0122] For example, a secondary transform (e.g., NSST or ROT) index and / or flag may be signaled for a block having at least one non-zero transform coefficient in a signaling unit, and if a block larger than the signaling unit or not included in the signaling unit has at least two non-zero coefficients, a secondary transform index may be signaled for the block. When the secondary transform index is not signaled, video decoder 30 infers the value of the secondary transform index to be, for example, equal to a default value (e.g., 0). The same technique may be applied to the first transform or any other transform.
[0123] These signaling unit specific parameters may also differ depending on the slice type and / or frequency block to which the signaling unit belongs. For example, I slices, P slices, and B slices may have different signaling unit parameters, different ranges of values, or different semantics / interpretations.
[0124] The signaling unit parameters described above are not limited to transforms but can be used with or introduced into any decoding mode.
[0125] Video encoder 20 may further send syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, to video decoder 30, e.g., in a picture header, a block header, a slice header, or other syntax data, such as a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS).
[0126] Where applicable, video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware, or any combination thereof. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined video encoder / decoder (codec). A device including video encoder 20 and / or video decoder 30 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.
[0127] Figure 21 is a block diagram illustrating an example of a video encoder 20 that may implement techniques for binarizing secondary transform indices. Video encoder 20 may perform intra-coding and inter-coding of video blocks within a video slice. Intra-coding relies on spatial prediction to reduce or remove video spatial redundancy within a given video frame or picture. Inter-coding relies on temporal prediction to reduce or remove video temporal redundancy within adjacent frames or pictures of a video sequence. Intra-mode (I-mode) may refer to any of several spatial-based coding modes. Inter-mode (e.g., unidirectional prediction (P-mode) or bidirectional prediction (B-mode)) may refer to any of several temporal-based coding modes.
[0128] like Figure 2 As shown, video encoder 20 receives a current video block within a video frame to be encoded. Figure 2 In the example of , video encoder 20 includes mode select unit 40, reference picture memory 64 (which may also be referred to as a decoded picture buffer (DPB)), summer 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Mode select unit 40, in turn, includes motion compensation unit 44, motion estimation unit 42, intra-prediction unit 46, and partition unit 48. For video block reconstruction, video encoder 20 also includes inverse quantization unit 58, inverse transform unit 60, and summer 62. Deblocking filter ( Figure 2 62 to filter the output of summer 50 (not shown in the figure) to filter block boundaries to remove blocking artifacts from the reconstructed video. If desired, a deblocking filter will typically filter the output of summer 62. In addition to the deblocking filter, additional filters (in-loop or post-loop) may also be used. These filters are not shown for brevity, but if desired, these filters may filter the output of summer 50 (as an in-loop filter).
[0129] During the encoding process, video encoder 20 receives a video frame or slice to be decoded. The frame or slice may be divided into multiple video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-frame predictive encoding of the received video block relative to one or more blocks in one or more reference frames to provide temporal prediction. Intra-frame prediction unit 46 may alternatively perform intra-frame predictive encoding of the received video block relative to one or more neighboring blocks in the same frame or slice as the block to be decoded to provide spatial prediction. Video encoder 20 may perform multiple decoding passes, for example, to select an appropriate decoding mode for each block of video data.
[0130] Furthermore, partition unit 48 may partition blocks of video data into sub-blocks based on evaluation of previous partitioning schemes in previous coding passes. For example, partition unit 48 may initially partition a frame or slice into CTUs, and partition each of the CTUs into sub-CUs based on rate-distortion analysis (e.g., rate-distortion optimization). Mode select unit 40 may further generate a quadtree data structure indicating the partitioning of the CTU into sub-CUs. A leaf node CU of the quadtree may include one or more PUs and one or more TUs.
[0131] Mode select unit 40 may select one of the prediction modes (intra or inter) (e.g., based on the error result) and provide the resulting predicted block to summer 50 to generate residual data and to summer 62 to reconstruct the encoded block for use as a reference frame. Mode select unit 40 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy encoding unit 56.
[0132] Motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are depicted separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors, which estimate the motion of video blocks. For example, a motion vector may indicate the displacement of a PU of a video block within a current video frame or picture relative to a predictive block within a reference frame (or other coded unit) relative to the current block being coded within the current frame (or other coded unit). A predictive block is a block that is found to closely match the block to be coded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. In some examples, video encoder 20 may calculate values for sub-integer pixel positions of a reference picture stored in reference picture memory 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference picture. Thus, motion estimation unit 42 may perform motion searches relative to full-pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.
[0133] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-coded slice by comparing the position of the PU to the position of a predictive block of a reference picture. The reference picture may be selected from a first reference picture list (List 0) or a second reference picture list (List 1), each of which identifies one or more reference pictures stored in reference picture memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.
[0134] Motion compensation performed by motion compensation unit 44 may involve extracting or generating a predictive block based on a motion vector determined by motion estimation unit 42. Again, in some examples, motion estimation unit 42 and motion compensation unit 44 may be functionally integrated. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may find the location of the predictive block to which the motion vector points in one of the reference picture lists. Summer 50 forms a residual video block by subtracting the pixel values of the predictive block from the pixel values of the current video block being coded, forming pixel difference values, as discussed below. In general, motion estimation unit 42 performs motion estimation relative to a luma component, and motion compensation unit 44 uses the motion vector calculated based on the luma component for both chroma components and luma components. Mode select unit 40 may also generate syntax elements associated with video blocks and video slices for use by video decoder 30 to decode video blocks of video slices.
[0135] As described above, intra-prediction unit 46 may intra-predict a current block, as an alternative to the inter-prediction performed by motion estimation unit 42 and motion compensation unit 44. In particular, intra-prediction unit 46 may determine an intra-prediction mode to use to encode the current block. In some examples, intra-prediction unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction unit 46 (or mode select unit 40 in some examples) may select an appropriate intra-prediction mode to use from the tested modes.
[0136] For example, intra-prediction unit 46 may calculate rate-distortion values using a rate-distortion analysis for various tested intra-prediction modes, and select the intra-prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis typically determines the amount of distortion (or error) between an encoded block and an original, unencoded block (which was encoded to produce the encoded block), as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra-prediction unit 46 may calculate ratios from the distortions and rates for the various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0137] After selecting the intra-prediction mode for the block, intra-prediction unit 46 may provide information indicating the selected intra-prediction for the block to entropy encoding unit 56. Entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode. Video encoder 20 may include the following in the transmitted bitstream: configuration data, which may include a plurality of intra-prediction mode index tables and a plurality of modified intra-prediction mode index tables (also referred to as codeword mapping tables); definitions of contexts for encoding various blocks; and indications of the most probable intra-prediction mode, the intra-prediction mode index tables, and the modified intra-prediction mode index tables to be used for each of the contexts.
[0138] Video encoder 20 forms a residual video block by subtracting the prediction data from mode select unit 40 from the original video block being coded. Summer 50 represents a component that performs this subtraction. Transform processing unit 52 applies a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to the residual block, producing a video block comprising transform coefficient values. A wavelet transform, an integer transform, a sub-band transform, a discrete sine transform (DST), or other type of transform may be used instead of a DCT. In any case, transform processing unit 52 applies a transform to the residual block, producing a block of transform coefficients. The transform may convert the residual information from the pixel domain to a transform domain, such as the frequency domain.
[0139] In addition, in some examples, e.g., when the block is intra predicted, transform processing unit 52 may apply a secondary transform, such as a non-separable secondary transform (NSST), to the transform coefficients produced by the first transform. Transform processing unit 52 may also pass one or more values of the secondary transform syntax elements for the block to entropy encoding unit 56 for entropy encoding. Entropy encoding unit 56 may entropy encode the following according to the techniques of this disclosure: Figure 3 These and / or other syntax elements (eg, secondary transform syntax elements or other signaling unit syntax elements) are discussed in greater detail.
[0140] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter.
[0141] After quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients (and any corresponding values for related syntax elements, such as secondary transform syntax elements, signaling unit syntax elements, coding tool syntax elements, enhanced multiple transform (EMT) syntax elements, etc.). For example, entropy encoding unit 56 may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. In the case of context-based entropy coding, the context may be based on neighboring blocks. After entropy coding by entropy encoding unit 56, the encoded bitstream may be transmitted to another device (e.g., video decoder 30) or archived for later transmission or retrieval.
[0142] According to the techniques of this disclosure, video encoder 20 may encode certain syntax elements at a signaling unit level. A signaling unit typically includes syntax elements for two or more blocks of video data, such as coding tree blocks (CTBs) or coding units (CUs). For example, the blocks may correspond to different branches / nodes of a common QTBT structure, or to distinct QTBT structures.
[0143] As discussed above, in one example, video encoder 20 may defer signaling syntax elements for a signaling unit until video encoder 20 encounters blocks to which those signaling unit syntax elements are associated. In this manner, if a signaling unit ultimately does not include any blocks to which the signaling unit syntax elements are associated, video encoder 20 may avoid encoding the signaling unit syntax elements altogether. If a signaling unit does contain blocks to which the signaling unit syntax elements are associated, video encoder 20 may encode these syntax elements to form a portion of the bitstream that follows the blocks to which the signaling unit syntax elements are not associated and precedes the blocks to which the signaling unit syntax elements are associated in encoding / decoding order. The signaling unit syntax elements may include any or all of NSST information (NSST flag and / or index), EMT information (EMT flag and / or index), and the like.
[0144] For example, mode select unit 40 may determine whether an intra-predicted block results in a zero or non-zero residual (as calculated by summer 50). Mode select unit 40 may await determination of a signaling unit syntax element for a signaling unit until an intra-predicted block having a non-zero residual (i.e., a residual block having at least one non-zero coefficient) has been encoded. After identifying an intra-predicted block having a non-zero residual, mode select unit 40 may determine one or more signaling unit syntax elements to be encoded for the signaling unit including the intra-predicted block, and further, entropy encoding unit 56 may entropy encode a value of the signaling unit syntax element at a location after other blocks of the signaling unit but before the intra-predicted block of the signaling unit in encoding / decoding order.
[0145] Inverse quantization unit 58 and inverse transform unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain. Specifically, summer 62 adds the reconstructed residual block to the motion compensated prediction block earlier generated by motion compensation unit 44 or intra-prediction unit 46 to produce a reconstructed video block for storage in reference picture memory 64. The reconstructed video block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block to inter-code a block in a subsequent video frame.
[0146] Figure 220 represents an example of a video encoder that may be configured to: determine a maximum value of a secondary transform (e.g., a non-separable secondary transform (NSST)) syntax element for a block of video data; and binarize a value of the secondary transform (e.g., NSST) syntax element based on the determined maximum value. Video encoder 20 may further entropy encode the value of the secondary transform (e.g., NSST) syntax element.
[0147] Figure 3 1 is a block diagram of an example entropy encoding unit 56 that may be configured to perform CABAC according to techniques of this disclosure. Entropy encoding unit 56 initially receives syntax element 118. If syntax element 118 is already a binary-valued syntax element, the binarization step may be skipped. If syntax element 118 is a non-binary-valued syntax element, binarizer 120 binarizes the syntax element.
[0148] Binarizer 120 performs a mapping of non-binary values to a series of binary decisions. These binary decisions may be referred to as "bins." For example, for a transform coefficient level, the value of the level may be decomposed into consecutive bins, each bin indicating whether the absolute value of the coefficient level is greater than a certain value. For example, for a transform coefficient, a binary 0 (sometimes referred to as a significance flag) indicates whether the absolute value of the transform coefficient level is greater than 0; a binary 1 indicates whether the absolute value of the transform coefficient level is greater than 1; and so on. A unique mapping may be generated for each non-binary valued syntax element.
[0149] Binarizer 120 passes each bin to the binary arithmetic encoding side of entropy encoding unit 56. That is, for a set of predetermined non-binary valued syntax elements, each bin type (e.g., bin 0) is encoded before the next bin type (e.g., bin 1). According to the techniques of this disclosure, when binarizing the values of secondary transform syntax elements (e.g., non-separable secondary transform (NSST) syntax elements) for an intra-predicted block of video data, binarizer 120 may determine a maximum possible value for the secondary transform (e.g., NSST) syntax element for the block, e.g., based on the intra-prediction mode used to predict the block and / or other parameters (e.g., the size of the block).
[0150] In one example, if the intra prediction mode for the block is DC, planar, or LM mode for chroma components, the binarizer 120 determines that the maximum possible value of the NSST index is equal to 3, and otherwise the maximum possible value of the NSST index is equal to 4. The binarizer 120 then binarizes the actual value of the NSST index based on the determined maximum possible value using a common binarization technique regardless of the determined maximum possible value (e.g., using truncated unary binarization regardless of whether the determined maximum possible value of the NSST index is 3 or 4).
[0151] Entropy coding may be performed in a normal mode or a bypass mode. In bypass mode, the bypass coding engine 126 performs arithmetic coding using a fixed probability model (eg, using Golomb-Rice or exponential Golomb coding). The bypass mode is typically used for more predictable syntax elements.
[0152] Entropy coding in normal mode CABAC involves performing context-based binary arithmetic coding. Normal mode CABAC is typically performed to predict the value of a bin given the value of a previously coded bin. The context modeler 122 determines the probability that the bin is a least likely symbol (LPS). The context modeler 122 outputs the bin value and a context model (e.g., a probability state σ) to the normal encoding engine 124. The context model may be an initial context model for a series of bins, or the context modeler 122 may determine the context model based on the coded values of previously coded bins. The context modeler 122 may update the context state based on whether the previously coded bin was an MPS or an LPS.
[0153]
[0063] In accordance with the techniques of this disclosure, context modeler 122 may be configured to determine a context model for entropy encoding a secondary transform syntax element (eg, a NSST syntax element) based on the determined maximum possible value of the secondary transform syntax element discussed above.
[0154] After the context modeler 122 determines the context model and the probability state σ, the conventional encoding engine 124 performs BAC on the binary values using the context model. Alternatively, in bypass mode, the bypass encoding engine 126 bypass encodes the binary values from the binarizer 120. In either case, the entropy encoding unit 56 outputs an entropy encoded bitstream including entropy encoded data.
[0155] In this way, Figure 1 and 2 Video encoder 20 (and about Figure 3 The entropy encoding unit 56 described therein represents an example of a video encoder that includes a memory configured to store video data; and one or more processors implemented in a circuit system and configured to: transform intermediate transform coefficients of a block of video data using a secondary transform; determine a maximum possible value for a secondary transform syntax element for the block, the value of the secondary transform syntax element representing the secondary transform; binarize the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value; and entropy encode the binarized values of the secondary transform syntax elements of the block to form binarized values representing the secondary transform for the block.
[0156] Figure 4 is a block diagram illustrating an example of a video decoder 30 that may implement techniques for binarizing secondary transform indices. Figure 4 In an example of , video decoder 30 includes entropy decoding unit 70, motion compensation unit 72, intra prediction unit 74, inverse quantization unit 76, inverse transform unit 78, reference picture memory 82, and summer 80. In some examples, video decoder 30 may perform the same operations as those performed with respect to video encoder 20 ( Figure 2 ) The encoding pass described is roughly the inverse of the decoding pass.
[0157] In some examples, entropy decoding unit 70 decodes certain syntax elements of a signaling unit. For example, video decoder 30 may determine that two or more blocks of video data correspond to a common signaling unit. Entropy decoding unit 70 may entropy decode syntax elements for a signaling unit according to the techniques of this disclosure. For example, entropy decoding unit 70 may entropy decode secondary transform syntax elements (e.g., non-separable secondary transform (NSST) index and / or flag), enhanced multiple transform (EMT) syntax elements (e.g., EMT index and / or flag), and so on. Entropy decoding unit 70 may entropy decode signaling unit syntax elements that follow one or more blocks of the signaling unit but precede one or more other blocks of the signaling unit, and apply the value of the signaling unit syntax element only to blocks that follow the syntax element in decoding order.
[0158] Furthermore, video decoder 30 may infer certain data from the presence of syntax elements, such as that the block immediately following these signaling unit syntax elements is inter-predicted and has a non-zero residual. Thus, the video decoder may determine that the relevant block-level syntax elements (e.g., indicating that the block is intra-predicted and the block is coded (i.e., has a non-zero residual value)) are not present in the bitstream, and thereby determine that subsequent data of the bitstream applies to other syntax elements.
[0159] In addition, entropy decoding unit 70 may be as follows with respect to Figure 5 For example, according to the techniques of this disclosure, entropy decoding unit 70 may inversely binarize secondary transform syntax element values using a common binarization scheme (eg, truncated unary binarization) regardless of the maximum possible value of the secondary transform syntax element values.
[0160] Motion compensation unit 72 may generate prediction data based on the motion vectors received from entropy decoding unit 70 , while intra-prediction unit 74 may generate prediction data based on the intra-prediction mode indicator received from entropy decoding unit 70 .
[0161] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements from video encoder 20. Entropy decoding unit 70 of video decoder 30 entropy decodes the bitstream to produce quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 70 transfers the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0162] When the video slice is coded as an intra-coded (I) slice, intra-prediction unit 74 may generate prediction data for a video block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is coded as an inter-coded (i.e., B or P) slice, motion compensation unit 72 generates predictive blocks for the video blocks of the current video slice (assuming the video blocks are inter-predicted) based on motion vectors and other syntax elements received from entropy decoding unit 70. The inter-predictive blocks may be generated from one of the reference pictures within one of the reference picture lists. Video decoder 30 may construct reference frame lists: List 0 and List 1 using a default construction technique based on reference pictures stored in reference picture memory 82. Blocks of P and B slices may also be intra-predicted.
[0163] Motion compensation unit 72 determines prediction information for a video block of the current video slice by analyzing motion vectors and other syntax elements, and uses the prediction information to produce a predictive block for the current video block being decoded. For example, motion compensation unit 72 uses some of the received syntax elements to determine a prediction mode (e.g., intra- or inter-prediction) used to code a video block of the video slice, an inter-prediction slice type (e.g., B slice or P slice), construction information for one or more of the reference picture lists for the slice, motion vectors for each inter-coded video block of the slice, inter-prediction status for each inter-coded video block of the slice, and other information used to decode a video block in the current video slice.
[0164] Motion compensation unit 72 may also perform interpolation based on interpolation filters. Motion compensation unit 72 may use interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks. In this case, motion compensation unit 72 may determine the interpolation filters used by video encoder 20 from the received syntax elements and use the interpolation filters to produce the predictive blocks.
[0165] Inverse quantization unit 76 inverse quantizes (ie, dequantizes) the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 70. The inverse quantization process may include using a quantization parameter QP calculated by video decoder 30 for each video block in a video slice. Y to determine the degree of quantization that should be applied and likewise the degree of inverse quantization.
[0166] Inverse transform unit 78 applies an inverse transform (eg, an inverse DCT, an inverse integer transform or a conceptually similar inverse transform process) to the transform coefficients to produce a residual block in the pixel domain.
[0167] After motion compensation unit 72 generates a predictive block for the current video block based on the motion vector and other syntax elements, video decoder 30 forms a decoded video block by summing the residual block from inverse transform unit 78 with the corresponding predictive block generated by motion compensation unit 72. Summer 80 represents a component that performs this summing operation. If desired, a deblocking filter may also be applied to filter the decoded blocks in order to remove blocking artifacts. Other loop filters (either in the decoding loop or after the decoding loop) may also be used to smooth pixel transitions, or otherwise improve video quality. The decoded video blocks in a given frame or picture are then stored in reference picture memory 82, which stores reference pictures used for subsequent motion compensation. Reference picture memory 82 also stores the decoded video for later presentation to a display device (e.g., Figure 1 on a display device 32).
[0168] Figure 4 30 represents an example of a video decoder that may be configured to: determine a maximum value of a secondary transform (e.g., a non-separable secondary transform (NSST)) syntax element for a block of video data; and binarize the value of the NSST syntax element based on the determined maximum value. Video decoder 30 may further entropy decode the value of the NSST syntax element.
[0169] Figure 5 is a block diagram of an example entropy decoding unit 70 that may be configured to perform CABAC in accordance with the techniques of this disclosure. Figure 5 The entropy decoding unit 70 is Figure 5 CABAC is performed in a manner reciprocal to that described for entropy encoding unit 56. Entropy decoding unit 70 receives entropy encoded bits from bitstream 218. Entropy decoding unit 70 provides the entropy encoded bits to context modeler 220 or bypass decoding engine 222 based on whether the entropy encoded bits were entropy encoded using bypass mode or normal mode. If the entropy encoded bits were entropy encoded in bypass mode, then bypass decoding engine 222 uses bypass decoding (e.g., Golomb-Rice or exponential Golomb decoding) to entropy decode the entropy encoded bits.
[0170] If the entropy coded bits are entropy coded in a conventional mode, the context modeler 220 may determine a probability model for the entropy coded bits and the conventional decoding engine 224 may entropy decode the entropy coded bits to produce bins for a non-binary valued syntax element (or the syntax element itself in the case of a binary value).
[0171] Context modeler 220 may determine context models and probability states for certain syntax elements, such as secondary transform syntax elements and / or enhanced multiple transform (EMT) syntax elements (e.g., NSST index, NSST flag, EMT index, EMT flag, etc.), using the techniques of this disclosure. For example, context modeler 220 may determine the context model based on a determined maximum possible value for the NSST syntax element. Entropy decoding unit 70 may determine the maximum possible value for the NSST syntax element based on, for example, an intra-prediction mode for a block to which the NSST syntax element corresponds and / or a size of the block.
[0172] After the context modeler 220 determines the context model and the probability state σ, the conventional decoding engine 224 performs binary arithmetic decoding on the binary values based on the determined context model.
[0173] After the conventional decoding engine 224 or the bypass decoding engine 222 entropy decodes the binary, the inverse binarizer 230 may perform a reverse mapping to convert the binary back to a value for a non-binary valued syntax element. According to the techniques of this disclosure, the inverse binarizer 230 may inverse binarize secondary transform syntax element values (e.g., NSST, ROT, and / or EMT values) using a common binarization scheme (e.g., truncated unary binarization) regardless of the maximum possible value of the secondary transform syntax element value.
[0174] For example, when inverse binarizing the value of a secondary transform syntax element (e.g., a non-separable secondary transform (NSST) syntax element) of an intra-predicted block of video data, the inverse binarizer 230 may determine a maximum possible value for the secondary transform (e.g., NSST) syntax element for the block, e.g., based on the intra-prediction mode used to predict the block and / or other parameters (e.g., the size of the block).
[0175] In one example, if the intra prediction mode for the block is DC, planar, or LM mode for chroma components, the inverse binarizer 230 determines that the maximum possible value of the NSST index is equal to 3, and otherwise the maximum possible value of the NSST index is equal to 4. The inverse binarizer 230 then inversely binarizes the actual value of the NSST index from the entropy decoded binary string using a common binarization technique regardless of the determined maximum possible value based on the determined maximum possible value (e.g., using truncated unary inverse binarization regardless of whether the determined maximum possible value of the NSST index is 3 or 4).
[0176] In this way, Figure 1 and 4 The video decoder 30 (including Figure 5 The entropy decoding unit 70 described represents an example of a video decoder that includes a memory configured to store video data; and one or more processors implemented in a circuit system and configured to: determine a maximum possible value for a secondary transform syntax element for a block of video data; entropy decode the value of the secondary transform syntax element for the block to form a binarized value representing a secondary transform for the block; inversely binarize the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value to determine a secondary transform for the block; and inverse transform transform coefficients for the block using the determined secondary transform.
[0177] Figure 6 1 is a flowchart of an example method of encoding video data according to the techniques of the present invention. Figure 1 , 2 3 and the video encoder 20 and its components are explained Figure 6 However, it should be understood that in other examples, other video encoding devices may perform this method or similar methods consistent with the techniques of this disclosure.
[0178] Initially, video encoder 20 receives a block to be encoded (250). In this example, assume that mode select unit 40 of video encoder 20 determines to intra-predict the block (252). Although Figure 6 , but this decision may include predicting the block using various prediction modes (including intra- or inter-prediction modes), and ultimately determining that the block is to be intra-predicted using a particular intra-prediction mode (e.g., an angular mode or a non-angular mode, such as DC, planar, or LM mode). Intra-prediction unit 46 of video encoder 20 then intra-predicts the block using the intra-prediction mode, producing a predicted block.
[0179] Summer 50 then calculates a residual block (254). Specifically, summer 50 calculates the pixel-by-pixel differences between the original block and the predicted block to calculate a residual block, where each value (sample) of the residual block represents a corresponding pixel difference.
[0180] Transform processing unit 52 then transforms the residual block (256) using a first transform (e.g., DCT or EMT) to produce intermediate transform coefficients. In this example, transform processing unit 52 also applies a secondary transform (e.g., NSST or ROT) to the intermediate transform coefficients resulting from the first transform (258). In some examples, transform processing unit 52 may select a secondary transform from a plurality of available secondary transforms. Thus, transform processing unit 52 may generate values for one or more secondary transform syntax elements (e.g., NSST flag, NSST index, ROT flag, ROT index, EMT flag, and / or EMT index) and provide these syntax element values to entropy encoding unit 56.
[0181] Quantization unit 54 quantizes the final transform coefficients resulting from the secondary (or any subsequent) transform, and entropy encoding unit 56 entropy encodes the quantized transform coefficients (260), along with other syntax elements for the block (e.g., syntax elements representing prediction modes, partition syntax elements representing the size of the block, etc.). In some examples, entropy encoding unit 56 also entropy encodes signaling unit syntax elements that include signaling units for the block. If the block is the first block to which these signaling unit syntax elements are applied, entropy encoding unit 56 may encode the signaling unit syntax elements and output the entropy encoded signaling unit syntax elements before outputting other block-based syntax elements for the block, as discussed above.
[0182] Entropy encoding unit 56 also entropy encodes secondary transform syntax as discussed above. Specifically, binarizer 120 binarizes secondary transform syntax elements according to the techniques of this disclosure (264). For example, binarizer 120 may perform a particular binarization scheme (e.g., truncated unary binarization) regardless of the maximum possible value of the secondary transform syntax elements.
[0183] The binarizer 120 may determine the maximum possible value of the secondary transform syntax element based on, for example, the intra prediction mode used to intra-predict the block, as discussed above. For example, if the intra prediction mode is a non-angular mode, the binarizer 120 may determine that the maximum possible value of the secondary transform syntax element is 3, but if the intra prediction mode is an angular mode, the binarizer 120 may determine that the maximum possible value of the secondary transform syntax element is 4. Although this determination may be used during binarization, in some examples, this determination does not affect the actual binarization scheme (e.g., truncated unary binarization) that the binarizer 120 performs to binarize the secondary transform syntax element values.
[0184] After binarization, context modeler 122 may determine a context to entropy encode the secondary transform syntax element (266). In some examples, context modeler 122 selects a context based on a maximum possible value of the secondary transform syntax element determined as discussed above. Conventional encoding engine 124 may then entropy encode the binarized value of the secondary transform syntax element using the determined context (268).
[0185] In this way, Figure 6 The method of the present invention represents an example of a method of encoding video data, the method comprising: transforming intermediate transform coefficients of a block of video data using a secondary transform; determining a maximum possible value of a secondary transform syntax element for the block, the value of the secondary transform syntax element representing the secondary transform; binarizing the value of the secondary transform syntax element using a common binarization scheme regardless of the maximum possible value; and entropy encoding the binarized value of the secondary transform syntax element of the block to form a binarized value representing the secondary transform for the block.
[0186] Figure 7 1 is a flowchart of an example of a method of decoding video data according to the techniques of the present invention. Figure 1 , 4 5 and the video decoder 30 and its components are explained Figure 7 However, it should be understood that in other examples, other video encoding devices may perform this method or similar methods consistent with the techniques of this disclosure.
[0187] Initially, entropy decoding unit 70 entropy decodes prediction information and quantized transform coefficients for a block of video data (280). In accordance with the techniques of this disclosure, entropy decoding unit 70 also entropy decodes a secondary transform syntax element for the block. Specifically, context modeler 220 determines a context to entropy decode the secondary transform syntax element (282). Context modeler 220 may determine the context based on a maximum possible value of the secondary transform syntax element. For example, if the intra-prediction mode is a non-angular mode (e.g., DC, planar, or LM mode), context modeler 220 may determine that the maximum possible value of the secondary transform syntax element is 3, but otherwise, if the intra-prediction mode is an angular mode, context modeler 220 may determine that the maximum possible value is 4. Context modeler 220 may then determine the context from the maximum possible value of the secondary transform syntax element. Conventional decoding engine 224 may then use the determined context to entropy decode data for the secondary transform syntax element (284).
[0188] Debinarizer 230 may then debinarize the entropy decoded data for the secondary transform syntax element (286) to generate a value for the secondary transform syntax element. This value may indicate, for example, whether a secondary transform is to be applied (e.g., an NSST flag or a ROT flag), and if so, which of a plurality of secondary transforms is to be applied (e.g., an NSST index or a ROT index).
[0189] Inverse quantization unit 76 may then inverse quantize the entropy decoded coefficients for the block (288).Inverse transform unit 78 may use the value of the secondary transform syntax element to determine whether to perform a secondary transform, and if so, which of a plurality of secondary transforms to apply. Figure 7 Thus, inverse transform 78 initially inverse transforms the transform coefficients (290) using the secondary transform to produce intermediate transform coefficients, and then inverse transforms the intermediate transform coefficients (292) using the first transform (eg, DCT or EMT) to regenerate a residual block for the block.
[0190] Intra-prediction unit 74 also intra-predicts the block using the indicated intra-prediction mode (294) to produce a predicted block for the block. Summer 80 then combines the predicted block with the residual block on a pixel-by-pixel basis to produce a decoded block (296). Ultimately, video decoder 30 outputs the decoded block. Video decoder 30 may also store the decoded block in reference picture memory 82, e.g., for use in intra- or inter-predicting subsequently decoded blocks.
[0191] In this way, Figure 7 The method represents an example of a method including the following operations: determining a maximum possible value of a secondary transform syntax element for a block of video data; entropy decoding the value of the secondary transform syntax element of the block to form a binarized value representing a secondary transform for the block; inversely binarizing the value of the secondary transform syntax element based on the determined maximum possible value to determine a secondary transform for the block; and inverse transforming transform coefficients of the block using the determined secondary transform.
[0192] It should be recognized that, depending on the example, certain actions or events of any of the techniques described herein may be performed in a different sequence, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the techniques). Furthermore, in some examples, actions or events may be performed concurrently (e.g., via multithreading, interrupt handling, or multiple processors) rather than sequentially.
[0193] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media may include: computer-readable storage media, which corresponds to tangible media (such as data storage media); or communication media, which includes any media that facilitates the transfer of a computer program from one place to another (for example, according to a communication protocol). In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present invention. A computer process product may include a computer-readable medium.
[0194] As examples and not limitations, these computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using a coaxial cable, optical cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio and microwave), then the coaxial cable, optical cable, twisted pair, DSL, or wireless technology (such as infrared, radio and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other temporary media, but rather refer to non-temporary tangible storage media. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, while optical disks employ lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.
[0195] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Thus, the term "processor" as used herein may refer to any of the aforementioned structures or any other structures suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.
[0196] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, or provided by a set of interoperable hardware units (including one or more processors as described above) in conjunction with appropriate software and / or firmware.
[0197] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for decoding video data, the method comprising: decoding transform coefficients of a transform block of a current block of video data; determining positions of non-zero-valued transform coefficients of the decoded transform coefficients in the transform block; determining a type of transform to be performed based on the positions of non-zero-valued transform coefficients in the transform block; transforming the transform coefficients in the transform block using the type of transform to generate a residual block; as well as Decoding the current block using the residual block; Wherein, determining the type of the transformation includes: determining a number of non-zero-valued transform coefficients following an Nth transform coefficient in the transform block in scan order, where N is an integer value; and Whether the type of the transform is a separable transform or a non-separable transform is determined according to the number of non-zero-valued transform coefficients following the Nth transform coefficient in a scan order.
2. The method according to claim 1, wherein: Determining the type of the transform includes determining the type of the transform without decoding a value of a syntax element of the video data indicating the type of the transform. The method according to claim 1 , further comprising determining the value of N to be a predefined value.
4. The method according to claim 1, wherein: Determining the type of the transform includes: when the number of non-zero-valued transform coefficients after the Nth transform coefficient is zero, determining the type of the transform to be a non-separable transform.
5. The method according to claim 1, wherein: Determining the type of the transformation includes: When the number of non-zero-valued transform coefficients after the Nth transform coefficient is greater than zero: decoding a value of a syntax element indicating a type of transform coefficient; and Determining the type of transform coefficient according to the value of the syntax element; and When the number of non-zero-valued transform coefficients after the Nth transform coefficient is equal to zero, the type of the transform is determined without decoding the value of the syntax element.
6. The method according to claim 1, wherein: Decoding the transform coefficients comprises: determining, based on a value of a syntax element corresponding to the type of the transform, a context model for entropy decoding a value of one of the transform coefficients; and A value of one of the transform coefficients is entropy decoded using the context model.
7. The method according to claim 1, wherein: Decoding the current block comprises: forming a prediction block for the current block; and The samples of the prediction block are combined with the samples of the residual block.
8. The method of claim 1, further comprising encoding the current block before decoding the current block.
9. The method according to claim 8, wherein: Transforming the transform coefficients comprises applying a secondary transform of the type of transform to the transform coefficients.
10. An apparatus for decoding video data, the apparatus comprising: a memory configured to store video data; as well as One or more processors implemented as circuits and configured to: decoding transform coefficients of a transform block of a current block of video data; determining positions of non-zero-valued transform coefficients of the decoded transform coefficients in the transform block; determining a type of transform to be performed based on the positions of non-zero-valued transform coefficients in the transform block; transforming transform coefficients in the transform block using the type of transform to generate a residual block; and Decoding the current block using the residual block; In order to determine the type of the transformation, the one or more processors are configured to: determining a number of non-zero-valued transform coefficients following an Nth transform coefficient in the transform block in scan order, where N is an integer value; as well as Whether the type of the transform is a separable transform or a non-separable transform is determined according to the number of non-zero-valued transform coefficients following the Nth transform coefficient in a scan order.
11. The device according to claim 10, wherein: The one or more processors are configured to determine the type of the transform without decoding a value of a syntax element of the video data indicating the type of the transform.
12. The device according to claim 10, wherein: The one or more processors are further configured to determine the value of N to be a predefined value.
13. The device according to claim 10, wherein: The one or more processors are configured to determine that the type of the transform is a non-separable transform when the number of non-zero-valued transform coefficients after the Nth transform coefficient is zero.
14. The device according to claim 10, wherein: To determine the type of the transformation, the one or more processors are configured to: When the number of non-zero-valued transform coefficients after the Nth transform coefficient is greater than zero: decoding a value of a syntax element indicating a type of transform coefficient; and Determining the type of transform coefficient according to the value of the syntax element; as well as When the number of non-zero-valued transform coefficients after the Nth transform coefficient is equal to zero, the type of the transform is determined without decoding the value of the syntax element.
15. The device according to claim 10, wherein: To decode the current block, the one or more processors are configured to forming a prediction block for the current block; and The samples of the prediction block are combined with the samples of the residual block.
16. The device according to claim 10, wherein: The one or more processors are further configured to encode the current block prior to decoding the current block.
17. The apparatus of claim 10, further comprising a display configured to display the decoded video data.
18. The device according to claim 10, wherein: The device includes one or more of: a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
19. The apparatus of claim 10, wherein the one or more processors are configured to apply a secondary transform of the type of transform to the transform coefficients.
20. An apparatus for decoding video data, the apparatus comprising means for decoding transform coefficients of a transform block of a current block of video data; means for determining positions of non-zero-valued transform coefficients of the decoded transform coefficients in the transform block; means for determining a type of transform to be performed based on the positions of non-zero-valued transform coefficients in the transform block; means for transforming transform coefficients in said transform block using said type of transform to produce a residual block; as well as means for decoding the current block using the residual block; Wherein, the component for determining the type of the transformation includes: means for determining a number of non-zero-valued transform coefficients following an Nth transform coefficient in said transform block in scan order, wherein N is an integer value; and Means for determining whether the type of transform is a separable transform or a non-separable transform based on the number of non-zero-valued transform coefficients following the Nth transform coefficient in scan order.
21. The device according to claim 20, wherein: The means for determining the type of transform comprises means for determining the type of transform without decoding a value of a syntax element of the video data indicating the type of transform.
22. The device according to claim 20, wherein: The means for transforming the transform coefficients comprises means for applying a secondary transform of said type of transform to the transform coefficients.
23. A computer-readable storage medium having stored thereon instructions which, when executed, cause a processor of an apparatus for decoding video data to: decoding transform coefficients of a transform block of a current block of video data; determining positions of non-zero-valued transform coefficients of the decoded transform coefficients in the transform block; determining a type of transform to be performed based on the positions of non-zero-valued transform coefficients in the transform block; transforming transform coefficients in the transform block using the transform of the type to generate a residual block; as well as Decoding the current block using the residual block; in, To determine the type of the transform, the instructions, when executed, further cause a processor of the apparatus for decoding video data to: determining a number of non-zero-valued transform coefficients following an Nth transform coefficient in the transform block in scan order, where N is an integer value; as well as Whether the type of the transform is a separable transform or a non-separable transform is determined according to the number of non-zero-valued transform coefficients following the Nth transform coefficient in a scan order.
Citation Information
Patent Citations
Video coding using transforms bigger than 4x4 and 8x8
CN102204251A
Coding of significance maps and transform coefficient blocks
CN102939755A
Video coding using mapped transforms and scanning modes
CN103329523A
Signaling syntax elements for transform coefficients for sub-sets of leaf-level coding unit
CN103636225A