Parity encoding of quantized transform coefficients by interleaving
By partitioning quantized transform coefficients into chunks with parity leaders and modifying coefficients to satisfy parity constraints, the method enhances video encoding efficiency by reducing bit rate and computational complexity through effective parity encoding and sign omission.
Patent Information
- Application Number
- PCT/US2024/057538
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Existing video encoding techniques face challenges in efficiently encoding and decoding quantized transform coefficients, particularly in managing parity encoding and probability distributions, which leads to increased complexity and limited ability to parity encode multiple coefficients.
The method involves partitioning quantized transform coefficients into chunks with respective parity leaders, where a parity leader is a coefficient with a magnitude greater than or equal to a threshold T. Each chunk's sum is checked against a parity constraint, and coefficients are modified as needed to satisfy the constraint. The method encodes parity-encoded values for parity leaders and omits encoding signs for parity leaders, allowing for inference at the decoder.
This approach reduces bit rate by omitting sign encoding for parity leaders and allows for efficient parity encoding of multiple coefficients, thereby improving encoding efficiency and reducing computational complexity.
Smart Images

Figure US2024057538_05062025_PF_FP_ABST
Abstract
Description
PARITY ENCODING OF QUANTIZED TRANSFORM COEFFICIENTS BY INTERLEAVINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 604,249, filed November 30, 2023, the entire disclosure of which is incorporated herein by reference.BACKGROUND
[0002] Digital video streams may represent video using a sequence of frames or still images. Digital video can be used for various applications including, for example, video conferencing, high definition video entertainment, video advertisements, or sharing of usergenerated videos. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various approaches have been proposed to reduce the amount of data in video streams, including encoding or decoding techniques.SUMMARY
[0003] These and other aspects of the present disclosure are disclosed in the following detailed description of the embodiments, the appended claims and the accompanying figures.
[0004] One general aspect includes a method. The method includes decoding a vector of values indicative of magnitudes of quantized transform coefficients from a compressed bitstream. The method also includes partitioning the vector into chunks that include respective parity leaders, where a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T. The method also includes, for each parity leader obtaining a magnitude of a corresponding quantized transform coefficient based on a sum of the coefficients of the chunk that includes the parity leader. The method also includes decoding, from the compressed bitstream, respective signs for quantized transform coefficients other than the parity leaders. The method also includes inferring respective signs of quantized transform coefficients corresponding to the parity leaders. The method also includes obtaining the quantized transform coefficients based on themagnitudes and the respective signs. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. Other embodiments include a non-transitory computer-readable storage medium having stored thereon an encoded bitstream that is configured for decoding by the actions of the method.
[0005] Implementations may include one or more of the following features.
[0006] The method where the threshold T is set based on a quantization parameter.
[0007] A number of the chunks can be inversely related to a quantization parameter.
[0008] Obtaining the magnitude of the corresponding quantized transform coefficient may include decoding the magnitude of the corresponding quantized transform coefficient using a range-folding technique.
[0009] The chunks include the respective parity leaders, and remaining coefficients of the quantized transform coefficients other than the respective parity leaders may be assigned to the chunks so that the chunks have roughly a same number of quantized transform coefficients.
[0010] The remaining coefficients may be assigned to the chunks based on a lexicographical ordering and sequential assignment performed by an encoder.
[0011] Decoding the respective signs for quantized transform coefficients other than the parity leaders may include determining whether parity encoding is applied for a current coefficient; if parity encoding is applied to the current coefficient, not decoding a sign of the current coefficient; and if parity encoding is not applied to the current coefficient, decoding the sign of the current coefficient.
[0012] Determining whether parity encoding is applied for the current coefficient may include determining whether a magnitude of the current coefficient is greater than or equal to the threshold T.
[0013] One general aspect includes a method for encoding a vector of quantized transform coefficients. The method includes partitioning the quantized transform coefficients into chunks, where each chunk includes one parity leader, where a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T, and where the chunks include different parity leaders. The method also includes for each chunk: determining whether a sum of coefficients in the chunk satisfies a parity constraint; and responsive to determining that the sum does not satisfy the parity constraint, modifying at least one coefficient of the chunk so that the sum satisfies the parity constraint. The method also includes encoding the quantized transform coefficients in acompressed bitstream by: obtaining respective parity encoded values for the parity leaders according to specific equations depending on the threshold T and the parity constraint, and encoding the respective parity encoded values. The method also includes encoding respective coefficient signs of all of the quantized transform coefficients except for respective signs of the parity leaders. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. Other embodiments may include a non-transitory computer-readable storage medium having stored thereon an encoded bitstream that is generated by an encoder performing the operations of the method.
[0014] Implementations may include one or more of the following features.
[0015] The method where the quantized transform coefficients other than the different parity leaders constitute remaining coefficients, and where the remaining coefficients may be assigned to the chunks such that the chunks have roughly a same number of coefficients.
[0016] The remaining coefficients may be assigned to the chunks by lexicographically ordering the remaining coefficients and sequentially assigning the remaining coefficients to the chunks.
[0017] Modifying the at least one coefficient may include increasing or decreasing a value of the at least one coefficient.
[0018] Modifying the at least one coefficient of the chunk may include shifting at least some of the quantized transform coefficients qi by a number to obtain qi*, while ensuring that if lqil<t then a shift of qi also satisfies lqi*l<t or if lqd>t then the shift of the qi also satisfies lqi*l>t, where z is an index of the at least some of the quantized transform coefficients; and identifying the at least one coefficient that results in a smallest cost shift based on ratedistortion analysis; and modifying the at least one coefficient.
[0019] The threshold T can be set based on a quantization parameter.
[0020] A number of the chunks can be inversely related to a quantization parameter.
[0021] The parity constraint can be an evenness parity constraint or an oddness parity constraint.
[0022] Obtaining the respective parity encoded values for the different parity leaders may use a range folding technique that maps negative values of coefficients to positive odd values.
[0023] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects may be implemented by appropriate computer programs which may be carried on appropriate carrier media which may be tangible carrier media (e.g., disks) or intangible carrier media (e.g., communications signals). Aspects may also be implementedusing suitable apparatus which may take the form of programmable computers running computer programs arranged to implement the methods and / or techniques disclosed herein. Aspects can be combined such that features described in the context of one aspect may be implemented in another aspect.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The description herein makes reference to the accompanying drawings described below, wherein like reference numerals refer to like parts throughout the several views.
[0025] FIG. 1 is a schematic of a video encoding and decoding system.
[0026] FIG. 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0027] FIG. 3 is a diagram of a typical video stream to be encoded and subsequently decoded.
[0028] FIG. 4 is a block diagram of an encoder according to implementations of this disclosure.
[0029] FIG. 5 is a block diagram of a decoder according to implementations of this disclosure.
[0030] FIG. 6 is a flowchart of a technique for encoding a quantized transform block.
[0031] FIG. 7 is a flowchart of a technique for decoding a quantized transform block.
[0032] FIG. 8 is a flowchart of a technique for encoding a vector of quantized transform coefficients.
[0033] FIG. 9 is a flowchart of a technique for decoding a quantized transform block.DETAIEED DESCRIPTION
[0034] As mentioned above, compression schemes related to coding video streams may include breaking images into blocks and generating a digital video output bitstream using one or more techniques to limit the information included in the output. A received encoded bitstream can be decoded to re-create the blocks and the source images from the limited information. Encoding a video stream, or a portion thereof, such as a frame or a block, can include using temporal or spatial similarities in the video stream to improve coding efficiency. For example, a current block of a video stream may be encoded based on identifying a difference (residual) between the previously coded pixel values and those in the current block. In this way, only the residual and parameters used to generate the residual needbe added to the encoded bitstream. The residual may be encoded using a lossy quantization step.
[0035] As further described below, the residual block can be in the pixel domain. The residual block can be transformed into the frequency domain resulting in a transform block of transform coefficients. The transform coefficients can be quantized resulting into a quantized transform block of quantized transform coefficients. The quantized coefficients can be entropy encoded and added to an encoded bitstream. A decoder can receive the encoded bitstream, entropy decode the quantized transform coefficients to reconstruct the original block.
[0036] A quantized transform block is a two-dimensional structure that is linearized into a one dimensional vector, denoted q herein, based on a scan order. The quantized transform block may have a size of PxQ, wherein P*Q=N. Thus, the vector q has a size of Nxl (i.e., q(N l )). At the decoder, a corresponding scan order is used to read the vector q of coefficients from the encoded bitstream and covert it back into the two-dimensional quantized transform block.
[0037] In an encoded video bitstream, many of the bits are used for one of two things: either content prediction (e.g., inter mode / motion vector coding, intra prediction mode coding, etc.) or residual coding (e.g., coding of the quantized transform coefficients). Encoders may use techniques to decrease the number of bits spent on coefficient coding. One such technique may be referred to as information hiding or using parity to infer some aspect (e.g., a value) related to at least one of the quantized transform coefficients.
[0038] To illustrate, a codec (i.e., an encoder and a decoder) may impose an evenness constraint on the sum of quantized transform coefficients within a block. Symbolically, the evenness constraint can be stated as: Z^o1Qi=even integer. This constraint results in a rate advantage when encoding the vector q. Equation (1) illustrates the operations of this constraint. Equation (1) states that if the sum of coefficients from the second to the last ( q to qw_i) is even, then the first coefficient (qr0) must be even. Conversely, if the sum of qrto qw_i is odd, then qQmust be odd. It is noted that while the first coefficient, q0, is used for illustrative purposes, the description herein can be applied to any of the coefficients qc, where c = 0, ... , A - 1.
[0039] The constraint enables encoding only half of the potential values for q0, effectively compressing its range and reducing the number of bits required for its representation. The decoder can in turn deduce the value of qQbased on the sum of the other coefficients (qrto QN-I). Instead of encoding q0, qo' given by equation (2) can be encoded instead. Equation (2) implements what may be referred to as “range compression” whereby the range of possible values that a data element can assume is reduced, thereby enabling more efficient coding by using fewer bits to represent the data.
[0040] Encoding qo' in place of qQcan be expected to save, on average, approximately one bit during the encoding of the vector q. This is because each of the quantized transform coefficients qLis typically assumed to be Laplacian or Gaussian for which a range compression of two will result in an average rate reduction by one bit. So that the constraint is always satisfied, in the case that the coefficients q to qN-^ sum to an odd number, the encoder will adjust (e.g., by +1 or -1) at least one of these coefficients so that the sum becomes even. As such, the one-bit reduction may be offset by an increase in distortion for vectors of quantized transform coefficients that sum to an odd value.
[0041] However, such a technique (e.g., the range compression technique) suffers from several disadvantages.
[0042] Firstly, the transformation from qcto qc' alters the probability distribution of the variable. In the context of signal, image, or video compression where the vector q represents quantized transform coefficients, this alteration presents a significant challenge. An image comprises numerous transform blocks, each represented by a respective sequence of q vectors. Selectively (i.e., turning it on and off for different blocks) implementing the parity encoding scheme described above necessitates distinct handling of qc' coefficients in parity- encoded blocks compared to qccoefficients in non-parity-encoded blocks. This is due to the fact that parity-encoded coefficients exhibit a range different from that of their non-parity- encoded counterparts. Consequently, this requires maintaining two separate sets of conditional distribution functions (CDFs) to optimize rate reduction effectively. However, increasing the number of CDFs in a codec is undesirable.
[0043] Secondly, the modification in probability distribution (as mentioned above) often mandates either selecting a particular quantized transform coefficient (i.e., fixing the value of c) or employing conditional entropy coding for a multitude of coefficients (e.g., coefficientpositions), each potentially having different values of c. Such a requirement can lead to excessive complexity in the encoding and decoding processes.
[0044] Thirdly, there is an essential requirement for the qccoefficient to be nonzero, and ideally, its absolute value should be at least 2 (i.e., | qc| > 2). To illustrate, a qcvalue of -1 would result in qc' being -1, yielding no advantage in terms of rate gain. Ensuring a nonzero typically compels the encoding system to Constantly encode the DC coefficient (i.e., setting c=0 corresponding to the quantized transform coefficient at location (0, 0) of the quantized transform block), as the DC coefficient is most likely to be nonzero; and limit the encoding approach to avoid multiple parity encodings within a single q vector. For instance, if the decision (by an encoder) is to partition the vector q into two subsets and apply parity encoding to a coefficient in each subset, identifying a second, fixed non-DC coefficient for encoding becomes challenging. This is due to the frequent quantization of non-DC coefficients to zero. Additionally, the application of parity encoding to multiple coefficients within a vector and allowing for variable coefficient locations re-introduces the increased complexity issue described above.
[0045] As such, implementations according to the above process are limited to using parity encoding with respect to one coefficient and that coefficient has to be the DC coefficient since it has the highest probability of being none zero. As such, given a large quantized transform block of size PxQ (e.g., 32x32), which includes P*Q (e.g., 32*32=1024) coefficients, it is not possible to parity encode more than one of these coefficients.
[0046] Implementations of disclosure solve problems such as those by applying a parity constraint to coefficients identified as “parity leaders.” Parity leaders can be identified and coded based on their respective magnitudes relative to a predetermined threshold. That is, a parity leader can be any coefficient qcwhere the magnitude (i.e., the absolute value) is greater than or equal to a positive integer threshold T (i.e., |qc| > T). As such, it is possible to parity code more than one coefficient of the vector q.
[0047] Coding a parity leader, and as further described herein, can mean encoding a value qc' which may be different from the actual value qcof the parity leader. The coding process described herein is invertible. As such, the actual value of the parity leader is recoverable(such as by a decoder) from the encoded value. The coding process can vary depending on whether the threshold T is even or odd, and whether qcis even or odd. This coding process ensures that the encoded value qc' maintains a magnitude greater than or equal to the threshold T.
[0048] Coding quantized transform coefficients typically includes coding respective magnitudes and signs of the quantized transform coefficients. According to implementations of this disclosure, sign bits of the parity leaders are omitted by the encoder and can be inferred at the decoder based on the encoded values of parity leaders, therewith resulting in bit rate reduction.
[0049] Further details of techniques of parity encoding of quantized transform coefficients by interleaving are described herein with initial reference to a system in which they can be implemented. FIG. 1 is a schematic of a video encoding and decoding system 100. A transmitting station 102 can be, for example, a computer having an internal configuration of hardware such as that described in FIG. 2. However, other implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.
[0050] A network 104 can connect the transmitting station 102 and a receiving station 106 for encoding and decoding of the video stream. Specifically, the video stream can be encoded in the transmitting station 102, and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), wide area network (WAN), virtual private network (VPN), cellular telephone network, or any other means of transferring the video stream from the transmitting station 102 to, in this example, the receiving station 106.
[0051] The receiving station 106, in one example, can be a computer having an internal configuration of hardware such as that described in FIG. 2. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0052] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit the network 104. In another implementation, a video stream can be encoded and then stored for transmission at a later time to the receiving station 106 or any other device having memory. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and / or some communication pathway) the encoded video stream and stores the video stream for later decoding. In an example implementation, a real-time transport protocol (RTP) is used for transmission of the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used (e.g., a Hypertext Transfer Protocol based (HTTP based) video streaming protocol).
[0053] When used in a video conferencing system, for example, the transmitting station 102 and / or the receiving station 106 may include the ability to both encode and decode a video stream as described below. For example, the receiving station 106 could be a video conference participant who receives an encoded video bitstream from a video conference server (e.g., the transmitting station 102) to decode and view and further encodes and transmits his or her own video bitstream to the video conference server for decoding and viewing by other participants.
[0054] FIG. 2 is a block diagram of an example of a computing device 200 that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement one or both of the transmitting station 102 and the receiving station 106 of FIG. 1. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of one computing device, for example, a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and the like.
[0055] A processor 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the processor 202 can be another type of device, or multiple devices, capable of manipulating or processing information now existing or hereafter developed. For example, although the disclosed implementations can be practiced with one processor as shown (e.g., the processor 202), advantages in speed and efficiency can be achieved by using more than one processor.
[0056] A memory 204 in computing device 200 can be a read only memory (ROM) device or a random access memory (RAM) device in an implementation. However, other suitable types of storage device can be used as the memory 204. The memory 204 can include code and data 206 that is accessed by the processor 202 using a bus 212. The memory 204 can further include an operating system 208 and application programs 210, the application programs 210 including at least one program that permits the processor 202 to perform the techniques described herein. For example, the application programs 210 can include applications 1 through N, which further include a video coding application that performs the techniques described herein. The computing device 200 can also include a secondary storage 214, which can, for example, be a memory card used with a mobile computing device.Because the video communication sessions may contain a significant amount of information, they can be stored in whole or in part in the secondary storage 214 and loaded into the memory 204 as needed for processing.
[0057] The computing device 200 can also include one or more output devices, such as a display 218. The display 218 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 218 can be coupled to the processor 202 via the bus 212. Other output devices that permit a user to program or otherwise use the computing device 200 can be provided in addition to or as an alternative to the display 218. When the output device is or includes a display, the display can be implemented in various ways, including by a liquid crystal display (LCD), a cathode-ray tube (CRT) display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.
[0058] The computing device 200 can also include or be in communication with an image-sensing device 220, for example, a camera, or any other image-sensing device 220 now existing or hereafter developed that can sense an image such as the image of a user operating the computing device 200. The image-sensing device 220 can be positioned such that it is directed toward the user operating the computing device 200. In an example, the position and optical axis of the image-sensing device 220 can be configured such that the field of vision includes an area that is directly adjacent to the display 218 and from which the display 218 is visible.
[0059] The computing device 200 can also include or be in communication with a soundsensing device 222, for example, a microphone, or any other sound-sensing device now existing or hereafter developed that can sense sounds near the computing device 200. The sound-sensing device 222 can be positioned such that it is directed toward the user operating the computing device 200 and can be configured to receive sounds, for example, speech or other utterances, made by the user while the user operates the computing device 200.
[0060] Although FIG. 2 depicts the processor 202 and the memory 204 of the computing device 200 as being integrated into one unit, other configurations can be utilized. The operations of the processor 202 can be distributed across multiple machines (wherein individual machines can have one or more processors) that can be coupled directly or across a local area or other network. The memory 204 can be distributed across multiple machines such as a network-based memory or memory in multiple machines performing the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 can be composed of multiple buses. Further, the secondary storage 214 can be directly coupled to the other components of the computing device 200 or can be accessed via a network and can comprise an integrated unit such as a memory card ormultiple units such as multiple memory cards. The computing device 200 can thus be implemented in a wide variety of configurations.
[0061] FIG. 3 is a diagram of an example of a video stream 300 to be encoded and subsequently decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of adjacent frames 304. While three frames are depicted as the adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, for example, a frame 306. At the next level, the frame 306 can be divided into a series of planes or segments 308. The segments 308 can be subsets of frames that permit parallel processing, for example. The segments 308 can also be subsets of frames that can separate the video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.
[0062] Whether or not the frame 306 is divided into segments 308, the frame 306 may be further subdivided into blocks 310, which can contain data corresponding to, for example, 16x16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of pixel data. The blocks 310 can also be of any other suitable size such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger. Unless otherwise noted, the terms block and macroblock are used interchangeably herein.
[0063] FIG. 4 is a block diagram of an encoder 400 according to implementations of this disclosure. The encoder 400 can be implemented, as described above, in the transmitting station 102, such as by providing a computer software program stored in memory, for example, the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the processor 202, cause the transmitting station 102 to encode video data in the manner described in FIG. 4. The encoder 400 can also be implemented as specialized hardware included in, for example, the transmitting station 102. In one particularly desirable implementation, the encoder 400 is a hardware encoder.
[0064] The encoder 400 has the following stages to perform the various functions in a forward path (shown by the solid connection lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy encoding stage 408. The encoder 400 may also include a reconstruction path (shown by the dotted connection lines) to reconstruct a frame for encoding of future blocks. In FIG. 4, the encoder 400 has the following stages to perform the various functions in the reconstruction path: a dequantizationstage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0065] When the video stream 300 is presented for encoding, respective adjacent frames 304, such as the frame 306, can be processed in units of blocks. At the intra / inter prediction stage 402, respective blocks can be encoded using intra-frame prediction (also called intraprediction) or inter- frame prediction (also called inter-prediction). In any case, a prediction block can be formed. In the case of intra-prediction, a prediction block may be formed from samples in the current frame that have been previously encoded and reconstructed. In the case of inter-prediction, a prediction block may be formed from samples in one or more previously constructed reference frames.
[0066] Next, the prediction block can be subtracted from the current block at the intra / inter prediction stage 402 to produce a residual block (also called a residual). The transform stage 404 transforms the residual into transform coefficients in, for example, the frequency domain using block-based transforms. The quantization stage 406 converts the transform coefficients into discrete quantum values, which are referred to as quantized transform coefficients, using a quantizer value or a quantization level. For example, the transform coefficients may be divided by the quantizer value and truncated.
[0067] The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy-encoded coefficients, together with other information used to decode the block (which may include, for example, syntax elements such as used to indicate the type of prediction used, transform type, motion vectors, a quantizer value, or the like), are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques, such as variable length coding (VLC) or arithmetic coding. The compressed bitstream 420 can also be referred to as an encoded video stream or encoded video bitstream, and the terms will be used interchangeably herein.
[0068] The reconstruction path (shown by the dotted connection lines) can be used to ensure that the encoder 400 and a decoder 500 (described below with respect to FIG. 5) use the same reference frames to decode the compressed bitstream 420. The reconstruction path performs functions that are similar to functions that take place during the decoding process (described below with respect to FIG. 5), including dequantizing the quantized transform coefficients at the dequantization stage 410 and inverse transforming the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also called a derivative residual). At the reconstruction stage 414, the prediction blockthat was predicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filtering stage 416 can be applied to the reconstructed block to reduce distortion such as blocking artifacts.
[0069] Other variations of the encoder 400 can be used to encode the compressed bitstream 420. In some implementations, a non-transform based encoder can quantize the residual signal directly without the transform stage 404 for certain blocks or frames. In some implementations, an encoder can have the quantization stage 406 and the dequantization stage 410 combined in a common stage.
[0070] FIG. 5 is a block diagram of a decoder 500 according to implementations of this disclosure. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the processor 202, cause the receiving station 106 to decode video data in the manner described in FIG. 5. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.
[0071] The decoder 500, similar to the reconstruction path of the encoder 400 discussed above, includes in one example the following stages to perform various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512, and a post filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.
[0072] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by the quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derivative residual that can be identical to that created by the inverse transform stage 412 in the encoder 400. Using header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter prediction stage 508 to create the same prediction block as was created in the encoder 400 (e.g., at the intra / inter prediction stage 402).
[0073] At the reconstruction stage 510, the prediction block can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be appliedto the reconstructed block to reduce blocking artifacts. Other filtering can be applied to the reconstructed block. In this example, the post filtering stage 514 is applied to the reconstructed block to reduce blocking distortion, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream, and the terms will be used interchangeably herein. Other variations of the decoder 500 can be used to decode the compressed bitstream 420. In some implementations, the decoder 500 can produce the output video stream 516 without the post filtering stage 514.
[0074] FIG. 6 is a flowchart of a technique 600 for encoding a quantized transform block. The technique 600 can be used to encode the quantized transform coefficients of the quantized transform block in a compressed bitstream, such as the compressed bitstream 420 of FIG. 4. The quantized transform coefficients can be arranged into a vector q.
[0075] The technique 600 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 600. The technique 600 may be implemented in whole or in part in the entropy encoding stage 408 of the encoder 400 of FIG. 4. The technique 600 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
[0076] The technique 600 implements a parity encoding of quantized transform coefficients by interleaving, an overview is now described. The parity encoding scheme disclosed herein can be applied to one or more coefficients that satisfy the condition of equation (3), where T is a predetermined positive integer (e.g., T=2, 3, . . .). Such coefficients are referred to as potential parity leaders. At least a subset of the potential parity leaders can then be used as parity leaders.\qc\ > T (3)
[0077] Encoding a parity leader qcmeans encoding a value qc' from which qccan be derived. In the case that the threshold T is selected to be an even number, then the encoded value qc' of a parity leader qccan be obtained using equation (4) in the case of an evenness constraint (i.e., in the case that qcis constrained to be even) and using equation (5) in the case of an oddness constraint (i.e., in the case that qcis constrained to be odd). With respect to equation (4), since the magnitude of qcis greater than the threshold T, then so is the magnitudeof qc' (i.e., Iq > T). With respect to equation (5), since the threshold T is even and qcis odd, then it follows that | qc| > T and that that > T., ( qcif qcT , , qr= . .rand T is even and qris assumed to be even (4){-qc+ 1 if qc< -T, ( qcif qc> T , , qr= ,rand T is even and qris assumed to be odd (5){-qc- 1 if qc< -T
[0078] In the case that the threshold T is selected to be an odd number, then the encoded value qc' of a parity leader qccan be obtained using equation (6) in the case of an evenness constraint (i.e., in the case that qcis constrained to be even) and using equation (7) in the case of an oddness constraint (i.e., in the case that qcis constrained to be odd). With respect to both equations (6) and (7), it can be observed thatT.
[0079] The following observations with the respect to parity encoding scheme follow from the foregoing.
[0080] The mapping (e.g., encoding / decoding) between qcand qc' is invertible. If the decoder were expecting an even (odd) qc' , but receives an odd (even) value, then the received value can be folded into the negative side, based on the above equations.
[0081] Since qc' is greater than or equal to T (for T>0), the, qc' is always positive. Therefore, approximately one bit of rate savings can be realized by not encoding the sign of qc. That is, if the decoder were expecting an even (odd) qc' , but receives an odd (even) value, then the sign of qccan be inferred to be negative. That is, if the oddness or evenness of a received qc' is not consistent with what is expected at the decoder, then the qcis inferred to be a negative value; on the other hand, if the oddness or evenness of a received qc' is consistent with what is expected at the decoder, then the qcis inferred to be a positive. In either case, the sign of qcneed not be encoded by the encoder therewith resulting in a one bit rate saving by not encoding the sign of qc.
[0082] As most transform-based image / video codecs encode magnitudes and signs of quantized transform coefficients separately, the teachings herein can be used within such codecs by changing the coding semantics to not transmit the sign of qc. Decoding a sign bit essentially includes the steps of determining whether parity encoding is applied for a currentcoefficient. If parity encoding is applied to a current coefficient, then the compressed bitstream would not include the sign of the current coefficient.
[0083] As further described herein, a protocol (between an encoder and a decoder) for determining whether parity encoding is applied for a current coefficient is implemented. If parity encoding is applied to a current coefficient, the decoder does not decode and deduces the sign of the current coefficient.
[0084] The sign bits associated with quantized transform coefficients are typically encoded using a probability distribution. According to the described scheme there is no need for separate conditional probabilities (e.g., based on whether a coefficient is parity encoded or not). Rather, the same distribution can be used for the encoded signs and the encoder and decoder omit encoding and decoding, respectively, signs associated with certain coefficients (i.e., the parity leaders).
[0085] Encoding certain values based on the teachings herein may result in a slight rate increase and encoding certain other values may result in a slight rate decrease. To illustrate, assuming that qc= —4 and T=2, then according to equation (4), the value transmitted would be qc' = 5, therewith resulting in a slight rate increase (e.g., as compared to coding the magnitude 4 since the value 5 requires more bits than 4). On the other hand, if qc= —3, then according to equation (5), the value transmitted would be qc' = 2, therewith resulting in a slight rate gain (e.g., as compared to coding the magnitude 3 since the value 2 requires fewer bits than 3). Even if these rate gains and penalties do not cancel each other, any rate penalties incurred are offset by the 1 -bit rate gains achieved by omitting encoding the signs of the parity leaders.
[0086] Since the encoding can be defined for any index c for which |qc| > T, then parity encoding multiple coefficients within the vector q of quantized transform coefficients is possible. Whereas, as described above, conventional techniques may apply parity encoding to a coefficients that is greater than 0 using different probability distributions (e.g., based on whether the coefficient is range compressed or not), the techniques described herein may be applied to coefficients that meet the condition | qc| > T without resorting to multiple distributions. To be clear, any quantized transform coefficient qcthat meets the condition | qc| > T can be a potential parity leader.
[0087] The described scheme does not use range compression. Rather, the described scheme uses or implements what is referred to herein as “range folding.” Thus, the same probability distribution can be used to encode / decode all coefficients. To describe the conceptof folding, reference is made to equation (4). With equation (4), the parity leaders are assumed to be even. Thus, assuming that T=2, the possible values of q are{. . . , —14, —10, —8, —6, —4, —2, 2, 4, 6, 8, 10, 12, 14, . . . } which are mapped, via equation (4) to the values q' = {15, 13, 11, 9, 7, 5, 3, 2, 4, 6, 8, 10, 12, 14, . . .As can be seen, the negative even values of q are folded into (e.g., mapped to) the empty, odd slots of the number line.
[0088] Turning now to the technique 600, the technique 600 encodes a vector q of quantized transform coefficients into a vector q’ . The technique 600 encodes magnitudes of the coefficients followed by their signs.
[0089] Any number of techniques can be used to code the magnitudes, and the disclosure herein is not limited to or by any particular way of encoding quantized transform coefficient magnitudes. In an example, quantized transform coefficient values may be coded with multiple level maps together with sign values. The level maps may be coded as three level planes, namely, lower-level, middle-level, and higher-level planes, and the sign is coded as a separate plane. The lower-, middle-, and higher-level planes correspond to different ranges of coefficient magnitudes (0-2, 3-14, 15 and above, respectively). While the disclosure herein refers to encoding or decoding a magnitude, it is noted that the full magnitudes may not be necessary. As long as whether a coefficient magnitude is above or below the threshold T is conveyed before sign decoding, the technique described herein can be implemented with insignificant changes during entropy coding.
[0090] At 602, the potential parity leaders in the vector q of quantized coefficients are identified. As described above, the potential parity leaders are those coefficients whose magnitudes are greater than or equal to a threshold T. Symbolically, the number of potential parity leaders, nq, can be determined using equation ().
[0091] In an example, the threshold T can be set to predefined value (e.g., 2, 3, or some other value). In n example, the threshold T can be set based on the quantization parameter associated with the block. The quantization parameter (QP) may be set for the frame that includes the block corresponding to the quantized transform block. The QP may be set for a group of frames (also referred to as a group of pictures (GOP)). The QP may be set at the sequence (e.g., the whole video sequence) level.
[0092] At 604, the number of chunks (e.g., groups), mq, to partition the vector q into is determined. The number of chunks mqcan be determined using a function f() (e.g., a “chunk numbers determining function”) that takes, as input, at least the number of potential parityleaders, nq. The chunk numbers determining function returns a positive or zero number that is less than or equal to the number of potential parity leaders, as given by inequality (9). To illustrate, assuming that there are 5 potential parity leaders (e.g., nq= 5), then there could be 0, 1, 2, 3, 4, or 5 chunks; but 6 or more chunks are not possible. If the function returns 0, then parity encoding is not applied to the block.
[0093] The chunk numbers determining function is a deterministic function between the encoder and the decoder. In an example, the chunk numbers determining function can be fixed or signaled by the encoder to the decoder at the sequence, frame, or block granularity. In an example, the chunk numbers determining function can be selected based on the QP. The number of chunks can be inversely related to the QP. For example, for high quality, high bitrate areas, the function may return mqvalues that agree (e.g., are closer to) nq. That is, the higher the bit rate (e.g., the lower the QP), the higher the number of chunks and the closer the number of chunks is to the number of potential parity leaders. In an example, for low QP values, the number of chunks will equal the number of potential parity leaders (e.g., mq= nqOn the other hand, in low bit rate cases (e.g., high QP), whereas nqmay be 5, for example, the chunk numbers determining function may use only one chunk (e.g., mq= 1). That is, all coefficients will be considered to be one group.
[0094] In an example, the chunk numbers determining function may be a machine learning model that is trained to output an optimal number of chunks based on the QP.
[0095] As further described below, one parity leader will be associated with each chunk and parity encoding will be performed with respect to the parity leaders associated with chunks. As such, mqnumber of bits will be saved since the sign bits associated with the mqparity leaders need not be encoded.
[0096] At 606, the indexes S={0, 1, . . . , N-l } of the coefficients in the vector q (or, equivalently, the coefficients themselves) are partitioned into mqnon-interesting chunks (e.g., subsets), 51?S2, . . . , Sm, whose union is the set S of indexes.
[0097] In an example partitioning, the first mqparity leaders can be lexicographically assigned to chunks. Then the remaining coefficients can be assigned to the chunks so that each chunk has roughly the same number of coefficients. In an example, the remaining coefficients can be assigned by lexicographically ordering them and sequentially assigning them to chunks so that the first remaining coefficient is added to the first chunk, the secondcoefficient to the second chunk 2, and so on. Once a coefficient is added the mnthchunk, assigning of coefficients to chunks resumes with the first chunk. To illustrate, assume that q=[3, 7, 2, -5, 9, 1, 6], T=5, and mq= 2. Thus, the potential parity leaders are [7, -5, 9, 6] and the two chunks can be [7, 3, 9, 6] and [-5, 2, 1]. Other partitionings (e.g., chunkifications) are possible so long as each chunk includes only one parity leader.
[0098] At 608, parity encoding in performed for each chunk to encode the coefficient magnitudes in the compressed bitstream. The same following steps can be performed for each chunk.
[0099] For a chunk of the chunks, a sum of the coefficients in the chunk is obtained. It is then determined whether the sum satisfies a parity constraint (e.g., condition). The encoder and decoder use the same parity constraint. The parity constraint can be an evenness parity constraint or an oddness parity constraint so long as the same parity constraint is used by the encoder and the decoder. The evenness parity constraint indicates that the sum of the coefficients of the chunk is an even number; and the oddness parity constraint indicates that the sum of the coefficients is an odd number.
[0100] If the sum of the coefficients does not satisfy a parity constraint, then at least one of the coefficients of the chunk is modified (e.g., increased or decreased in value) so that the parity constraint is satisfied. Identifying the at least one of the coefficients to be modified includes finding the least cost shift that will establish the parity constraint by shifting each coefficient qLby an odd number (e.g., 1, -1, 3, -3, or some other odd value) to obtain qt*, while ensuring that if |q < T then its shift also satisfies |qt* | < T, or if |q > T then its shift also satisfies > T. The coefficient that obtains the least cost shift can be found by using standard rate-distortion optimization metrics. It is noted that the parity leader of the chunk can itself be shifted.
[0101] The quantized transform coefficients can now be encoded in the compressed bitstream. However, instead of encoding the original values of the parity leaders (i.e., the qcvalues), their parity-encoded values (i.e., the qc' values) are encoded according to one of equations (4)-(7), depending on whether the threshold T is selected to be odd or even and whether the parity constraint is the evenness or the oddness parity constraint.
[0102] To illustrate, and continuing from the above example, the first sum of coefficients of the first chunk [7, 3, 9, 6] is 25 (an odd value) and the second sum of the second chunk [-5, 2, 1] is -2 (an even value). Thus, at least one of the coefficients of the first chunk is shifted. For the sake of this example, assume that the 3 will be shifted to 4. Thus, the vector q ismodified to [4, 7, 2, -5, 9, 1, 6]. Since the parity leader 7 satisfies the condition qc> T of equation (4), its actual magnitude (i.e., 7) is encoded. However, since the parity leader -5 satisfies the condition qc< —T of equation (4), the value —(—5) + 1 = 6 is encoded in its place. As such, instead of encoding the values [3, 7, 2, -5, 9, 1, 6], the values [4, 7, 2, 6, 9, 1, 6] are encoded instead. It is noted that it is the magnitudes of these values that are encoded at this step.
[0103] Under certain conditions, the values of certain potential parity leaders that are not used in partitioning the vector q can be changed (e.g., shifted) across the threshold T: If the f (nq— 1)= mqthen the last parity leader’ s magnitude can be allowed to transition across the threshold T when value shifting is carried out; likewise, if f(nq— 2) == mq then the 'ast tw0P^ity leaders can be allowed to transition across T and; and so on.
[0104] Under certain conditions, the values of certain non-potential parity leaders can be changed (e.g., shifted) across the threshold T: If f(nq+ 1)then any one of the non-parity leader coefficients following the mqthcoefficient in lexicographic order can be allowed to transition (e.g., shift) across the threshold T; likewise, if f(nq+ 2) =mq, then any two non-parity leader coefficients following the mqthcoefficient in lexicographic order can be allowed to transition across the threshold T; and so on.
[0105] At 610, coefficient signs are encoded in the compressed bitstream. More specifically, during entropy coding the signs of the parity leaders are skipped (i.e. not encoded) since they can be deduced at the decoder. Thus, no sign information are coded for the second (e.g., 7) and the fifth (e.g., 9) quantized transform coefficients of the vector q.
[0106] FIG. 7 is a flowchart of a technique 700 for decoding a quantized transform block. The technique 700 decodes the quantized transform coefficients of the quantized transform block from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5. The technique 700 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 700. The technique 700 may be implemented in whole or in part in the entropy decoding stage 502 of the decoder 500 of FIG. 5. The technique 700 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
[0107] At 702, a vector q' of values indicative of magnitudes of quantized transform coefficients is decoded from the compressed bitstream. As indicated above, the magnitudes can be encoded in the compressed bitstream in any number of ways, such as into planes. Thus, and continuing with the example from above, the technique 700 may decode the q' = [4, 7, 2, 6, 9, 1, 6] from the compressed bitstream.
[0108] At 704, potential parity leaders are identified. The potential parity leaders can be identified as described above with respect to 602 of FIG. 6. In this example, the potential parity leaders are [7, 6, 9, 6]. It is noted that the locations of the potential parity leaders in the vector q' are the same as the locations of the potential parity leaders identified by the encoder in the vector q.
[0109] At 706, a number of chunks is determined and, at 708, the vector q' is partitioned into chunks. Determining the number of chunks and partitioning the vector q' into the chunks can be as described with respect to 604 and 606 of FIG. 6, respectively. In this example, the number of chunks is 2 (since the encoder and the decoder use the same chunk numbers determining function) and the chunks of the q' can be a first chunk [7, 4, 9, 6] and a second chunk [6, 2, 1].
[0110] At 710, the technique 700 performs parity decoding on the chunks. That is, for each parity leader of the chunks, the technique 700 determines whether the decoded value should be adjusted and, if so, adjusts the value. Since the sum of the values in the first chunk is even (e.g., 7+4+9+6=26), then no adjustments are necessary to the parity leader (e.g., 7). Said another way, the parity leader 7, being odd and greater than T=5, does not need any modification as it was not modified (mapped to a different value) for encoding. On the other hand, that the sum of the values in the second chunk is odd (e.g., 6+2+1 =9), and since the parity leader (e.g., qc' = 6) is even and greater than the threshold T=5 suggests that the value 6 was encoded from a negative value. The original value qccan be obtained, based on equation (4) as qc= 1 — qc' = 1 — 6 = — 5. As such, the vector of magnitudes of the quantized transform coefficients is recovered as [4, 7, 2, -5, 9, 1, 6].
[0111] At 712, the technique 700 decodes signs for all but the parity leaders used in the chunks. That is, the technique 700 does not decode signs for the first mqpotential parity leaders. The signs of the parity leaders 7 and -5 are inferred to be positive and negative. The quantized transform coefficients are then obtained based on the decoded magnitudes and the decoded signs.
[0112] FIG. 8 is a flowchart of a technique 800 for encoding a vector of quantized transform coefficients. The technique 800 can be used to encode the quantized transform coefficients of a quantized transform block in a compressed bitstream, such as the compressed bitstream 420 of FIG. 4. The quantized transform coefficients can be arranged into a vector q.
[0113] The technique 800 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 800. The technique 800 may be implemented in whole or in part in the entropy encoding stage 408 of the encoder 400 of FIG. 4. The technique 800 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
[0114] At 802, the quantized transform coefficients are partitioned into chunks. Each chunk includes one parity leader. A parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T. The chunks include different parity leaders. The threshold T can be set (e.g., selected) based on a quantization parameter. The number of the chunks can be inversely related to the quantization parameter.
[0115] For ease of reference, the quantized transform coefficients other than the different parity leaders are referred to as the “remaining coefficients.” The quantized transform coefficients can be partitioned into chunks in a way that the remaining coefficients are assigned to the chunks such that the chunks have roughly the same number of coefficients. In an example, the remaining coefficients can be assigned to the chunks by lexicographically ordering them and sequentially assigning them to the chunks. Stated another way, the remaining coefficients are distributed among chunks to maintain a roughly equal number of coefficients in each chunk.
[0116] For each of the chunks, the technique 800 performs the steps 806-810. As such, at 804, it is determined whether more chunks remain to be processed. If not, the technique 800 proceeds to 812; otherwise, the technique 800 proceeds to 806 to start the processing of the next chunk.
[0117] At 806, it is determined whether a sum of coefficients in the chunk satisfies a parity constraint. The parity constraint can be an evenness parity constraint or an oddness parity constraint, as described above. At 808, it is determined whether the sum satisfies the parity constraint. If the parity constraint is not satisfied, the technique 800 proceeds to 810;otherwise, the technique 800 proceeds back to 804. At 810, at least one coefficient of the chunk is modified so that the sum satisfies the parity constraint.
[0118] The at least one coefficient of the chunk can be modified by increasing or decreasing its value. Modifying the at least one coefficient of the chunk can include shifting at least some (e.g., each) of the quantized transform coefficients qi by a number to obtain qi*, while ensuring that if lqd<T then a shift of qi also satisfies lqi*l<T , or if lqd>T then the shift of the qi also satisfies lqi*l>T, where z is used to index the at least some of the quantized transform coefficients; identifying the at least one coefficient that results in the smallest cost shift, such as based on rate-distortion analysis; and modifying the at least one coefficient. Stated another way, if the sum of coefficients within a chunk fails to meet the required parity constraint, a coefficient within the chunk is adjusted. This adjustment may involve increasing or decreasing the value of the coefficient to achieve the target parity. By altering the coefficient’s value in this minimal way, the encoding process adheres to the parity constraint without introducing significant distortion into the quantized transform block.
[0119] At 812, the quantized transform coefficients are encoded in the compressed bitstream. Encoding the quantized transform coefficients includes, at 814, encoding parity encoded values of the parity leaders. A parity encoded value for a parity leader is obtained according to specific equations depending on the threshold T and the parity constraint. In an example, the respective parity encoded values for the parity leaders are derived (e.g., the specific equations implement) a range folding technique that maps negative values of coefficients to positive odd values. At 816, respective coefficient signs of all of the quantized transform coefficients except for respective signs of the parity leaders are encoded in the compressed bitstream.
[0120] FIG. 9 is a flowchart of a technique 900 for decoding a quantized transform block. The technique 900 decodes the quantized transform coefficients of the quantized transform block from a compressed bitstream, such as the compressed bitstream 420 of FIG. 5. The technique 900 can be implemented, for example, as a software program that may be executed by computing devices such as transmitting station 102 or receiving station 106. The software program can include machine-readable instructions that may be stored in a memory such as the memory 204 or the secondary storage 214, and that, when executed by a processor, such as CPU 202, may cause the computing device to perform the technique 900. The technique 900 may be implemented in whole or in part in the entropy decoding stage 502 of the decoder 500 of FIG. 5. The technique 900 can be implemented using specialized hardware or firmware. Multiple processors, memories, or both, may be used.
[0121] At 902, a vector of values indicative of magnitudes of quantized transform coefficients is decoded from a compressed bitstream. At 904, the vector is partitioned into chunks that include respective parity leaders. A parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T. In an example, the threshold set may be preset (e.g., to 3). In an example, the threshold T may be set based on a quantization parameter, which may be decoded from the compressed bitstream. The number of the chunks can be inversely related to the quantization parameter. In an example, the threshold T may be decoded from the compressed bitstream. The number of the chunks may be inversely related to a quantization parameter. As such, the number of the chunks can be interpreted (e.g., set) based on the quantization parameter.The chunks include the respective parity leaders, and remaining coefficients are assigned to the chunks so that the chunks have roughly the same number of coefficients. Each chunk can include one of the respective parity leaders. Said another way, the decoder interprets the structure of the chunks as including respective parity leaders, with the remaining coefficients distributed across the chunks such that each chunk contains roughly the same (e.g., exactly the same) number of coefficients. Said another way, the technique 900 interprets the structure of the chunks as including respective parity leaders, with the remaining coefficients distributed across the chunks such that each chunk contains roughly the same number of coefficients. The remaining coefficients can be assigned to the chunks based on a lexicographical ordering and sequential assignment performed during an encoder.
[0122] At 906, for each parity leader, a magnitude of a corresponding quantized transform coefficient is obtained based on a sum of the coefficients of the chunk. In an example, the technique 900 obtains the magnitude of the corresponding quantized transform coefficient from the compressed bitstream using a range-folding technique. For example, the technique 900 obtains the magnitude by interpreting a range-folded value decoded from the compressed bitstream using a range-folding technique that maps negative values of coefficients to positive odd values. As such, the magnitude of the corresponding quantized transform coefficient can be said to be decoded using a range-folding technique.
[0123] At 908, respective signs for coefficients other than the parity leaders are decoded from the compressed bitstream. Decoding the respective signs for coefficients other than the parity leaders can include determining whether parity encoding is applied for a current coefficient; if parity encoding is applied to the current coefficient, not decoding a sign of the current coefficient; and if parity encoding is not applied to the current coefficient, decoding the sign of the current coefficient. Determining whether parity encoding is applied for thecurrent coefficient can be based on the threshold T. For example, if the magnitude of the current coefficient is greater than or equal to the threshold T, then it can be determined that the current coefficient is parity encoded.
[0124] At 910, respective signs of quantized transform coefficients corresponding to the parity leaders are inferred. At 912, the quantized transform coefficients are obtained based on the decoded magnitudes and the respective signs (i.e., the decoded and the inferred signs).
[0125] For simplicity of explanation, the techniques 600, 700, 800, and 900 of FIGS. 6, 7, 8, and 9, respectively, are each depicted and described as respective series of steps or operations. However, the steps or operations in accordance with this disclosure can occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, not all illustrated steps or operations may be required to implement a technique in accordance with the disclosed subject matter.
[0126] It will be appreciated that the present disclosure may include any one and up to all of the following examples.
[0127] Example 1 is a method that includes decoding a vector of values indicative of magnitudes of quantized transform coefficients from a compressed bitstream; partitioning the vector into chunks that include respective parity leaders, wherein a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T; obtaining, for each parity leader, a magnitude of a corresponding quantized transform coefficient based on a sum of the coefficients of the chunk that includes the parity leader; decoding, from the compressed bitstream, respective signs for quantized transform coefficients other than the parity leaders; inferring respective signs of quantized transform coefficients corresponding to the parity leaders; and obtaining the quantized transform coefficients based on the magnitudes and the respective signs.
[0128] Example 2 is the method of Example 1, wherein the threshold T is set based on a quantization parameter.
[0129] Example 3 is the method of Example 1, wherein a number of the chunks is inversely related to a quantization parameter.
[0130] Example 4 is the method of any one of Examples 1 to 3, wherein obtaining the magnitude of the corresponding quantized transform coefficient includes decoding the magnitude of the corresponding quantized transform coefficient using a range-folding technique.
[0131] Example 5 is the method of any one of Examples 1 to 4, wherein the chunks include the respective parity leaders, and remaining coefficients of the quantized transformcoefficients other than the respective parity leaders are assigned to the chunks so that the chunks have roughly the same number of quantized transform coefficients.
[0132] Example 6 is the method of Example 5, wherein the remaining coefficients are assigned to the chunks based on a lexicographical ordering and sequential assignment performed by an encoder.
[0133] Example 7 is the method of any one of Examples 1 to 6, wherein decoding the respective signs for quantized transform coefficients other than the parity leaders includes determining whether parity encoding is applied for a current coefficient, not decoding a sign of the current coefficient if parity encoding is applied, and decoding the sign of the current coefficient if parity encoding is not applied.
[0134] Example 8 is the method of Example 7, wherein determining whether parity encoding is applied for the current coefficient includes determining whether a magnitude of the current coefficient is greater than or equal to the threshold T.
[0135] Example 9 is a method for encoding a vector of quantized transform coefficients that includes partitioning the quantized transform coefficients into chunks, wherein each chunk includes one parity leader, wherein a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T, and wherein the chunks include different parity leaders; for each chunk, determining whether a sum of coefficients in the chunk satisfies a parity constraint, and, responsive to determining that the sum does not satisfy the parity constraint, modifying at least one coefficient of the chunk so that the sum satisfies the parity constraint; and encoding the quantized transform coefficients in a compressed bitstream by obtaining respective parity encoded values for the parity leaders according to specific equations depending on the threshold T and the parity constraint, encoding the respective parity encoded values, and encoding respective coefficient signs of all of the quantized transform coefficients except for respective signs of the parity leaders.
[0136] Example 10 is the method of Example 9, wherein the quantized transform coefficients other than the different parity leaders constitute remaining coefficients, and wherein the remaining coefficients are assigned to the chunks such that the chunks have roughly the same number of coefficients.
[0137] Example 11 is the method of Example 10, wherein the remaining coefficients are assigned to the chunks by lexicographically ordering the remaining coefficients and sequentially assigning the remaining coefficients to the chunks.
[0138] Example 12 is the method of any one of Examples 9 to 11, wherein modifying the at least one coefficient includes increasing or decreasing a value of the at least one coefficient.
[0139] Example 13 is the method of any one of Examples 9 to 11, wherein modifying the at least one coefficient of the chunk includes shifting at least some of the quantized transform coefficients qi by a number to obtain qi*, while ensuring that if lqd<T then a shift of qi also satisfies lqi*l<T , or if lqd>T then the shift of the qi also satisfies lqi*l>T, wherein z is an index of the at least some of the quantized transform coefficients; identifying the at least one coefficient that results in a smallest cost shift based on rate-distortion analysis; and modifying the at least one coefficient.
[0140] Example 14 is the method of any one of Examples 9 to 13, wherein the threshold T is set based on a quantization parameter.
[0141] Example 15 is the method of any one of Examples 9 to 14, wherein a number of the chunks is inversely related to a quantization parameter.
[0142] Example 16 is the method of any one of Examples 9 to 15, wherein the parity constraint is an evenness parity constraint or an oddness parity constraint.
[0143] Example 17 is the method of any one of Examples 9 to 16, wherein obtaining the respective parity encoded values for the different parity leaders uses a range-folding technique that maps negative values of coefficients to positive odd values.
[0144] Example 18 is a device that includes a processor configured to perform the method of any one of Examples 1 to 17.
[0145] Example 19 is a device that includes a memory and a processor, the processor configured to execute instructions stored in the memory to perform the method of any one of Examples 1 to 17.
[0146] Example 20 is a non-transitory computer-readable storage medium that includes executable instructions that, when executed by a processor, facilitate performance of operations, including operations that perform the method of any one of Examples 1 to 17.
[0147] Example 21 is a non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is configured for decoding by the method of any one of Examples 1 to 8.
[0148] Example 22 is a non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is generated by an encoder performing the method of any one of Examples 9 to 17.- l-
[0149] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it is to be understood that encoding and decoding, as those terms are used in the claims, could mean compression, decompression, transformation, or any other processing or change of data.
[0150] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” is not necessarily to be construed as being preferred or advantageous over other aspects or designs. Rather, use of the word “example” is intended to present concepts in a concrete fashion. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise or clearly indicated otherwise by the context, the statement “X includes A or B” is intended to mean any of the natural inclusive permutations thereof. That is, if X includes A; X includes B; or X includes both A and B, then “X includes A or B” is satisfied under any of the foregoing instances. In addition, the articles “a” and “an” as used in this application and the appended claims should generally be construed to mean “one or more,” unless specified otherwise or clearly indicated by the context to be directed to a singular form. Moreover, use of the term “an implementation” or the term “one implementation” throughout this disclosure is not intended to mean the same embodiment or implementation unless described as such.
[0151] Implementations of the transmitting station 102 and / or the receiving station 106 (and the algorithms, methods, instructions, etc., stored thereon and / or executed thereby, including by the encoder 400 and the decoder 500) can be realized in hardware, software, or any combination thereof. The hardware can include, for example, computers, intellectual property (IP) cores, application- specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors, or any other suitable circuit. In the claims, the term “processor” should be understood as encompassing any of the foregoing hardware, either singly or in combination. The terms “signal” and “data” are used interchangeably. Further, portions of the transmitting station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.
[0152] Further, in one aspect, for example, the transmitting station 102 or the receiving station 106 can be implemented using a general purpose computer or general purpose processor with a computer program that, when executed, carries out any of the respective methods, algorithms, and / or instructions described herein. In addition, or alternatively, forexample, a special purpose computer / processor can be utilized which can contain other hardware for carrying out any of the methods, algorithms, or instructions described herein.
[0153] The transmitting station 102 and the receiving station 106 can, for example, be implemented on computers in a video conferencing system. Alternatively, the transmitting station 102 can be implemented on a server, and the receiving station 106 can be implemented on a device separate from the server, such as a handheld communications device. In this instance, the transmitting station 102, using an encoder 400, can encode content into an encoded video signal and transmit the encoded video signal to the communications device. In turn, the communications device can then decode the encoded video signal using a decoder 500. Alternatively, the communications device can decode content stored locally on the communications device, for example, content that was not transmitted by the transmitting station 102. Other suitable transmitting and receiving implementation schemes are available. For example, the receiving station 106 can be a generally stationary personal computer rather than a portable communications device, and / or a device including an encoder 400 may also include a decoder 500.
[0154] Further, all or a portion of implementations of the present disclosure can take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transport the program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable mediums are also available.
[0155] The above-described embodiments, implementations, and aspects have been described in order to facilitate easy understanding of this disclosure and do not limit this disclosure. On the contrary, this disclosure is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation as is permitted under the law so as to encompass all such modifications and equivalent arrangements.
Claims
What is claimed is:
1. A method, comprising: decoding a vector of values indicative of magnitudes of quantized transform coefficients from a compressed bitstream; partitioning the vector into chunks that include respective parity leaders, wherein a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T; for each parity leader, obtaining a magnitude of a corresponding quantized transform coefficient based on a sum of the coefficients of the chunk that includes the parity leader; decoding, from the compressed bitstream, respective signs for quantized transform coefficients other than the parity leaders; inferring respective signs of quantized transform coefficients corresponding to the parity leaders; and obtaining the quantized transform coefficients based on the magnitudes and the respective signs.
2. The method of claim 1, wherein the threshold T is set based on a quantization parameter.
3. The method of claim 1, wherein a number of the chunks is inversely related to a quantization parameter.
4. The method of any one of claims 1 to 3, wherein obtaining the magnitude of the corresponding quantized transform coefficient comprises: decoding the magnitude of the corresponding quantized transform coefficient using a range-folding technique.
5. The method of any one of claims 1 to 4, wherein the chunks include the respective parity leaders, and remaining coefficients of the quantized transform coefficients other than the respective parity leaders are assigned to the chunks so that the chunks have roughly a same number of quantized transform coefficients.
6. The method of claim 5, wherein the remaining coefficients are assigned to thechunks based on a lexicographical ordering and sequential assignment performed by an encoder.
7. The method of any one of claims 1 to 6, wherein decoding the respective signs for quantized transform coefficients other than the parity leaders comprises: determining whether parity encoding is applied for a current coefficient; if parity encoding is applied to the current coefficient, not decoding a sign of the current coefficient; and if parity encoding is not applied to the current coefficient, decoding the sign of the current coefficient.
8. The method of claim 7, wherein determining whether parity encoding is applied for the current coefficient comprises: determining whether a magnitude of the current coefficient is greater than or equal to the threshold T.
9. A method for encoding a vector of quantized transform coefficients, comprising: partitioning the quantized transform coefficients into chunks, wherein each chunk includes one parity leader, wherein a parity leader is a coefficient of the quantized transform coefficients whose magnitude is greater than or equal to a threshold T, and wherein the chunks include different parity leaders; for each chunk: determining whether a sum of coefficients in the chunk satisfies a parity constraint; and responsive to determining that the sum does not satisfy the parity constraint, modifying at least one coefficient of the chunk so that the sum satisfies the parity constraint; and encoding the quantized transform coefficients in a compressed bitstream by: obtaining respective parity encoded values for the parity leaders according to specific equations depending on the threshold T and the parity constraint; and encoding the respective parity encoded values; and encoding respective coefficient signs of all of the quantized transform coefficients except for respective signs of the parity leaders.
10. The method of claim 9, wherein the quantized transform coefficients other than the different parity leaders constitute remaining coefficients, and wherein the remaining coefficients are assigned to the chunks such that the chunks have roughly a same number of coefficients.
11. The method of claim 10, wherein the remaining coefficients are assigned to the chunks by lexicographically ordering the remaining coefficients and sequentially assigning the remaining coefficients to the chunks.
12. The method of any one of claims 9 to 11, wherein modifying the at least one coefficient comprises: increasing or decreasing a value of the at least one coefficient.
13. The method of any one of claims 9 to 11, wherein modifying the at least one coefficient of the chunk comprises: shifting at least some of the quantized transform coefficients qi by a number to obtain qi*, while ensuring that if lqd<T then a shift of qi also satisfies lqi*l<T , or if lqd>T then the shift of the qi also satisfies lqi*l>T, wherein z is an index of the at least some of the quantized transform coefficients; and identifying the at least one coefficient that results in a smallest cost shift based on rate-distortion analysis; and modifying the at least one coefficient.
14. The method of any one of claims 9 to 13, wherein the threshold T is set based on a quantization parameter.
15. The method of any one of claims 9 to 14, wherein a number of the chunks is inversely related to a quantization parameter.
16. The method of any one of claims 9 to 15, wherein the parity constraint is an evenness parity constraint or an oddness parity constraint.
17. The method of any one of claims 9 to 16, wherein obtaining the respectiveparity encoded values for the different parity leaders uses a range folding technique that maps negative values of coefficients to positive odd values.
18. A device, comprising: a processor that is configured to perform the method of any one of claims 1-17.
19. A device, comprising: a memory; and a processor, the processor configured to execute instructions stored in the memory to perform the method of any one of claims 1-17.
20. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate performance of operations, comprising operations that perform the method of any one of claims 1-17.
21. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is configured for decoding by the method of any one of claims 1-8.
22. A non-transitory computer-readable storage medium having stored thereon an encoded bitstream, wherein the encoded bitstream is generated by an encoder performing the method of any one of claims 9-17.
Citation Information
Patent Citations
Multiple sign bit hiding within a transform unit
EP3644611A1
Method and apparatus for harmonizing multiple sign bit hiding and residual sign prediction
US20200404257A1
Method and apparatus for detecting blocks suitable for multiple sign bit hiding
US20200404308A1