Context model and scan order for joint design of transform coefficient code processing
By using wavefront scanning order and probability distribution entropy code processing in video encoding, the quantized transform blocks are optimized to solve the problem of low encoding efficiency of video streams in the prior art, and higher compression efficiency and coding performance are achieved.
Patent Information
- Application Number
- CN202280101535.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-02
- Filing Date
- 2022-12-15
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively reduce the amount of data and improve the compression efficiency when encoding a digital video stream, especially when processing quantized transform blocks.
The quantized transform block is coded using wavefront scanning order, the appropriate probability distribution is selected for entropy code processing, and the context coefficients are weighted to optimize code processing.
Improves the compression efficiency of the video stream, reduces the bit rate, and enhances the performance of the encoder, especially when processing transform blocks.
Smart Images

Figure CN120153653A_ABST
Abstract
Description
Background Art
[0001] A digital video stream can represent video using a sequence of frames or still images. Digital video can be used for various applications, including, for example, video conferencing, high-definition video entertainment, video advertising, or sharing of user-generated video. A digital video stream can contain a large amount of data and consume a significant amount of computing or communication resources of a computing device for processing, transmitting, or storing the video data. Various methods have been proposed to reduce the amount of data in a video stream, including compression and other coding techniques.
[0002] Coding based on motion estimation and compensation can be performed by decomposing a frame or image into blocks that are predicted based on one or more prediction blocks of a reference frame. The difference between a block and a prediction block (i.e., the residual error) is compressed and encoded in a bitstream. A decoder uses the difference and the reference frame to reconstruct the frame or image. Summary of the Invention
[0003] A system of one or more computers can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which, in operation, cause the system to perform these actions. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by a data processing apparatus, cause the apparatus to perform these actions. One general aspect includes a method for code processing of a quantized transform block. The method further includes selecting a wavefront scan order for code processing quantized transform coefficients of the quantized transform block, where the quantized transform block has a size N×N, and where the wavefront scan order is such that for at least one x and at least one y, where 2≤x<N and 2≤y<N, positions (x, y−1), (x, y−2), and (x, y−3) are sequentially code processed, and positions (x−1, y), (x−2, y), and (x−3, y) are sequentially code processed. The method further includes selecting a probability distribution for code processing quantized transform coefficients among the quantized transform coefficients, where a context model for selecting the probability distribution includes at least two immediate neighbors of the quantized transform coefficient in the wavefront scan order. The method further includes entropy coding the quantized transform coefficients using the probability distribution. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0004] Implementations may include one or more of the following features. A method, wherein the wavefront scan order is characterized by coding processed quantization transform coefficients of a flipped L-shaped region, and wherein the first quantized transform coefficients along the first axis of the flipped L-shaped region are coded processed, followed by coding processing of the second quantized transform coefficients along the second axis of the flipped L-shaped region. A first weight may be used with a first context coefficient of the same dimension as the quantized transform coefficients in the flipped L-shaped region, and a second weight lower than the first weight may be used with a second context coefficient not in the flipped L-shaped region. The context coefficients include a first context coefficient and a second context coefficient, the first context coefficient being the immediate neighbor of the quantized transform coefficient in the wavefront scan order, the second context coefficient not being an immediate neighbor, and wherein the first weight used with the first context coefficient may be greater than the second weight used with the second context coefficient. The quantized transform coefficients may be located on the diagonal of the quantized transform block, and the method may further include obtaining a context that is the sum of the context coefficients of the quantized transform coefficients. The number of immediate neighbors of the quantized transform coefficients used as context coefficients depends on the position of the quantized transform coefficients. The flipped L-shaped region includes the quantized transform coefficients at the position (p, p) of the quantized transform block and all other quantized transform coefficients having coordinates (p, y) and (x, p) such that y ≤ p and x ≤ p. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.
[0005] An overall aspect includes an apparatus for coding processing a quantized transform block. The apparatus further includes a processor configured to select a wavefront scan order for coding processed quantized transformed coefficients of the quantized transform block, wherein the quantized transform block has a size N×N, wherein the wavefront scan order is such that for at least one x and at least one y, where 2 ≤ x < N and 2 ≤ y < N, the positions (x, y - 1), (x, y - 2) and (x, y - 3) are coded processed in sequence, and the positions (x - 1, y), (x - 2, y) and (x - 3, y) are coded processed in sequence. The processor is further configured to select a probability distribution for coding processed the quantized transform coefficients among the quantized transform coefficients. The context model for selecting the probability distribution may include at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order. The processor may also be configured to perform entropy coding on the quantized transformed coefficients using the probability distribution. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0006] The implementation may include one or more of the following features. An apparatus, wherein the wavefront scanning order is characterized by coding the corresponding quantized transform coefficients of the flipped L-shaped region, and wherein the first quantized transform coefficients along the first axis of the flipped L-shaped region are coded, and subsequently the second quantized transform coefficients along the second axis of the flipped L-shaped region are coded. The processor is further configured to obtain a context that is a weighted combination of context coefficients of the quantized transform coefficients. A first weight is used with a first context coefficient along the same dimension as the quantized transform coefficient in the flipped L-shaped region, and a second weight that may be lower than the first weight is used with a second context coefficient that is not in the flipped L-shaped region. The processor is further configured to obtain a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein the context coefficients include a first context coefficient and a second context coefficient, the first context coefficient being the immediate neighbor of the quantized transform coefficient in the wavefront scanning order, the second context coefficient not being an immediate neighbor, and wherein the first weight used with the first context coefficient is greater than the second weight used with the second context coefficient. The quantized transform coefficients may be located on the diagonal of the quantized transform block, and the processor may be further configured to obtain a context that is the sum of the context coefficients of the quantized transform coefficients. The number of immediate neighbors of the quantized transform coefficients used as context coefficients may depend on the location of the quantized transform coefficients. The flipped L-shaped region may include the quantized transform coefficient at the location (p,p) of the quantized transform block and all other quantized transform coefficients having coordinates (p,y) and (x,p) such that y ≤ p and x ≤ p. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.
[0007] One general aspect includes a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can include executable instructions that facilitate performing operations for coding processed quantized transform blocks. The operations can include selecting a wavefront scan order for coding processed quantized transform coefficients of a quantized transform block, where the quantized transform block has a size N×N, and where the wavefront scan order is such that for at least one x and at least one y, where 2≤x<N and 2≤y<N, positions (x, y-1), (x, y-2), and (x, y-3) are coded in sequence, and positions (x-1, y), (x-2, y), and (x-3, y) are coded in sequence. The operations can also include selecting a probability distribution for coding processed quantized transform coefficients among the quantized transform coefficients, where a context model for selecting the probability distribution includes at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order. The operations can further include entropy coding the quantized transform coefficients using the probability distribution. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0008] Implementations can include one or more of the following features. A non-transitory computer-readable storage medium, where the wavefront scan order is characterized by coding processed corresponding quantized transform coefficients of a flipped L-shaped region, and where a first quantized transform coefficient along a first axis of the flipped L-shaped region is coded, followed by coding a second quantized transform coefficient along a second axis of the flipped L-shaped region. A first weight is used with a first context coefficient along a dimension of the flipped L-shaped region that is the same as the dimension of the quantized transform coefficient, and a second weight that can be lower than the first weight is used with a second context coefficient that is not in the flipped L-shaped region. The context coefficients include a first context coefficient and a second context coefficient, the first context coefficient being an immediate neighbor of the quantized transform coefficient in the wavefront scan order, the second context coefficient not being an immediate neighbor, and where the first weight used with the first context coefficient is greater than the second weight used with the second context coefficient. The quantized transform coefficient can be located on a diagonal of the quantized transform block, and the operation further includes obtaining a context that is a sum of context coefficients of the quantized transform coefficient. The number of immediate neighbors of the quantized transform coefficient used as context coefficients depends on the position of the quantized transform coefficient. The flipped L-shaped region includes the quantized transform coefficient at position (p,p) of the quantized transform block and all other quantized transform coefficients having coordinates (p,y) and (x,p) such that y≤p and x≤p. Implementations of the described techniques can include hardware, a method or process, or computer software on a computer-accessible medium.
[0009] These and other aspects of the disclosure are set forth in the following detailed description of the embodiments, the appended claims, and the drawings. It is understood that the aspects can be implemented in any convenient form. For example, the aspects can be implemented by a suitable computer program that can be carried on a suitable carrier medium, which can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). The aspects can also be implemented using a suitable device, which can take the form of a programmable computer running a computer program arranged to implement the methods and / or techniques disclosed herein. The aspects can be combined such that features described in the context of one aspect can be implemented in another aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The description herein refers to the drawings described below, in which like reference numerals refer to like parts throughout the several views.
[0011] Figure 1 is a schematic diagram of a video encoding and decoding system.
[0012] Figure 2 is a block diagram of an example of a computing device that can implement a transmitting station or a receiving station.
[0013] Figure 3 is a diagram of a video stream to be encoded and then decoded.
[0014] Figure 4 is a block diagram of an encoder according to an implementation of the disclosure.
[0015] Figure 5 is a block diagram of a decoder according to an implementation of the disclosure.
[0016] Figure 6 is a flowchart of a technique for performing code processing on quantized transform blocks using a wavefront scan order.
[0017] Figure 7 illustrates an example of performing code processing on the position of the end of a block.
[0018] Figure 8 illustrates an example of performing code processing on quantized transform coefficients using a wavefront scan order.
[0019] Figure 9 illustrates an example of a backward (i.e., reverse) zigzag scan order and context selection.
[0020] Figure 10 illustrates an example of grid optimization for determining the quantization level of the transform coefficients of a transform block.
[0021] Figure 11It is a flowchart of a technique for performing code processing on quantized transform blocks using a wavefront scanning order. Detailed implementation
[0022] As mentioned above, compression schemes related to performing code processing on a video stream may include: decomposing an image into blocks and using one or more techniques for restricting the information included in the output to generate a digital video output bitstream. The received encoded bitstream may be decoded to recreate the blocks and the source image from the restricted information. Encoding a video stream or a portion thereof (such as a frame or a block) may include using temporal and spatial similarities in the video stream to improve code processing efficiency. For example, the current block of the video stream may be encoded based on identifying the difference (residual) between the pixel values of a previously code processed pixel and the pixel values in the current block. In this way, only the residual and the parameters used to generate the residual need to be added to the encoded bitstream. The residual may be encoded using a lossy quantization step.
[0023] As further described below, the residual block may be in the pixel domain. The residual block may be transformed into the frequency domain, resulting in a transform block of transform coefficients. The transform coefficients may be quantized, resulting in a quantized transform block of quantized transform coefficients. The quantized coefficients may be entropy encoded and added to the encoded bitstream. The decoder may receive the encoded bitstream and perform entropy decoding on the quantized transform coefficients to reconstruct the original block.
[0024] Entropy coding is a technique for lossless coding that relies on a probability model that models the distribution of values that occur in the encoded video bitstream. By using a probability model based on the measured or estimated distribution of values, entropy coding can reduce the number of bits required to represent the video data to near the theoretical minimum. In practice, the actual reduction in the number of bits required to represent the video data may depend on the accuracy of the probability model, the number of bits through which the coding is performed, and the computational accuracy of the fixed-point arithmetic used to perform the coding.
[0025] The probability model as used herein may be a lossless (entropy) coding or may be a parameter in a lossless (entropy) coding. An arithmetic coder (AC) may be used to encode the symbols (also referred to as syntax elements) corresponding to the transform coefficients losslessly. The model may be any parameter or method that affects the probability estimate for the purpose of entropy coding. In an example, a two-pass process may be used to learn the probabilities of the current frame. In another example, the model may define a certain context derivation method.
[0026] In an encoded video bitstream, many of the bits in the bitstream are used for one of two things: content prediction (e.g., inter-frame mode / motion vector code processing, intra-frame prediction mode code processing, etc.) or residual code processing (e.g., transform coefficients). An encoder can use techniques to reduce the number of bits spent on coefficient code processing.
[0027] In some codecs, to encode a quantized transform block, a scan order is selected for traversing the block according to that scan order. When accessing the quantized transform coefficients, a probability distribution is selected for coding the quantized transform coefficients. A context for selecting the probability distribution is determined (selected according to a context model). An indicator for the end-of-block coefficient (EOB) can also be coded. The EOB is the last non-zero quantized transform coefficient in the scan order. The quantized transform coefficients after the EOB in the scan order do not need to be coded because, by definition of the EOB, such coefficients are known to be zero.
[0028] The algorithms for encoding transform coefficients (including how to encode the EOB, how to traverse the quantized transform block, and the accuracy of the selected probability model) have a great impact on the compression efficiency, throughput, and memory consumption of the software and hardware implementations of the codec.
[0029] Implementations according to the present disclosure use a wavefront scan order to code (e.g., encode or decode) quantized transform coefficients. The wavefront scan order can reduce the bitrate and achieve higher compression efficiency. The techniques disclosed herein for coding quantized transform blocks using a wavefront scan order include techniques for signaling the EOB, techniques for scanning (e.g., traversing) the quantized transform coefficients of the block, and techniques for context model design.
[0030] The wavefront scan order is characterized by dividing the quantized transform block into flipped L-shaped regions. First, the transform coefficients along the first axis (e.g., the vertical axis) of the flipped L-shaped region are coded, and then the transform coefficients along the second axis (e.g., the horizontal axis) of the flipped L-shaped region are coded. The coding of the transform coefficients along each axis starts at the transform coefficient corresponding to the intersection of the two axes and proceeds in reverse order. The transform coefficient corresponding to the intersection of the two axes can be the transform coefficient along the diagonal of the quantized transform block.
[0031] Since there can be signal correlations in the transform coefficients, adjacent information (i.e., context) can help in coding each transform coefficient. Higher compression ratios can be achieved with AC when a good estimate of the symbol probability is available. AC can use the context to better estimate the probability of the quantized transform coefficients. Thus, a good design for context-aware transform coefficient coding is also described herein.
[0032] Additionally, and as further described below, in the process of obtaining a quantized transform block from a transform block, the encoder may use optimization techniques (such as lattice optimization) to jointly determine the quantization levels of the transform coefficients. Lattice optimization may perform best when there is a first-order correlation (explained below) between the transform coefficients used for context modeling. As further described below, the conventional scan order does not yield a first-order correlation. The wavefront scan order described herein can address problems such as these because it yields a first-order correlation for at least some (if not most) of the transform coefficients in the transform block.
[0033] Further details of the techniques for coding a transform block using the wavefront scan order are described herein first with reference to a system in which these techniques may be implemented. Figure 1 is a schematic diagram of a video encoding and decoding system 100. The transmitting station 102 may be, for example, a computer such as Figure 2 described with an internal hardware configuration. However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 may be distributed among multiple devices.
[0034] The network 104 may connect the transmitting station 102 and the receiving station 106 for encoding and decoding of a video stream. Specifically, the video stream may be encoded at the transmitting station 102, and the encoded video stream may be decoded at the receiving station 106. The network 104 may be, for example, the Internet. The network 104 may also be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network, or any other component that conveys the video stream from the transmitting station 102 to the receiving station 106 (in this example).
[0035] In one example, the receiving station 106 may be a computer such as Figure 2 described with an internal hardware configuration. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 may be distributed among multiple devices.
[0036] Other implementations of the video encoding and decoding system 100 are possible. For example, one implementation may omit the network 104. In another implementation, the video stream may be encoded and then stored for transmission to the receiving station 106 or any other device with memory at a later time. In one implementation, the receiving station 106 receives (e.g., via the network 104, computer bus, and / or some communication means) the encoded video stream and stores the video stream for later decoding. In an example implementation, the Real-Time Transport Protocol (RTP) is used for the transmission of the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used, e.g., the Hypertext Transfer Protocol (HTTP) video streaming protocol.
[0037] When used in a video conferencing system, for example, the transmitting station 102 and / or the receiving station 106 may include the ability to both encode and decode video streams as described below. For example, the receiving station 106 may be a video conferencing participant who receives an encoded video bitstream from a video conferencing server (e.g., the transmitting station 102) for decoding and viewing, and further encodes his or her own video bitstream and transmits it to the video conferencing server for other participants to decode and view.
[0038] Figure 2 is a block diagram of an example of a computing device 200 that may implement the transmitting station or the receiving station. For example, the computing device 200 may implement Figure 1 either or both of the transmitting station 102 and the receiving station 106. The computing device 200 may be in the form of a computing system including multiple computing devices, or in the form of a single computing device (e.g., a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.).
[0039] The CPU 202 in the computing device 200 may be a conventional central processing unit. Alternatively, the CPU 202 may be any other type of device or devices, existing or later developed, capable of manipulating or processing information. Although the disclosed implementations may be practiced using one processor (e.g., the CPU 202) as shown, advantages in speed and efficiency may be achieved by using more than one processor.
[0040] In an implementation, the memory 204 in the computing device 200 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that are accessed by the CPU 202 using the bus 212. The memory 204 can further include an operating system 208 and an application 210 that includes at least one program that permits the CPU 202 to execute the methods described herein. For example, the application 210 can include Applications 1 through N, which further include video codec applications that execute the techniques described herein, such as techniques for coding transform blocks using a wavefront scan order. The computing device 200 can also include auxiliary storage 214, which can be, for example, a memory card used with a mobile computing device. Because video communication sessions can contain a substantial amount of information, they can be stored in whole or in part in the auxiliary storage 214 and loaded into the memory 204 as needed for processing.
[0041] The computing device 200 can also include one or more output devices, such as a display 218. In one example, the display 218 can be a touch-sensitive display that combines the display with touch-sensitive elements operable to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. In addition to or instead of the display 218, other output devices can be provided that permit a user to program or otherwise use the computing device 200. When the output device is or includes a display, the display can be implemented in a variety of ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, or a light-emitting diode (LED) display (such as an organic LED (OLED) display).
[0042] The computing device 200 can also include an image sensing device 220, such as a camera or any other image sensing device 220, present now or developed later, that can sense an image, such as an image of a user operating the computing device 200, or communicate with such an image sensing device. The image sensing device 220 can be positioned such that it is pointed at the user operating the computing device 200. In an example, the position and optical axis of the image sensing device 220 can be configured such that the field of view includes an area that is directly adjacent to and visible from the display 218.
[0043] The computing device 200 may also include a sound sensing device 222, such as a microphone or any other currently existing or later developed sound sensing device that can sense the presence of sound near the computing device 200, or communicate with the sound sensing device. The sound sensing device 222 may be positioned such that it is directed at the user operating the computing device 200, and may be configured to receive sound emitted by the user when the user operates the computing device 200, such as speech or other utterances.
[0044] Although Figure 2 The CPU 202 and the memory 204 of the computing device 200 are depicted as being integrated into one unit, other configurations may be utilized. The operation of the CPU 202 may be distributed across multiple machines (where each machine may have one or more processors), and the multiple machines may be directly coupled or coupled across a local area network or other network. The memory 204 may be distributed across multiple machines, such as network-based memory or memory in multiple machines that perform the operations of the computing device 200. Although depicted here as one bus, the bus 212 of the computing device 200 may consist of multiple buses. Further, the auxiliary storage 214 may be directly coupled to other components of the computing device 200 or may be accessed via a network, and may include an integrated unit (such as a memory card) or multiple units (such as multiple memory cards). Thus, the computing device 200 may be implemented in a wide variety of configurations.
[0045] Figure 3 is an illustration of an example of a video stream 300 to be encoded and then decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a number of adjacent frames 304. Although three frames are depicted as the adjacent frames 304, the video sequence 302 may include any number of adjacent frames 304. The adjacent frames 304 may then be further subdivided into individual frames, e.g., frame 306. At the next level, the frame 306 may be divided into a series of planes or segments 308. For example, the segment 308 may be a subset of the frame that permits parallel processing. The segment 308 may also be a subset of the frame that can separate video data into individual colors. For example, a frame 306 of color video data may include one luminance plane and two chrominance planes. The segments 308 may be sampled at different resolutions.
[0046] Regardless of whether frame 306 is divided into segments 308, frame 306 can be further subdivided into blocks 310, which can contain data corresponding to, for example, 16×16 pixels in frame 306. Blocks 310 can also be arranged to include data from one or more segments 308 of the pixel data. Blocks 310 can also be any other suitable size, such as 4×4 pixels, 8×8 pixels, 16×8 pixels, 8×16 pixels, 16×16 pixels, or larger. Unless otherwise stated, the terms block and macroblock are used interchangeably herein.
[0047] Figure 4 is a block diagram of an encoder 400 according to an implementation of the present disclosure. As described above, encoder 400 can be implemented in transmission station 102, such as by providing a computer software program stored in a memory (e.g., memory 204). The computer software program can include machine instructions that, when executed by a processor such as CPU 202, cause transmission station 102 to Figure 4 encode video data in the manner described. Encoder 400 can also be implemented as dedicated hardware included in, for example, transmission station 102. In a particularly desirable implementation, encoder 400 is a hardware encoder.
[0048] Encoder 400 has the following stages for performing various functions in the forward path (shown by solid connecting lines) to use video stream 300 as an input to produce an encoded or compressed bitstream 420: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy coding stage 408. Encoder 400 can also include a reconstruction path (shown by dashed connecting lines) for reconstructing frames for encoding future blocks. In Figure 4 it, encoder 400 has the following stages for performing various functions in the reconstruction path: an inverse quantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filter stage 416. Other structural variations of encoder 400 can be used to encode video stream 300.
[0049] When the video stream 300 is presented for encoding, the corresponding adjacent frames 304, such as frame 306, can be processed in units of blocks. At the intra-frame / inter-frame prediction stage 402, intra-frame prediction (also known as intra-prediction) or inter-frame prediction (also known as inter-prediction) can be used to encode the corresponding blocks. In any case, a prediction block can be formed. In the case of intra-frame prediction, the prediction block can be formed from samples that have been previously encoded and reconstructed in the current frame. In the case of inter-frame prediction, the prediction block can be formed from samples in one or more previously constructed reference frames. The following discusses the implementation of forming the prediction block with respect to Figure 6 , Figure 7 and Figure 8 , for example, using a parameterized motion model identified for encoding the current block of the video frame.
[0050] Next, still referring to Figure 4 , the prediction block can be subtracted from the current block at the intra-frame / inter-frame prediction stage 402 to produce a residual block (also known as the residue). The transform stage 404 uses a block-based transform to transform the residue into transform coefficients, for example, in the frequency domain. The quantization stage 406 uses a quantizer value or quantization level to convert the transform coefficients into discrete quantum values, which are called quantized transform coefficients. For example, the transform coefficients can be divided by the quantizer value and truncated. Then the quantized transform coefficients are entropy encoded by the entropy encoding stage 408. Then the entropy-encoded coefficients, along with other information for decoding the block (this other information can include, for example, the prediction type used, the transform type, the motion vector, and the quantizer value), are output to the compressed bitstream 420. Various techniques (such as variable length code processing (VLC) or arithmetic code processing) can be used to format the compressed bitstream 420. The compressed bitstream 420 can also be referred to as the encoded video stream or the encoded video bitstream, and the terms will be used interchangeably herein.
[0051] Figure 4The reconstruction path (shown by the dashed connection lines) therein can be used to ensure that the encoder 400 and the decoder 500 (described below) use the same reference frame to decode the compressed bitstream 420. The reconstruction path performs functions similar to those that occur during the decoding process (described below), including dequantizing the quantized transform coefficients at the dequantization stage 410, and performing an inverse transform on the dequantized transform coefficients at the inverse transform stage 412 to generate a derivative residual block (also referred to as a derivative residual). At the reconstruction stage 414, the predicted block that has been predicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filter stage 416 can be applied to the reconstructed block to reduce distortions such as blocking artifacts.
[0052] Other variants of the encoder 400 can be used to encode the compressed bitstream 420. For example, for some blocks or frames, a non-transform based encoder can directly quantize the residual signal without the transform stage 404. In another implementation, the encoder can have a quantization stage 406 and a dequantization stage 410 combined in a common stage.
[0053] Figure 5 is a block diagram of a decoder 500 according to an implementation of the present disclosure. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to Figure 5 decode the video data in the manner described. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.
[0054] Similar to the reconstruction path of the encoder 400 discussed above, in one example, the decoder 500 includes the following stages for performing various functions to generate an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filter stage 512, and a post-filter stage 514. Other structural variants of the decoder 500 can be used to decode the compressed bitstream 420.
[0055] When the compressed bitstream 420 is presented for decoding, the entropy decoding stage 502 can decode the data elements within the compressed bitstream 420 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by a quantizer value), and the inverse transform stage 506 performs an inverse transform on the dequantized transform coefficients to produce a derivative residual, which can be the same as the derivative residual created by the inverse transform stage 412 in the encoder 400. Using the header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter prediction stage 508 to create the same prediction blocks as those created in the encoder 400 (e.g., at the intra / inter prediction stage 402). At the reconstruction stage 510, the prediction blocks can be added to the derivative residual to create the reconstructed blocks. The loop filter stage 512 can be applied to the reconstructed blocks to reduce blocking artifacts.
[0056] Other filtering can be applied to the reconstructed blocks. In this example, the post-filtering stage 514 is applied to the reconstructed blocks to reduce blocking distortion or perform other post-processing on the frame, and the result is output as the output video stream 516. The output video stream 516 can also be referred to as the decoded video stream, and these terms will be used interchangeably herein. Other variants of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can produce the output video stream 516 without the post-filtering stage 514.
[0057] Figure 6 is a flowchart of a technique 600 for code processing quantized transform blocks using a wavefront scan order. The technique 600 can include code processing the positions of the block end coefficients in the quantized transform blocks; and code processing the quantized transformed coefficients of the quantized transform blocks using a wavefront scan order.
[0058] The technique 600 can be implemented in a decoder such as Figure 5 the decoder 500 or an encoder such as Figure 4 the encoder 400. When implemented by the decoder, "code processing" (and related terms) means "decoding", such as decoding from a compressed bitstream (e.g., Figure 5 the compressed bitstream 420). When implemented by the encoder, "code processing" (and related terms) means "encoding", such as encoding into a compressed bitstream (e.g., Figure 4 the compressed bitstream 420).
[0059] The technique 600 can be implemented as, for example, something that can be performed by, such as Figure 1A software program executed by a computing device of the transmission station 102 or the receiving station 106. The software program may include machine-readable instructions (e.g., executable instructions) that may be stored in a memory such as the memory 204 or the auxiliary storage 214, and the machine-readable instructions may be executed by a processor such as the CPU 202 to cause the computing device to execute the technique 600. In at least some implementations, the technique 600 may be performed in whole or in part by Figure 4 of the encoder 400 Figure 4 the entropy encoding stage 408 of or Figure 5 the entropy decoding stage 502 of the decoder 500. Thus, the technique 600 can be used by the decoder to decode the quantized transform block from the compressed bitstream to be input (e.g., processed, dequantized, etc.) by the dequantization stage 504. The technique 600 can be used by the encoder to encode the quantized transform block received from the quantization stage 406 into a compressed bitstream.
[0060] The technique 600 can be implemented using dedicated hardware or firmware. Some computing devices may have multiple memories, multiple processors, or both. The steps or operations of the technique 600 can be distributed using different processors, memories, or both. The use of the singular terms "processor" or "memory" encompasses a computing device having one processor or one memory and a device having multiple processors or multiple memories that can be used to execute some or all of the listed steps.
[0061] In an example, the technique 600 can be used to perform code processing on the quantized transform block regardless of the transform type (e.g., one-dimensional vertical, one-dimensional, two-dimensional transform type) used to obtain the transform block from the residual block (and vice versa). In an example, when the transform used to obtain the transform block is a two-dimensional transform type, the technique 600 can be used to perform code processing on the quantized transform block; otherwise, a different scan order is used. Thus, the wavefront scan order can be selected in response to determining that a two-dimensional transform is used to obtain the quantized transform block from the pixel domain residual block.
[0062] At 602, code processing is performed on the position of the EOB coefficient in the transform block. The position of the EOB coefficient can be defined such that the coordinates of any non-zero quantized transform coefficient of the quantized transform block must be less than or equal to the coordinates of the EOB. Thus, if (x 0 , y 0 ) represents the horizontal and vertical coordinates of the EOB (if Cartesian coordinates are used), and (x i , y i ) represents the coordinates of any quantized transform coefficient in the block, and if x i > x 0 or y i > y 0, then this quantized transform coefficient must be zero. In other words, the flipped L-shaped region includes the quantized transform coefficient at the position (p, p) of the quantized transform block and all other quantized transform coefficients having coordinates (p, y) and (x, p) such that y ≤ p and x ≤ p.
[0063] Many different techniques can be used to code the position of the EOB. Refer to Figure 7 which describes several examples. However, other ways of coding the position of the EOB using at least two syntax elements are possible, and the present disclosure is not limited to the examples Figure 7 described.
[0064] Figure 7 Examples 700, 720, and 740 of coding the position of the EOB are shown. Figure 7 The quantized transform blocks 701, 721, and 741 are shown as having a size of 8×8. However, the present disclosure is not limited thereto, and the transform blocks can have any size.
[0065] In example 700, the EOB 702 shows the position of the last non-zero coefficient of the quantized transform block 701. The position of the last non-zero coefficient is defined as above. The quantized transform block 701 is (at least logically) divided into wavefronts with the upper-left pixel 704 as the origin. In the example, if the quantized transform block has a size of N×N, the quantized transform block can include N wavefronts, such as wavefronts 706, 708, and 710. The coding of the EOB described with reference to example 700 is referred to herein as the wavefront design for EOB coding.
[0066] As mentioned, a wavefront is a flipped L-shaped set of coefficients that includes a first set of coefficients along a first axis and a second set of coefficients along a second axis such that the first axis and the second axis intersect at a transform coefficient on the diagonal of the transform block. For illustration, the wavefront 710 includes a first set of coefficients 712 along the vertical axis and a second set of coefficients 714 along the horizontal axis. These axes intersect at coefficient 716, which is on the diagonal of the quantized transform block. The diagonal element (i.e., coefficient 716) is included in one of the first set of coefficients or the second set of coefficients but not both. As Figure 7 shown, the diagonal element is included in the vertical set of coefficients.
[0067] In the example, two syntax elements can be used to code the position of the EOB coefficient. The first syntax element (e.g., a symbol) can indicate the wavefront that includes the end-of-block coefficient. That is, the first syntax element can represent in which wavefront the EOB is. For illustration, the first syntax element can be or represent a value 4 indicating that the EOB 702 is in the 5th wavefront of the quantized transform block.
[0068] The second syntax element may indicate an offset of the EOB relative to a predetermined position in the wavefront. For example, the second syntax element may indicate an offset of the EOB relative to the topmost position of the wavefront. Thus, the second syntax element may be a value of 5 (indicating that the EOB is the 6th quantized transform coefficient in the wavefront). The predetermined position in the wavefront may be a diagonal coefficient (e.g., coefficient 716). Other predetermined positions in the wavefront are possible.
[0069] In another example, three syntax elements may be used to code the position of the EOB coefficient. The first syntax element may indicate the wavefront including the EOB, as described above. The second syntax element may indicate which one of a first coefficient set or a second coefficient set of the wavefront includes the EOB. For illustration, a value of 0 may indicate that the EOB is in the first coefficient set, and a value of 1 may indicate that the EOB is in the second coefficient set. The third syntax element may indicate an offset of the EOB within a subset of one of the first coefficient set or the second coefficient set. Thus, in the example, the second syntax element is coded to indicate whether the end-of-block coefficient is in a column or a row of the wavefront; and the third syntax element is coded to indicate an offset of the end-of-block coefficient within one of the column or the row of the wavefront. The offset may be measured relative to a predetermined position such as a diagonal position (e.g., coefficient 716).
[0070] In example 720, the Cartesian coordinates of the EOB 722 of the quantized transform block 721 may be used to code the position of the EOB 722. Thus, the first syntax element may be used to code the horizontal offset of the EOB, and the second syntax element may be used to code the vertical offset of the EOB in a Cartesian coordinate system having an origin at the direct current (DC) coefficient. That is, the origin may be the upper left corner of the quantized transform block 721, which is the block position (0,0) of the quantized transform block. Thus, with respect to the EOB 722, the first syntax element and the second syntax element may be used to code values 5 and 4, respectively. Coding the EOB as described with respect to example 720 is herein referred to as a diagonal design for EOB coding.
[0071] Example 740 shows coding the position of the end-of-block (EOB) 742 of a quantized transform block 741 using a coordinate system having an origin at the DC coefficient (i.e., coefficient 744) and characterized by anti-diagonals (such as anti-diagonals 746, 748, 750). Each coefficient of the quantized transform block 741 can be located at a Cartesian position (column, row). For example, the EOB 742 is at the Cartesian position (1, 5). The anti-diagonals of example 740 can be lines such that the quantized transform coefficients of the quantized transform block 741 having the same value column + row are on the same anti-diagonal. For example, anti-diagonal 746 includes those coefficients having column + row = 1. Thus, the quantized transform coefficients at positions (0, 1) and (1, 0) are on anti-diagonal 746. As another example, anti-diagonal 748 includes those quantized transform coefficients having column + row = 4. Thus, anti-diagonal 746 includes the quantized transform coefficients at the Cartesian positions (4, 0), (3, 1), (2, 2), (1, 3), and (0, 4).
[0072] Coding the position of the EOB can include coding a first syntax element that indicates the anti-diagonal (i.e., the index for it) that includes the EOB. In an example, a second syntax element can indicate an offset from a predetermined position (e.g., the lower left position) on the line relative to the anti-diagonal. Thus, with respect to the EOB 742, the first and second syntax elements can indicate values 6 and 1, respectively. In another example, the second syntax element can be used to code the distance of the EOB to the center of the anti-diagonal, and a third syntax element can be used to indicate whether the EOB is in the lower left region or the upper right region of the diagonal.
[0073] Entropy coding can be used for the syntax elements used to code the position of the EOB using cumulative distribution functions (CDFs) representing different probability models conditioned on different factors. The factor can include one or more of the following: the size of the quantized transform block, the color channel that the quantized transform block corresponds to, and the transform type used to obtain the transform block from which the quantized transform block is obtained. The color channel can be, for example, the luminance (Y) channel or one of the chrominance channels (U or V). The transform type can be, for example, a two-dimensional transform, a one-dimensional vertical transform, or a one-dimensional horizontal transform. Other transform types are possible. In an example, each transform block can be associated with a fixed combination of these factors. For illustration, for a quantized transform block having size 8×8, Y channel, 2D transform, both the encoder and decoder can use the corresponding CDF to write / read (i.e., encode / decode) the symbol indicating the value of the EOB.
[0074] To further refine the probability model selection, additional factors based on the technique used to code the position of the EOB can be used.
[0075] For example, in the wavefront design for EOB code processing, and as mentioned, the first syntax element may indicate in which flipped L-shaped region the EOB is; the second syntax element may be a boolean value indicating whether the EOB is on the horizontal axis or the vertical axis; and the third syntax element may represent an offset relative to the diagonal coefficient. In an example, the third syntax element may have a separate probability model for each of the horizontal and vertical axes. That is, if the second syntax element has a first value, one probability model may be selected for coding the third syntax element; and if the second syntax element has a second value, another probability model may be selected for coding the third syntax element. Thus, the context model of the third syntax element includes the second syntax element.
[0076] For example, in the diagonal design for EOB code processing, and as described above, the first symbol may indicate in which anti-diagonal the EOB is. Different ways may be used to indicate the offset of the EOB on the line. In an example, the offset may be defined as the distance to the lower left corner of the current row. In another example, the second syntax element may code the distance of the EOB to the center of the anti-diagonal, and the third syntax element may be used to indicate whether the EOB is in the lower left region or the upper right region. Thus, different probability models may be associated with these different design alternatives.
[0077] At Figure 6 604 of , the quantized transformed coefficients of the quantized transform blocks are coded using a wavefront scan order. The wavefront scan order is characterized by coding the corresponding quantized transform coefficients of the flipped L-shaped regions. That is, each flipped L-shaped region includes a subset of the quantized transformed coefficients of the quantized transform block; and the quantized transform coefficients are coded one flipped L-shaped region at a time. First, the transform coefficients along the first axis of the flipped L-shaped region are coded, and subsequently, the second quantized transform coefficients along the second axis of the flipped L-shaped region are coded. First, the quantized transform coefficients along the first axis of the flipped L-shaped region are coded, and subsequently, the second quantized transform coefficients along the second axis of the flipped L-shaped region are coded.
[0078] Entropy coding (encoding into a compressed bitstream and decoding from the compressed bitstream) is performed on each quantized transform coefficient. Different techniques can be used to code the quantized transform coefficients. In an example, the levels (e.g., values) of the quantized transform coefficients can be decomposed into different planes. The lower-level planes can correspond to coefficient levels between 0 and 2, while the higher-level planes can be used to code levels above 2. Separating into planes can be used to assign rich context models to at least the lower-level planes. In an example, the context can include one or more of the size of the quantized transform block and adjacent coefficient information. The higher-level planes can use a simplified context model for levels between 3 and 15, and can directly code the residuals above level 15 using ExpGolomb codes.
[0079] Now refer to Figure 8 to explain coding of the quantized transform block using a wavefront scan order. Figure 8 An example of coding the quantized transform coefficients using a wavefront scan order is shown. Figure 8 It includes quantized transform blocks 810, 830, and 850. Each of the quantized transform blocks 810, 830, 850 is shown as including a symbol 802 indicating the position of the EOB in this block, and is shown as including a symbol 804 indicating the position of the diagonal position (to be further described below). Thus, in the quantized transform block 810, the EOB is at coordinates (5,5) and the diagonal position is also at (5,5); in the quantized transform block 830, the EOB is at coordinates (5,3) and the diagonal position is at (5,5); and in the quantized transform block 850, the EOB is at coordinates (2,5) and the diagonal position is at (5,5). The coding of the quantized transform block using a wavefront scan order can be summarized as follows.
[0080] Let (x,y) represent the offset of the quantized transform coefficient with the upper left corner as the origin. The coordinates of the EOB are given by (x 0 ,y 0 ).
[0081] First, locate the diagonal position P. The coordinates of the diagonal position are given by (x 1 ,y 1 ) and satisfy the conditions: x 1 = y 1 , x 1 ≥ x 0 , y 1 ≥ y 0 . And for those that also satisfy x 2 = y 2 , x 2 >= x 0 , y 2 >= y0 of any other pixel (x 2 , y 2 ), it is required that the diagonal position P is closest to the EOB: x 1 < x 2 . Thus, the diagonal position is the diagonal position of the quantized transform block on the same horizontal axis or the same vertical axis as the EOB, taking either one with the larger value. That is, if x 0 ≥ y 0 , then the diagonal position P is at (x 0 , x 0 ); and if x 0 < y 0 , then the diagonal position P is at (y 0 , y 0 ).
[0082] The diagonal position P can be used to identify the flipped L-shaped region. Starting from the diagonal position P, the transform coefficients are sequentially coded in the first direction (e.g., the vertical direction) until the block boundary (e.g., the top boundary of the quantized transform block) is reached, and then sequentially coded in the second direction (e.g., the horizontal direction) until the block boundary (e.g., the left boundary of the quantized transform block) is reached. The quantized transform coefficients whose coordinates (x i , y i ) satisfy x i > x 0 or y i > y 0 are skipped (i.e., not coded). These coefficients are known to be zero.
[0083] After coding all the quantized transform coefficients of the current flipped L-shaped region, the coding proceeds to the next smaller flipped L-shaped region until there are no more flipped L-shaped regions available. If the current flipped L-shaped region is defined by the diagonal position (p, p), then the next smaller flipped L-shaped region is defined by the diagonal position (p - 1, p - 1).
[0084] The numbers shown at the positions of the quantized transform coefficients in each of the quantized transform blocks 810, 830, 850 indicate the coding order of the coefficients of this block. X and Y indicate the direction or the coding order. The shaded squares indicate that the coefficients are coded in the vertical (shaded with pattern 806) direction and the horizontal (shaded with pattern 808) direction.
[0085] Table I shows the pseudocode for code processing of a quantized transform block using a wavefront scan order. Other ways of implementing code processing of a quantized transform block using a wavefront scan order are possible, and the disclosure herein is not limited by the pseudocode shown in Table I. Table I shows a case where end-of-block (EOB) is code processed using two syntax elements (eob_symbol_1 and eob_symbol_2). However, as mentioned above, more than two syntax elements can be used. The function f(·) at line 3 takes a symbol and determines the coordinates of the EOB. The actual implementation of the function f(·) depends on the way the symbols for code processing the EOB are defined, as described with respect to Figure 7 as described.
[0086]
[0087] In an example, all the quantized transform coefficients in a block can share the same cumulative distribution function (CDF) (e.g., the same probability model). In another example, separate CDFs can be used for each axis in the scanned axes. For example, two separate CDFs can be used, one for the coefficients scanned vertically and the other for the coefficients scanned horizontally.
[0088] As mentioned above, code processing of a quantized transform block includes traversing the quantized transform block using a scan order and determining a context that can be used to select a probability distribution (e.g., an estimate, a model) for code processing a particular quantized transform block. The context model can include adjacent quantized transform coefficients that have already been code processed.
[0089] Traditionally, to encode a quantized transform block, the quantized transform coefficients can first be serialized in a particular scan order (from a 2D block). The scan order can be chosen such that in the serialized 1-D vector, the quantized transform coefficients will exhibit some correlation with their adjacent quantized transform coefficients, or the quantized transform coefficients will have certain characteristics in the order of the scan.
[0090] Then, the quantized transform coefficients are entropy coded according to the scan order (e.g., in the order given by the scan order). The scan order can scan the quantized transform block in a reverse manner (i.e., from higher frequencies to lower frequencies, ending at the DC coefficient). An appropriate probability distribution is assumed (e.g., used or selected) for each quantized transform coefficient for the entropy encoder. The probability distribution can also be updated adaptively as more quantized transform coefficients are code processed to adapt to the statistics of different video sequences.
[0091] For a more accurate estimate of the probability distribution, the current coefficient is first classified into different distributions using a context model. For example, when adjacent coefficients are large, it is very likely that the current coefficient may also have a larger variance.
[0092] Figure 9 An example 900 of a backward (i.e., reverse) zigzag scan order and context selection is shown. The quantized transform block 902 will be coded using the backward zigzag scan order 920. Although the quantized transform block 902 is shown as having a size of 8×8, the present disclosure is not limited thereto. The quantized transform block 902 may have other N×N sizes. The numbers in the backward zigzag scan order 920 show the order of accessing (e.g., traversing) the quantized transform block 902. Although the backward zigzag scan order 920 shows that the coding process starts at the coefficient located at (N-1,N-1), this may not be the case. The starting point for traversing the quantized transform block 902 depends on the position of the EOB.
[0093] The quantized transform block 902 includes a current coefficient 904 located at Cartesian coordinates (x,y) and corresponding to the scan order position 55. Line 906 shows the order of traversing (e.g., coding process) of the coefficients in the region of the current coefficient 904. The already coded adjacent quantized transform coefficients (also called context coefficients) for context selection include the coefficients at Cartesian positions (x+2,y), (x+1,y+1), (x,y+2), (x,y+1), and (x+1,y) corresponding to the scan order positions 44, 45, 46, 51, and 52, respectively. Thus, in order to identify (e.g., select, pick, etc.) the appropriate probability distribution for coding the current coefficient 904, the values of the indicated adjacent coefficients are used to provide an index to the appropriate probability distribution used by the encoder and decoder. In the example, the sum of the context coefficients is used to obtain the index of the appropriate probability distribution.
[0094] Thus, with context modeling, the coding process of the current quantized transform coefficient becomes dependent on the values of its previously coded neighbors. In other words, the quantized value of the current coefficient will also affect the compression of future coefficients in this transform block. "Future coefficients" refer to the quantized transform coefficients that follow the current quantized transform coefficient in the scan order.
[0095] To better account for such correlations, the encoder can implement (e.g., apply) one or more optimization schemes to jointly determine the corresponding quantized values (e.g., quantization levels) of the coefficients. One such and known optimization scheme is the lattice optimization that can be applied to a transform block to determine the optimal quantization levels of the transform coefficients of the transform block. Lattice optimization is now briefly described.
[0096] Different states can be maintained for each transform coefficient (e.g., different states can be associated with each transform coefficient), where each state indicates a quantization level. For example, a first state associated with a transform coefficient (e.g., state 0) can indicate that regular quantization will be applied to the transform coefficient; and a second state (e.g., state 1) can indicate that a reduced quantization level (e.g., subtracting one from the quantization step size) will be applied to the transform coefficient. Other states and / or other state semantics can be used. After associating states with coefficients, at each coefficient, the grid optimization considers the states of the coefficients from the previous code processing and selects, for the current coefficient, the best origin state for each state maintained for the current coefficient. The grid optimization can proceed backward from the EOB to the start of the scan order. When the optimization is complete, the best path through all the coefficients can be traced back from the starting point (e.g., the DC coefficient or the first scan order position) to the EOB (e.g., the last scan order position) to find the best path. Then the transform coefficients are quantized according to the states of the best path.
[0097] Figure 10 Example 1000 shows a grid optimization for determining the quantization levels of the transform coefficients of a transform block. Example 1000 shows the transform coefficients at the scan positions of the scan order, such as coefficients 1002, 1004, 1006 at respective scan positions 1, N - 3, and N - 1, where N corresponds to the number of scan positions. Example 1000 shows the use of two states: a regular quantization state 1008 and a reduced quantization state 1010.
[0098] For each transform coefficient, and for each of the possible states, the grid optimization can use two origins (states) from the immediately preceding transform coefficient to quantize the transform coefficient. By way of illustration, with respect to coefficient 1006, coefficient 1006 is quantized based on the quantization of coefficient 1004 using the regular quantization state 1008 (as shown by line 1012) and also based on the quantization of coefficient 1004 using the reduced quantization state 1010 (as shown by line 1014). The grid optimization can then retain the optimal state (as shown by the solid line such as line 1014). The grid optimization can then discard the other (e.g., non - optimal) states (e.g., the state corresponding to the dashed line such as line 1012). A better state can be a state that results in a better compression ratio. Which is the better state can be determined based on rate - distortion cost analysis. Finally, the best path (shown by the thick line 1016) can be retained because it corresponds to the optimized quantization result of the transform coefficient.
[0099] Grid optimization achieves better results when there are first-order correlations between transform coefficients. First-order correlation means that the coding (e.g., quantization) of the current transform coefficient depends only on the immediately preceding coded coefficient in the coding order (which may be a reverse scan order). When the correlation is not first-order, grid optimization may not be optimally optimized. Nevertheless, grid optimization may still be effective when the correlation is "local." Local correlation means that the context coefficient is not immediately preceding in the scan order, but is in the Cartesian neighborhood of the current coefficient. However, given a correlation such as about Figure 9 In the described case of scan order and context model, the context model neighbors, although very local in the 2-D sense, are actually far away from the current coefficient in the scan order, making the grid optimization invalid or at least less optimal than in the case of first-order correlation.
[0100] The foregoing shows that, in order to jointly optimize the quantized coefficients, the design of both the context model and the scan order should take into account the optimization method used (such as grid optimization). As mentioned above, the wavefront scan order described in this article can solve problems such as these because it obtains the first-order correlation of at least some (if not most) of the transform coefficients of the transform block.
[0101] In order to obtain better optimization results, it is desirable that the scanning order used places at least a portion of the contextual dependencies in a local area of the scan. The wavefront scanning order achieves such results by performing a scan in a first direction (e.g., vertical direction) and then performing a scan in a second direction (e.g., horizontal direction) in the backward L-shaped area as described above. Thus, the wavefront scanning order can be said to take into account the context model by placing at least a portion of the contextual dependencies in a local area of the scan.
[0102] To reiterate, traversing the quantized transform block using wavefront scan order includes coding the quantized transform coefficients in each backward L-shaped region starting from the outermost (e.g., largest) backward L-shaped region toward the inner (e.g., smaller) backward L-shaped regions. The outermost backward L-shaped region is identified based on the position of the EOB, as described above with respect to Figure 8 For each backward L-shaped region, the transform coefficients quantized at the intersection of the diagonal lines are coded, then the transform coefficients quantized along the first direction (e.g., on the vertical line) are coded, and finally the transform coefficients quantized along the second direction (e.g., on the horizontal line) are coded.
[0103] consider Figure 9The context model (i.e., the context of the coefficient at (x,y) includes the coefficients at positions (x+2,y), (x+1,y+1), (x,y+2), (x,y+1), and (x+1,y)), and the wavefront scan order described herein, at least some of these context coefficients are immediate neighbors in the scan order. For illustration, consider Figure 8 the scan position number 22 of the quantized transform block 810 of Figure 8 , whose context coefficients include the quantized transform coefficients at scan positions 21, 20, 14, 13, and 4. Thus, using the wavefront scan order, the coefficients at scan positions 21 and 20 are immediate neighbors in the scan of the quantized transform coefficient at scan position 22, which in turn makes the lattice optimization more effective.
[0104] Thus, even when using the same context model as the context model described with respect to Figure 9 the context model for many of the quantized transform coefficients in the quantized transform coefficients will include immediate neighbors in the scan order, which provides benefits over other scan orders in lattice code processing. This is true even for some quantized transform coefficients (e.g., the quantized transform coefficients at the diagonal intersections), where none of the context coefficients are immediate neighbors.
[0105] Figure 11 is a flowchart of a technique 1100 for code processing a quantized transform block using a wavefront scan order. Technique 1100 may include selecting a wavefront scan order for code processing the quantized transformed coefficients of the quantized transform block; selecting a probability distribution for code processing the quantized transform coefficients among the quantized transform coefficients; and entropy coding the quantized transformed coefficients using the probability distribution.
[0106] Technique 1100 may be implemented in a decoder such as Figure 5 decoder 500 or in an encoder such as Figure 4 encoder 400. When implemented by a decoder, "code processing" (and related terms) means "decoding", such as decoding from a compressed bitstream (e.g., Figure 5 the compressed bitstream 420 of Figure 5 ). When implemented by an encoder, "code processing" (and related terms) means "encoding", such as encoding into a compressed bitstream (e.g., Figure 4 the compressed bitstream 420 of Figure 4 ).
[0107] Technique 1100 may be implemented as, for example, may be performed by a device such as Figure 1A software program executed by a computing device of the transmission station 102 or the receiving station 106. The software program may include machine-readable instructions (e.g., executable instructions) that may be stored in a memory such as the memory 204 or the auxiliary storage 214, and the machine-readable instructions may be executed by a processor such as the CPU 202 to cause the computing device to execute the technique 1100. In at least some implementations, the technique 1100 may be performed in whole or in part by Figure 4 of the encoder 400 Figure 4 entropy encoding stage 408 of Figure 5 entropy decoding stage 502 of the decoder 500. Thus, the technique 1100 can be used by the decoder to decode the quantized transform blocks from the compressed bitstream to be input (e.g., processed, dequantized, etc.) to the dequantization stage 504. The technique 1100 can be used by the encoder to encode the quantized transform blocks received from the quantization stage 406 into a compressed bitstream.
[0108] The technique 1100 can be implemented using dedicated hardware or firmware. Some computing devices may have multiple memories, multiple processors, or both. The steps or operations of the technique 1100 can be distributed using different processors, memories, or both. The use of the singular terms "processor" or "memory" encompasses computing devices having one processor or one memory and devices having multiple processors or multiple memories that can be used to execute some or all of the listed steps.
[0109] At 1102, a wavefront scan order is selected for code processing the quantized transform coefficients of the quantized transform block. The quantized transform block may have a size N×N, where N is a positive integer. As described above, the wavefront scan order is such that for at least one x and at least one y, where 2≤a≤N - 1 and 2≤b≤N - 1, the positions (x + 1, y), (x + 1, y - 1), and (x + 1, x - 2) are sequentially code processed, and the positions (x, y + 1), (x - 1, y + 1), and (x - 2, y + 1) are sequentially code processed. As mentioned above, the wavefront scan order is characterized by code processing the corresponding quantized transform coefficients of the flipped L-shaped region, and code processing the first quantized transform coefficient along the first axis of the flipped L-shaped region, and then code processing the second quantized transform coefficient along the second axis of the flipped L-shaped region. Also as described above, the flipped L-shaped region includes the quantized transform coefficient at the position (p, p) of the quantized transform block and all other quantized transform coefficients having coordinates (p, y) and (x, p) such that y≤p and x≤p.
[0110] At 1104, a probability distribution for coding the quantized transform coefficients among the quantized transform coefficients is selected. As described above, the context model for selecting the probability model includes at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order. At 1106, the quantized transformed coefficients are coded using the probability distribution.
[0111] In some examples, the context may be selected in a manner that better supports quantization optimization algorithms such as lattice optimization. For example, since the immediate neighbors in the scan order are the immediate neighbors from which lattice coding can more benefit, when generating (computing) the context, instead of using the sum of 2D neighbors as described above, a weighted sum may alternatively be used. The weights of the immediate neighbors may be greater than the context coefficients that are not immediate neighbors in the wavefront scan order. Thus, technique 1100 may include obtaining a context that is a weighted combination of context coefficients that are quantized transform coefficients. The context coefficients include a first context coefficient and a second context coefficient, where the first context coefficient is an immediate neighbor of the quantized transform coefficient in the wavefront scan order and the second context coefficient is not an immediate neighbor. A first weight used with the first context coefficient may be greater than a second weight used with the second context coefficient.
[0112] The flipped L-shaped region may be considered to be divided into sub-regions by the wavefront scan order; each sub-region includes quantized transform coefficients that are coded in a particular direction, as described above. For illustration, in Figure 8 the quantized transform block 830 of, the first sub-region of the flipped L-shaped region 832 includes the coefficients at scan positions 8, 9, 10, and 11; and the second sub-region of the flipped L-shaped region 832 includes the coefficients at scan positions 12, 13, and 14. Thus, as described herein, if the two coefficients are in the same sub-region of the flipped L-shaped region, they are said to be "coded together".
[0113] In an example, the weight for a context coefficient may depend on the position of the context coefficient in the quantized transform block and whether the context coefficient is coded together with the current coefficient. When coding the quantized transform coefficients along the vertical axis of the flipped L-shaped region, a context coefficient that is in the same column as the quantized transform coefficient and is coded together with the quantized transform coefficient in the flipped L-shaped region may be assigned a higher weight than other context coefficients that are not in the same column.
[0114] For illustration, and referring again to Figure 8When coding the quantized transform coefficient at scan position 22 of the quantized transform block 810, a larger weight can be used with the context coefficients at scan positions 21 and 20 compared to the weights used with the context coefficients at scan positions 14, 13, and 4. Similarly, when coding the quantized transform coefficients along the horizontal axis of the flipped L-shaped region, the context coefficients that are in the same row as the quantized transform coefficient and are coded together with the quantized transform coefficient can be assigned a higher weight than other context coefficients that are not in the same row. Thus, a first weight is used with a first context coefficient that is along the same dimension as the quantized transform coefficient in the flipped L-shaped region and is coded together with the quantized transform coefficient, and a second weight lower than the first weight is used with a second context coefficient that is not in the flipped L-shaped region.
[0115] In an example, the same weight (e.g., 1) can be used for the context coefficients of the quantized transform coefficients on the diagonal of the transform block (i.e., the diagonal elements of the flipped L-shaped region). Thus, when the quantized transform coefficient is on the diagonal of the quantized transform block, the context can be the sum of the context coefficients of the quantized transform coefficient.
[0116] In an example, the position of the context neighbors can be changed for the scan order locality of the current quantized transform coefficient being coded (e.g., set the locality, adapt to the locality, select based on the locality, etc.). That is, which relative Cartesian neighbor is used as the context coefficient for the quantized transform coefficient depends on the position of the quantized transform coefficient in the sub-region of the flipped L-shaped region to which the quantized transform coefficient belongs. In an example, the number of immediate neighbors of the quantized transform coefficient that are used as context coefficients can depend on the position of the quantized transform coefficient. In other words, the first set of context coefficients for a first quantized transform coefficient can include different relative neighbors than the second set of context coefficients for a second quantized transform coefficient.
[0117] For illustration, for a location at Figure 8For the coefficient at scan position 13 of the quantized transform block 810, the context coefficients can include the coefficients at scan positions 12, 11, 3, and 2; and for the coefficient at scan position 14, the context coefficients can be those at scan positions 13, 12, 11, 4, and 3. That is, when available, more immediate neighbors can be used as context coefficients. The context coefficients are available for the current coefficient if they are co-processed with the context coefficients. As described above, if two quantized transform coefficients belong to the same sub-region of the flipped L-shaped region, they are co-processed. Thus, the context coefficient is available at least when the context coefficient is along the same axis (e.g., column or row) as the current quantized transform coefficient and is also in the same flipped L-shaped region. Similar adjustments can be made to the quantized transform coefficients along the horizontal axis.
[0118] For ease of explanation, techniques 600 and 1100 are depicted and described as a corresponding series of steps or operations. However, the steps or operations according to the present disclosure can occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein can be used. Moreover, not all of the illustrated steps or operations may be required to implement the method according to the disclosed subject matter.
[0119] The encoding and decoding aspects described above illustrate some examples of encoding and decoding techniques. However, it should be understood that when those terms are used in the claims, encoding and decoding can mean compressing data, decompressing data, transforming data, or any other processing or alteration of data.
[0120] The word "example" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as an "example" is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, the use of the word "example" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clearly indicated in the context, the statement "X includes A or B" is intended to mean any of its natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then "X includes A or B" is satisfied under any of the above instances. Additionally, unless otherwise specified or clearly indicated in the context for the singular form, the articles "a / an" used in this application and the appended claims should generally be construed to mean "one or more". Moreover, the use of the term "implementation" or the term "an implementation" throughout this disclosure is not intended to mean the same embodiment or implementation unless so described.
[0121] The implementation of the transmission station 102 and / or the receiving station 106 (as well as the algorithms, methods, instructions, etc. stored thereon and / or executed thereby (including by the encoder 400 and the decoder 500)) can be implemented in hardware, software, or any combination thereof. The hardware can include, for example, a computer, an intellectual property (IP) core, an application specific integrated circuit (ASIC), a programmable logic array, an optical processor, a programmable logic controller, microcode, a microcontroller, a server, a microprocessor, a digital signal processor, or any other suitable circuitry. In the claims, the term "processor" should be understood to cover any one of the foregoing hardware, either individually or in combination. The terms "signal" and "data" can be used interchangeably. Further, parts of the transmission station 102 and the receiving station 106 do not necessarily have to be implemented in the same way.
[0122] Further, in one aspect, for example, the transmission station 102 or the receiving station 106 can be implemented using a general-purpose computer or a general-purpose processor having a computer program that, when executed, performs any one of the corresponding methods, algorithms, and / or instructions described herein. Additionally or alternatively, for example, a special-purpose computer / processor can be utilized, which can include additional hardware for performing any one of the methods, algorithms, or instructions described herein.
[0123] The transmission station 102 and the receiving station 106 can be implemented, for example, on a computer in a video conferencing system. Alternatively, the transmission station 102 can be implemented on a server, and the receiving station 106 can be implemented on a device separate from the server (such as a handheld communication device). In this instance, the transmission station 102 can use the encoder 400 to encode the content into an encoded video signal and transmit the encoded video signal to the communication device. Subsequently, the communication device can then use the decoder 500 to decode the encoded video signal. Alternatively, the communication device can decode content locally stored on the communication device (e.g., content not transmitted by the transmission station 102). Other suitable transmission and reception implementations are available. For example, the receiving station 106 can be a generally stationary personal computer instead of a portable communication device, and / or the device including the encoder 400 can also include the decoder 500.
[0124] Further, all or part of the implementation of the present disclosure can take the form of a computer program product that can be accessed from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can tangibly contain, store, communicate, or transport a program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media are also available.
[0125] The above-described embodiments, implementations, and aspects have been described to facilitate an easy understanding of the present disclosure and not to limit the present disclosure. On the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope should be accorded the broadest interpretation permitted under the law so as to cover all such modifications and equivalent structures.
Claims
1. A method for coding a quantized transform block, comprising: selecting a wavefront scan order for coding the transformed coefficients of the quantization of the quantized transform block, wherein the quantized transform block has a size of N×N, wherein the wavefront scan order is such that for at least one x and at least one y, where 2≤x<N and 2≤y<N, the positions (x, y - 1), (x, y - 2) and (x, y - 3) are coded sequentially, and the positions (x - 1, y - 1), (x - 2, y) and (x - 3, y) are coded sequentially; selecting a probability distribution for coding the quantized transform coefficients among the quantized transform coefficients, wherein the context model for selecting the probability distribution includes at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order; and entropy coding the quantized transformed coefficients using the probability distribution.
2. The method according to claim 1, wherein, the wavefront scan order is characterized by coding the corresponding quantized transform coefficients of a flipped L-shaped region, and wherein the first quantized transform coefficient along the first axis of the flipped L-shaped region is coded, and then the second quantized transform coefficient along the second axis of the flipped L-shaped region is coded.
3. The method according to claim 2, further comprising: obtaining a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein a first weight is used together with a first context coefficient along the same dimension as the quantized transform coefficient in the flipped L-shaped region, and a second weight lower than the first weight is used together with a second context coefficient not in the flipped L-shaped region.
4. The method according to claim 2, wherein the flipped L-shaped region includes the quantized transform coefficient at the position (p, p) of the quantized transform block and all other quantized transform coefficients having coordinates (p, y) and (x, p) such that y≤p and x≤p.
5. The method according to any one of claims 1 to 4, further comprising: obtaining a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein the context coefficients include a first context coefficient and a second context coefficient, the first context coefficient being an immediate neighbor of the quantized transform coefficient in the wavefront scan order, the second context coefficient not being an immediate neighbor, and wherein a first weight used together with the first context coefficient is greater than a second weight used together with the second context coefficient.
6. The method according to any one of claims 1 to 4, wherein the quantized transform coefficient is located on the diagonal of the quantized transform block, the method further comprising: obtaining a context that is the sum of the context coefficients of the quantized transform coefficients.
7. The method according to any one of claims 1 to 4, wherein the number of immediate neighbors of the quantized transform coefficient used as a context coefficient depends on the position of the quantized transform coefficient.
8. An apparatus for coding a quantized transform block, comprising: A processor configured to: Select a wavefront scan order for coding the transformed coefficients of the quantization of the quantized transform block, wherein the quantized transform block has a size of N×N, wherein the wavefront scan order is such that for at least one x and at least one y, where 2≤x<N and 2≤y<N, the positions (x, y-1), (x, y-2) and (x, y-3) are coded sequentially, and the positions (x-1, y), (x-2, y) and (x-3, y) are coded sequentially; Select a probability distribution for coding the quantized transform coefficients among the quantized transform coefficients, wherein the context model for selecting the probability distribution includes at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order; and Entropy code the quantized transformed coefficients using the probability distribution.
9. The apparatus according to claim 8, wherein, the wavefront scan order is characterized by coding the corresponding quantized transform coefficients of the flipped L-shaped region, and wherein the first quantized transform coefficient along the first axis of the flipped L-shaped region is coded, and then the second quantized transform coefficient along the second axis of the flipped L-shaped region is coded.
10. The apparatus according to claim 9, wherein the processor is further configured to: Obtain a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein a first weight is used together with a first context coefficient along the same dimension as the quantized transform coefficient in the flipped L-shaped region, and a second weight lower than the first weight is used together with a second context coefficient not in the flipped L-shaped region.
11. The apparatus according to claim 9, wherein the flipped L-shaped region includes the quantized transform coefficient at the position (p, p) of the quantized transform block and all other quantized transform coefficients having coordinates (p, y) and (x, p) such that y≤p and x≤p.
12. The apparatus according to any one of claims 8 to 11, wherein the processor is further configured to: Obtain a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein the context coefficients include a first context coefficient and a second context coefficient, the first context coefficient is an immediate neighbor of the quantized transform coefficient in the wavefront scan order, the second context coefficient is not an immediate neighbor, and wherein a first weight used together with the first context coefficient is greater than a second weight used together with the second context coefficient.
13. The apparatus according to any one of claims 8 to 11, wherein the quantized transform coefficients are located on the diagonal of the quantized transform block, and the processor is further configured to: Obtain a context that is the sum of the context coefficients of the quantized transform coefficients.
14. The apparatus according to any one of claims 8 to 11, wherein the number of immediate neighbors of the quantized transform coefficients used as context coefficients depends on the position of the quantized transform coefficients.
15. A non - transitory computer - readable storage medium includes executable instructions that, when executed by a processor, facilitate the execution of operations for coding - processing quantized transform blocks, the operations including: selecting a wavefront scan order for coding - processing the quantized transform coefficients of the quantized transform block, wherein the quantized transform block has a size N×N, wherein the wavefront scan order is such that for at least one x and at least one y, where 2≤x<N and 2≤y<N, the positions (x,y - 1), (x,y - 2), and (x,y - 3) are coded - processed in sequence, and the positions (x - 1,y), (x - 2,y), and (x - 3,y) are coded - processed in sequence; selecting a probability distribution for coding - processing the quantized transform coefficients among the quantized transform coefficients, wherein the context model for selecting the probability distribution includes at least two immediate neighbors of the quantized transform coefficients in the wavefront scan order; and entropy - coding the quantized transform coefficients using the probability distribution.
16. The non - transitory computer - readable storage medium according to claim 15, wherein, the wavefront scan order is characterized by coding - processing the corresponding quantized transform coefficients of a flipped L - shaped region, and wherein the first quantized transform coefficient along the first axis of the flipped L - shaped region is coded - processed, and then the second quantized transform coefficient along the second axis of the flipped L - shaped region is coded - processed.
17. The non - transitory computer - readable storage medium according to claim 16, the operations further including: obtaining a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein a first weight is used with a first context coefficient along the same dimension as the quantized transform coefficient in the flipped L - shaped region, and a second weight lower than the first weight is used with a second context coefficient not in the flipped L - shaped region.
18. The non - transitory computer - readable storage medium according to claim 16, wherein the flipped L - shaped region includes the quantized transform coefficient at the position (p,p) of the quantized transform block and all other quantized transform coefficients having coordinates (p,y) and (x,p) such that y≤p and x≤p.
19. The non - transitory computer - readable storage medium according to any one of claims 15 to 18, the operations further including: obtaining a context that is a weighted combination of context coefficients of the quantized transform coefficients, wherein the context coefficients include a first context coefficient and a second context coefficient, the first context coefficient is an immediate neighbor of the quantized transform coefficient in the wavefront scan order, the second context coefficient is not an immediate neighbor, and wherein a first weight used with the first context coefficient is greater than a second weight used with the second context coefficient.
20. The non - transitory computer - readable storage medium according to any one of claims 15 to 18, wherein the quantized transform coefficients are located on the diagonal of the quantized transform block, the operations further Comprising: Obtaining a context that is a sum of context coefficients of the quantized transform coefficients as the context coefficients.
21. The non-transitory computer-readable storage medium according to any one of claims 15 to 18, wherein the number of immediate neighbors of the quantized transform coefficients used as context coefficients depends on the position of the quantized transform coefficients.