Dual Standard Block Partitioning Heuristics for Lossy Compression

Through the dual standard block segmentation heuristic, combining expected entropy and visual masking, the partitioning strategy of image blocks is determined, which solves the problem of ringing artifacts in lossy compression, improves image quality and maintains high compression efficiency.

CN115280772BActive Publication Date: 2025-07-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080098574.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-08
Publication Date
2025-07-22
Estimated Expiration
2040-04-08

AI Technical Summary

Technical Problem

The existing lossy compression technology is prone to introduce ringing artifacts in image processing, especially at transitions between fast-changing areas and slow-changing areas of the image, resulting in a degradation of image quality.

Method used

The double standard block segmentation heuristic is used to determine whether the image block is partitioned to reduce the propagation of quantized artifacts by estimating the expected entropy and visual masking of the block. This method partitions the block into sub-blocks, calculates the visual masking amount of each sub-block, selects the highest visual masking value as the visual masking feature of the block, and determines whether to divide the blocks with the expected entropy.

Benefits of technology

It effectively reduces the appearance of ringing artifacts in the image, improves image quality, and maintains high compression efficiency. It is suitable for various image contents including graphics, photos and screenshots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280772B_ABST
    Figure CN115280772B_ABST
Patent Text Reader

Abstract

A method for partitioning an image block to reduce quantization artifacts includes estimating the expected entropy of the block; partitioning the block into sub-blocks, where the size of each sub-block is the smallest possible partitioning size; calculating the corresponding visual masking amounts of the sub-blocks; selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; combining the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value; and determining whether to segment the block based on the segmentation indicator.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Image content (such as a still image or a video frame) represents a large amount of online content. For example, a web page may include multiple images, and most of the time and resources spent on rendering the web page are dedicated to rendering these images for display. The amount of time and resources required to receive and render an image for display depends in part on how the image is compressed. Thus, it is possible to render an image more quickly by reducing the total data size of the image using lossy compression and decompression techniques.

[0002] Lossy compression techniques attempt to represent image content using fewer bits than the number of bits in the original image. Lossy compression techniques may introduce visual artifacts such as ringing artifacts and banding artifacts into the decompressed image. Higher compression levels may result in more pronounced artifacts. It is desirable to minimize artifacts while maintaining a high compression level. Summary of the Invention

[0003] One aspect of the present disclosure is a method, and one aspect is a method for partitioning an image block to reduce quantization artifacts. The method includes: estimating the expected entropy of the block; partitioning the block into sub-blocks, where each sub-block has the smallest possible partition size; calculating the corresponding visual masking amounts of the sub-blocks; selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; combining the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value; and determining whether to segment the block based on the segmentation indicator.

[0004] Another aspect is an apparatus for partitioning an image block to reduce quantization artifacts. The apparatus includes a memory and a processor. The processor is configured to execute instructions stored in the memory to: estimate the expected entropy of the block; partition the block into sub-blocks, where each sub-block has the smallest possible partition size; calculate the corresponding visual masking amounts of the sub-blocks; select the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; combine the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value; and determine whether to segment the block based on the segmentation indicator.

[0005] Another aspect is a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including executable instructions that, when executed by a processor, facilitate the execution of operations, the operations including: estimating the expected entropy of the block; partitioning the block into sub-blocks, where each sub-block has the smallest possible partition size; calculating the corresponding visual masking amounts of the sub-blocks; selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; combining the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value; and determining whether to segment the block based on the segmentation indicator.

[0006] These and other aspects of the disclosure are disclosed in the following detailed description of the embodiments, the appended claims, and the drawings.

[0007] It will be appreciated that the aspects can be implemented in any convenient form. For example, the aspects can be implemented by a suitable computer program, which can be carried on a suitable carrier medium. The suitable carrier medium can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). The aspects can also be implemented using a suitable device, which can take the form of a programmable computer running a computer program, the computer program being arranged to implement the methods and / or techniques disclosed herein. The aspects can be combined such that features described in the context of one aspect can be implemented in another aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a diagram of a computing device according to an embodiment of the present disclosure.

[0009] Figure 2 is a diagram of a computing and communication system according to an embodiment of the present disclosure.

[0010] Figure 3 is a diagram of a video stream for use in encoding and decoding according to an embodiment of the present disclosure.

[0011] Figure 4 is a block diagram of an encoder according to an embodiment of the present disclosure.

[0012] Figure 5 is a block diagram of a decoder according to an embodiment of the present disclosure.

[0013] Figure 6 is a block diagram of a representation of a portion of a frame according to an embodiment of the present disclosure.

[0014] Figure 7 is a flowchart of a technique for partitioning an image block to reduce quantization artifacts according to an embodiment of the present disclosure.

[0015] Figure 8 is a flowchart of a technique 800 for partitioning an image block to reduce quantization artifacts according to an embodiment of the present disclosure.

[0016] Figure 9 is an example of calculating a Laplacian operator according to an embodiment of the present disclosure.

[0017] Figure 10 is an example of ringing artifacts according to an embodiment of the present disclosure.

[0018] Figure 11An example of a hypothetical comparison of a transform partition according to the conventional art and a transform partition according to an embodiment of the present disclosure is shown. Detailed Description

[0019] A video compression scheme may include decomposing each image (a single image or a video frame) into smaller parts such as blocks, and using techniques that limit the information included in the output for each block to generate an output bitstream. The encoded bitstream can be decoded to recreate the blocks and the source image based on the limited information. In some embodiments, the information included in the output for each block can be limited by reducing spatial redundancy, reducing temporal redundancy (in the case of video), or a combination thereof. For example, by predicting a frame based on information available to both the encoder and the decoder, and including information representing the difference or residual between the predicted frame and the original frame, temporal (in the case of video frames) or spatial redundancy can be reduced.

[0020] The residual information can be further compressed by transforming the residual information into transform coefficients. Transforming the residual information into transform coefficients can include a quantization step that introduces loss - hence the name or designation "lossy compression".

[0021] Lossy compression can be used to code the visual information of an image. Lossy compression techniques can be applied to the source image to produce a compressed image. The inverse of the lossy technique can be applied to the compressed image to produce a decompressed image. The lossy aspect of lossy compression techniques can be at least partially attributed to the quantization of frequency domain information (described further below). The amount of loss is specified by a quantization step using a quantization parameter (QP).

[0022] The quantization parameter can be used to control the trade - off between rate and distortion. Generally, a larger quantization parameter means higher quantization (such as of transform coefficients), resulting in a lower rate but higher distortion; and a smaller quantization parameter means lower quantization, resulting in a higher rate but lower distortion. The variables QP, q, and Q can be used interchangeably in the present disclosure to refer to the quantization parameter.

[0023] Generally, an image block (e.g., a luminance block or a chrominance block) can be divided into smaller blocks. One type of partitioning of the block (i.e., prediction partitioning) can be used for prediction purposes. Another type of partitioning (i.e., transform partitioning) can be used for the purpose of transforming the predicted residual into the transform domain. A transform type such as DCT is used to transform the predicted residual into the transform domain. The prediction partitioning and the transform partitioning do not necessarily result in the same partitioning of the source block. For illustration, the image block can be a 32×32 luminance block. In raster scan order, the prediction partitioning can be sub-blocks of sizes 8×16, 8×16, 16×16, 16×16, 16×8, and 16×8. On the other hand, in raster scan order, the transform partitioning can be 8×8, 8×8, 8×8, 8×8, 16×16, 16×16, 8×8, 8×8, 8×8, and 8×8. The sub-blocks of the prediction partitioning and / or the transform partitioning can be square or rectangular.

[0024] Traditionally, determining the transform partitioning for transforming residual information can include recursively determining whether the cost of using the current block size for transformation exceeds the cost of partitioning the current block into sub-blocks and using the sub-block size for transformation coding. If the cost of coding using the sub-block size is smaller, then the sub-block size is used as the current block size and the determination for each sub-block is repeated with an even smaller sub-block size.

[0025] Given a specific desired quality, the cost is typically based on a rate-distortion (RD) function defined to balance rate and distortion. The rate refers to the number of bits required for encoding (such as encoding a block, a frame, quantized transform block coefficients, etc.). The distortion measures the quality loss between the source image block and the reconstructed version of the source image block. By performing a rate-distortion optimization (RDO) process, the codec optimizes the amount of distortion with respect to the rate required for encoding the video (i.e., the specified quality).

[0026] Generally, quantization parameters are used to calculate the RD cost. More generally, whenever an encoder decision (e.g., mode decision) is based on the RD cost, the encoder can use the QP value to determine the RD cost.

[0027] The QP can be used to derive a multiplier for combining the rate and distortion values into one metric. Some codecs may refer to the multiplier as a Lagrangian multiplier (denoted as λ mode ); other codecs may use a similar multiplier called rdmult. Each codec can have a different method for calculating the multiplier. Unless the context clearly indicates otherwise, regardless of the codec, the multiplier will be referred to as the Lagrangian multiplier or Lagrangian parameter in this document.

[0028] For example, given a mode m corresponding to a certain transform partitioning, let r m represent the rate or cost (in bits) generated by using mode m and let dm represents the resulting distortion. The rate-distortion cost for selecting mode m can be calculated as a scalar value: d m + λ mode r m . By using the Lagrange parameter λ mode , the costs of two modes can then be compared and the mode with the lower combined RD cost can be selected. This technique for evaluating the rate-distortion cost is the basis of the mode decision process in at least some codecs.

[0029] As mentioned above, lossy compression aims to describe (i.e., code, compress, etc.) an image with the fewest number of bits while retaining as much of the image quality as possible when decompressing the compressed image. That is, lossy compression techniques attempt to compress an image without degrading the image quality to an unacceptable level.

[0030] During compression, fewer bits can be used to describe the slow-changing regions and / or objects of an image compared to the bits that can be used to describe the fast-changing regions and / or objects. "Slow-changing" and "fast-changing" refer to changes in the frequency domain in this context. Figure 10 is Example 1000 of ringing artifacts according to an embodiment of the present disclosure. In Example 1000, an image such as Image 1010 includes a sky that forms the background of Image 1010 and tree branches that occlude at least a portion of the background. Regions covered by the sky background, such as Region 1002, are slow-changing regions. That is, the slow-changing regions correspond to spatial low-frequency regions. Regions of Image 1010 that include the transition between the background sky and the tree branches, such as Region 1004, are fast-changing regions. Magnifying glass 1050 illustrates a magnified region of Image 1010 that includes Region 1004. The fast-changing regions correspond to spatial high-frequency regions. Region 1006 (i.e., a portion of the tree trunk) illustrates another example of a slow-changing region. Thus, Image 1010 illustrates that the background sky mainly includes slow-changing regions but also includes fast-changing regions.

[0031] As further described below, encoding an image or a picture (i.e., a video frame) of a video can include partitioning the image into blocks. As used herein, both "image" and "picture" refer to a single image or video frame. A block can include slow-changing regions (i.e., low-frequency signals) and fast-changing regions (i.e., high-frequency signals).

[0032] Lossy compression techniques may produce undesirable artifacts such as ringing artifacts, banding artifacts. Ringing artifacts appear at sharp transitions in an image (e.g., at the Figure 10 edge between the sky and the tree branches in Region 1004), thereby transforming a smooth image gradient into a perceptible discrete band, contour, color separation, or some other artifact.

[0033] For example, ringing artifacts can be generated by compressing high-frequency signals. Ringing artifacts may appear as bands and / or ghosts near the edges of objects in the decompressed image. Figure 10 Region 1052 in FIG. illustrates an example of ringing. Ringing artifacts are caused by undershoots and overshoots around the edges. "Undershoot" means that the value of a pixel in the decompressed image is less than the value of the same pixel in the source image. That is, "undershoot" can mean that the pixels around the edge are faded. "Overshoot" means that the value of a pixel in the decompressed image is greater than the value of the same pixel in the source image. That is, "overshoot" can mean that some of the pixels around the edge are intensified. That is, due to lossy compression, in the decompressed image, some parts of the bright (dark) background can become brighter (darker).

[0034] Overshoots and undershoots can be caused by sinusoidal oscillations in the frequency domain. For example, in an image including a bright (dark) background partially occluded by a dark (bright) foreground object, there is a step function at the edge of the background and the foreground object.

[0035] If the edge is compressed according to a frequency-based transform, due to the quantization frequency limitation characteristic, an increased quantization level will result in sinusoidal oscillations near the edge. As mentioned, undershoots and overshoots can be observed around the edge. As further described below, examples of frequency-based transforms (also called "block-based transforms") include discrete cosine transform (DCT), Fourier transform (FT), discrete sine transform (DST), etc.

[0036] It should be noted that larger transform blocks can allow larger regions of the source image to share common parameters and can be efficient for uniform textures and low bitrates. However, larger transform blocks come at the cost of artifacts that propagate further in the image. That is, the artifacts can statistically propagate all the way into the blocks in the image covered by the transform block.

[0037] For illustration, assume that an image includes a first 8×8 block containing a lot of texture (e.g., a patch of grass) and a smooth second adjacent 8×8 image block (e.g., a part of a car hood). If the first block and the second block are encoded separately (e.g., transformed), the reconstructed (e.g., decoded) version of the first block may include ringing, while the reconstructed version of the second block will not include ringing. On the other hand, if the two blocks are transformed together using an 8×16 (or 16×8, depending on how the first block and the second block are arranged) transform size, the reconstructed second block will also likely include ringing artifacts.

[0038] As described above, traditionally, given a target quality (e.g., a given QP value), the RD cost can be used to determine the transform partition. The RD cost is calculated based on two objective metrics (e.g., the cost in bits and the distortion).

[0039] However, using only the rate and the distortion can result in partitioning a block when the partition is not needed or not partitioning the block when the partition might be preferred.

[0040] This disclosure relates to transform partitioning. More specifically, this disclosure relates to determining an optimal partition of a block into sub-blocks such that each sub-block is then transformed using a transform type such as the DCT. As used herein, the optimal partition is a partition that balances reducing the number of bits used to compress the block while simultaneously improving the quality by restricting the propagation of quantization artifacts (e.g., ringing). The block partition (e.g., the transform block size) can be selected based on the sub-blocks of the block according to the amount of ringing suitable for the block.

[0041] Regarding restricting the propagation of artifacts, embodiments of this disclosure determine whether artifacts from one block are likely to be visible or masked in an adjacent block. If the artifacts in the first block are likely to be masked in the second block, the first block and the second block can be combined into one block for the purpose of transformation into the frequency domain. If the artifacts of the first block are not likely to be masked in the second block, the first block is not combined with the second block. Instead, the first block is transformed separately.

[0042] To determine whether a block should be partitioned, embodiments of a dual-criterion block splitting heuristic for lossy compression calculate a heuristic that combines two criteria to obtain a split indicator value. Whether the block should be further split depends on whether the split indicator value is above or below a threshold. Any number of values can be used to compare the split indicator value. In an example, the split indicator value can be compared to an empirically derived constant. In an example, the split indicator value can be compared to a value related (e.g., linearly related) to the target image quality and / or the allowable error. In an example, if the block is to be partitioned, the split indicator value can be compared to a value derived based on the resulting partition size.

[0043] The first criterion estimate (e.g., approximation, correlation, etc.) has an expected entropy required for a specified quality. The specific quality can be given by a specific QP value. The second criterion estimates how easily artifacts (e.g., ringing artifacts) will be visible in the block. The second criterion is an estimate of the impact of the artifacts in the sub-blocks on adjacent sub-blocks. That is, the second criterion estimates the amount of visual masking in the sub-blocks of the block. The size of each sub-block can be the smallest possible partition size typically supported by a codec that determines the segmentation. In an example, the block can be a luminance block of size 32×32, and the smallest possible partition size can be 8×8. However, other block sizes and smallest possible partition sizes are possible. In an example, the minimum size of each sub-block can be changed based on the target quality. For example, at the highest quality (e.g., butteraugli limit <1.0, or some other quality criterion), the visual masking sub-block size can be set to 5×5 or some other such small size; and at low quality (e.g., butteraugli limit >4 or 8, or some other quality criterion), the minimum size can be extended to a larger size (e.g., 12×12, 16×16, etc.). The butteraugli limit is explained below.

[0044] Visual masking is a property of the human eye. Visual masking allows some artifacts to disappear (e.g., become unobtrusive) and some artifacts to appear (e.g., become obtrusive). Thus, if the artifacts (e.g., ringing artifacts) in a sub-block are likely to disappear in other sub-blocks of the block (i.e., more precisely, the reconstructed version of the sub-block), this favors combining the sub-block with other adjacent sub-blocks for transformation using a transform of the combined sub-block size (such as DCT). On the other hand, if the artifacts are unlikely to disappear in adjacent sub-blocks, this favors transforming the sub-block without combining it with other adjacent sub-blocks.

[0045] Thus, for example, when determining whether to partition an N×N block (e.g., a 32×32 block), in addition to considering the entropy required to achieve a specified quality in the block (e.g., an indication of the rate), embodiments according to the present disclosure also consider the masking characteristics within the block.

[0046] For illustration, given an N×N (e.g., 32×32) block, the smallest possible sub-blocks of the block (e.g., 8×8) can be considered. The visual masking characteristics (or simply masking characteristics) of each sub-block can be considered (e.g., computed, determined, etc.). If any sub-block is flat (e.g., includes little texture), then that sub-block may be more likely to show artifacts propagating from another sub-block than in the case where the sub-block is noisy (e.g., includes a large amount of texture). If each sub-block of the N×N block has some noise (such as in the case of an image block of a carpet, a bush, or grass), then any artifacts in one sub-block can be hidden in the texture of other sub-blocks (e.g., masked by visual masking). On the other hand, if at least a portion of the remainder of the N×N block is smooth (e.g., in the case of the sky or car paint), then there is no visual masking and the visibility of the artifacts is revealed (e.g., propagated to at least that portion of the remainder of the N×N block).

[0047] Thus, the dual-criterion block splitting (i.e., transform partitioning) heuristic refers to using both a cost in bits (first criterion) and visual masking within sub-blocks (second criterion) to determine block partitioning. Determine the worst-case region of visual masking within the block (i.e., one or more sub-blocks of the block). Assume that artifacts propagate in an equivalent manner throughout the block. Thus, it can be assumed that the artifacts within the block are determined by the lowest visual masking region within the block.

[0048] Benefits of the dual-criterion block splitting heuristic for lossy compression include that a single heuristic (the dual-criterion block splitting heuristic) can be used for a wide range of images (such as graphics, photos, screenshots, user interface elements), while preserving the natural appearance of fragile textures (such as forests, marble, skin), while still using a smaller integral transform, where the smaller transform can be useful (such as in the case of screenshots, graphics, or the boundary between a smooth region and a non-smooth region).

[0049] Details of the dual-criterion block splitting heuristic for lossy compression are described herein with reference first to a system in which the teachings herein can be implemented. Figure 1 is a schematic diagram of a video encoding and decoding system 100. The transmitting station 102 can be, for example, a computer having an internal hardware configuration such as the computer described with respect to Figure 2 However, other suitable implementations of the transmitting station 102 are possible. For example, the processing of the transmitting station 102 can be distributed among multiple devices.

[0050] Network 104 can connect the transmitting station 102 and the receiving station 106 for encoding and decoding video streams. Specifically, the video stream can be encoded in the transmitting station 102, and the encoded video stream can be decoded in the receiving station 106. For example, network 104 can be the Internet. Network 104 can also be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network, or any other means of transmitting the video stream from the transmitting station 102 to (in this example) the receiving station 106.

[0051] In one example, the receiving station 106 can be a computer with an internal hardware configuration, such as the computer described with respect to Figure 2 However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.

[0052] Other implementations of the video encoding and decoding system 100 are possible. For example, an implementation can omit network 104. In another implementation, the video stream can be encoded and then stored for later transmission to the receiving station 106 or any other device with a memory. In one implementation, the receiving station 106 (e.g., via network 104, a computer bus, and / or some communication path) receives the encoded video stream and stores the video stream for later decoding. In an example implementation, the Real-Time Transport Protocol (RTP) is used to transmit the encoded video over network 104. In another implementation, a transport protocol other than RTP can be used (e.g., an HTTP-based video stream transport protocol).

[0053] For example, when used in a video conferencing system, the transmitting station 102 and / or the receiving station 106 can include the ability to encode and decode video streams as described below. For example, the receiving station 106 can be a video conferencing participant that receives an encoded video bitstream from a video conferencing server (e.g., the transmitting station 102) to decode and view and further encode its own video bitstream and transmit its own video bitstream to the video conferencing server for other participants to decode and view.

[0054] Figure 2 is a block diagram of an example of a computing device 200 that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement Figure 1 one or both of the transmitting station 102 and the receiving station 106. The computing device 200 can be in the form of a computing system including multiple computing devices, or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.

[0055] The CPU 202 in the computing device 200 can be a central processing unit. Alternatively, the CPU 202 can be any other type of device capable of manipulating or processing information, or multiple devices, whether existing or to be developed in the future. Although the disclosed embodiments can be practiced using a single processor (e.g., CPU 202) as shown, advantages in terms of speed and efficiency can be achieved by using more than one processor.

[0056] In an embodiment, the memory 204 in the computing device 200 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 204. The memory 204 can include code and data 206 that are accessed by the CPU 202 using the bus 212. The memory 204 can further include an operating system 208 and application programs 210, where the application programs 210 include at least one program that allows the CPU 202 to execute the methods described herein. For example, the application programs 210 can include Applications 1 through N, which further include video coding applications that execute the methods described herein. The computing device 200 can also include an auxiliary storage device 214, which can be, for example, a memory card used with the mobile computing device 200. Because video communication sessions can contain a large amount of information, they can be stored in whole or in part in the auxiliary storage device 214 and loaded into the memory 204 as needed for processing.

[0057] The computing device 200 can also include one or more output devices, such as a display 218. In one example, the display 218 can be a touch-sensitive display that combines a display with touch-sensitive elements that can be operated to sense touch inputs. The display 218 can be coupled to the CPU 202 via the bus 212. In addition to or in place of the display 218, other output devices can be provided that allow a user to program or otherwise use the computing device 200. When the output device is or includes a display, the display can be implemented in various ways, including as a liquid crystal display (LCD); a cathode ray tube (CRT) display; or a light-emitting diode (LED) display, such as an organic LED (OLED) display.

[0058] The computing device 200 can also include or communicate with an image sensing device 220, such as a camera or any other image sensing device, whether existing or to be developed in the future, that can sense an image, such as an image of a user operating the computing device 200. The image sensing device 220 can be positioned such that it is pointed at the user operating the computing device 200. In an example, the position and optical axis of the image sensing device 220 can be configured such that the field of view includes an area that is directly adjacent to and from which the display 218 is visible.

[0059] The computing device 200 can also include or communicate with a sound sensing device 222, such as a microphone or any other existing or later developed sound sensing device capable of sensing sound near the computing device 200. The sound sensing device 222 can be positioned such that it is pointed at the user operating the computing device 200 and can be configured to receive sound emitted by the user when the user operates the computing device 200, such as speech or other utterances.

[0060] Although Figure 2 The CPU 202 and the memory 204 of the computing device 200 are described as integrated into a single unit, but other configurations can be utilized. The operation of the CPU 202 can be distributed across multiple machines (each machine having one or more processors) that can be directly coupled or coupled across a local area network or other network. The memory 204 can be distributed across multiple machines, such as network-based memory or memory in multiple machines that execute the operations of the computing device 200. Although described herein as a single bus, the bus 212 of the computing device 200 can consist of multiple buses. Additionally, the auxiliary storage device 214 can be directly coupled to other components of the computing device 200 or can be accessed via a network and can include a single integrated unit such as a memory card or multiple units such as multiple memory cards. The computing device 200 can thus be implemented in a variety of configurations.

[0061] Figure 3 FIG. is an example of a video stream 300 to be encoded and then decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes a plurality of adjacent frames 304. Although three frames are depicted as adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. The adjacent frames 304 can then be further subdivided into individual frames, such as frame 306. At the next level, the frame 306 can be divided into a series of segments 308 or planes. For example, the segment 308 can be a subset of the frame that allows for parallel processing. The segment 308 can also be a subset of the frame that can separate video data into different colors. For example, a frame 306 of color video data can include one luminance plane and two chrominance planes. The segments 308 can be sampled at different resolutions.

[0062] Regardless of whether the frame 306 is divided into segments 308, the frame 306 can be further subdivided into blocks 310, which can contain data corresponding to, for example, 16×16 pixels in the frame 306. The blocks 310 can also be arranged to include data from one or more segments 308 of the pixel data. The blocks 310 can also have any other suitable size, such as 4×4 pixels, 8×8 pixels, 16×8 pixels, 8×16 pixels, 16×16 pixels, or larger.

[0063] Figure 4 is a block diagram of an encoder 400 according to an embodiment of the present disclosure. As described above, the encoder 400 can be implemented in the transmitting station 102, such as by providing a computer software program stored in a memory (e.g., memory 204). The computer software program can include machine instructions that, when executed by a processor such as CPU 202, cause the transmitting station 102 to encode video data in the manner described herein. The encoder 400 can also be implemented as dedicated hardware, for example, included in the transmitting station 102. The encoder 400 has the following stages to perform various functions in the forward path (shown by solid connecting lines) to generate an encoded or compressed bitstream 420 using the video stream 300 as an input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy coding stage 408. The encoder 400 may also include a reconstruction path (shown by dashed connecting lines) to reconstruct frames for encoding future blocks. In Figure 4 it, the encoder 400 has the following stages to perform various functions in the reconstruction path: an inverse quantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filter stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.

[0064] When presenting the video stream 300 for encoding, the frame 306 can be processed in units of blocks. In the intra / inter prediction stage 402, a block can be encoded using intra prediction (also known as intra-frame prediction) or inter prediction (also known as inter-frame prediction) or a combination of both. In any case, a prediction block can be formed. In the case of intra prediction, all or part of the prediction block can be formed by samples that have been previously encoded and reconstructed in the current frame. In the case of inter prediction, all or part of the prediction block can be formed by samples in one or more previously constructed reference frames determined using motion vectors.

[0065] Next, still referring to Figure 4 , the prediction block can be subtracted from the current block in the intra / inter prediction stage 402 to generate a residual block (also known as a residue). The transform stage 404 transforms the residue into transform coefficients, for example, in the frequency domain, using a block-based transform. Such block-based transforms (i.e., transform types) include, for example, the discrete cosine transform (DCT) and the asymmetric discrete sine transform (ADST). Other block-based transforms are possible. In addition, a combination of different transforms can be applied to a single residue. In one example of applying the transform, the DCT transforms the residue block into a frequency domain where the transform coefficient values are based on spatial frequency. The lowest frequency (DC) coefficient is at the upper left corner of the matrix, and the highest frequency coefficients are at the lower right corner of the matrix. It is worth noting that the size of the prediction block and thus the resulting residual block may be different from the size of the transform block. For example, the prediction block can be divided into smaller blocks to which individual transforms are applied.

[0066] Quantization stage 406 converts the transform coefficients into discrete quantum values using a quantizer value or quantization level, and the discrete quantum values are referred to as quantized transform coefficients. For example, the transform coefficients can be divided by the quantizer value and the transform coefficients can be truncated. Then, the entropy coding stage 408 performs entropy coding on the quantized transform coefficients. Any number of techniques including tokens and binary trees can be used to perform the entropy coding. Then, the entropy-coded coefficients are output to the compressed bitstream 420 together with other information for decoding the block, which can include, for example, the type of prediction used, the transform type, the motion vector, and the quantizer value. The information for decoding the block can be entropy-coded into the block, frame, slice, and / or segment headers within the compressed bitstream 420. The compressed bitstream 420 can also be referred to as an encoded video stream or an encoded video bitstream; these terms will be used interchangeably herein.

[0067] Figure 4 The reconstruction path (shown by the dashed connection lines) in can be used to ensure that both the encoder 400 and the decoder 500 (described below) use the same reference frames and blocks to decode the compressed bitstream 420. The reconstruction path performs functions similar to those that occur during the decoding process and are discussed in more detail below, including dequantizing the quantized transform coefficients at the dequantization stage 410 and performing an inverse transform on the dequantized transform coefficients at the inverse transform stage 412 to produce a derivative residual block (also referred to as a derivative residual). At the reconstruction stage 414, the predicted block predicted at the intra / inter prediction stage 402 can be added to the derivative residual to create a reconstructed block. The loop filter stage 416 can be applied to the reconstructed block to reduce distortions such as block artifacts.

[0068] Other variants of the encoder 400 can be used to encode the compressed bitstream 420. For example, a non-transform-based encoder 400 can directly quantize the residual signal without the transform stage 404 for some blocks or frames. In another embodiment, the encoder 400 can combine the quantization stage 406 and the dequantization stage 410 into a single stage.

[0069] Figure 5 is a block diagram of a decoder 500 according to an embodiment of the present disclosure. The decoder 500 can be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program can include machine instructions that, when executed by a processor such as the CPU 202, cause the receiving station 106 to decode video data in the manner described below. The decoder 500 can also be implemented in hardware included in, for example, the transmitting station 102 or the receiving station 106.

[0070] Similar to the reconstruction path of the encoder 400 discussed above, in one example, the decoder 500 includes the following stages to perform various functions to generate an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512, and a post-filtering stage 514. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 420.

[0071] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by a quantizer value), and the inverse transform stage 506 inverse-transforms the dequantized transform coefficients using a selected transform type to produce a derivative residual, which can be the same as the derivative residual created by the inverse transform stage 412 in the encoder 400. Using the header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter prediction stage 508 to create the same prediction blocks as those created in the encoder 400, e.g., at the intra / inter prediction stage 402. At the reconstruction stage 510, the prediction blocks can be added to the derivative residual to create a reconstructed block. The loop filtering stage 512 can be applied to the reconstructed block to reduce block artifacts. Other filtering can be applied to the reconstructed block. In an example, the post-filtering stage 514 is applied to the reconstructed block to reduce block distortion, and the output result is used as the output video stream 516. The output video stream 516 can also be referred to as a decoded video stream; these terms will be used interchangeably herein.

[0072] Other variations of the decoder 500 can be used to decode the compressed bitstream 420. For example, the decoder 500 can generate the output video stream 516 without the post-filtering stage 514. In some embodiments of the decoder 500, the post-filtering stage 514 is applied after the loop filtering stage 512. The loop filtering stage 512 can include an optional deblocking filtering stage. Additionally or alternatively, the encoder 400 includes an optional deblocking filtering stage in the loop filtering stage 416.

[0073] The codec can use multiple transform types. For example, the transform type can be Figure 4 The transform type used by the transform stage 404 for generating the transform block. For example, the transform type (i.e., the inverse transform type) can be Figure 5The transform type used for the dequantization stage 504. Available transform types can include a one-dimensional discrete cosine transform (1D DCT) or an approximation thereof, a one-dimensional discrete sine transform (1D DST) or an approximation thereof, a two-dimensional DCT (2D DCT) or an approximation thereof, a two-dimensional DST (2D DST) or an approximation thereof, and the identity transform. Other transform types can be available. In an example, a one-dimensional transform (1D DCT or 1D DST) can be applied in one dimension (e.g., a row or a column), and the identity transform can be applied in the other dimension.

[0074] In the case of using a 1D transform (e.g., 1D DCT, 1D DST) (e.g., 1D DCT is applied to a column (or row) of a transform block), the quantized coefficients can be coded by using a row-by-row (i.e., raster) scan order or a column-by-column scan order. In the case of using a 2D transform (e.g., 2D DCT), different scan orders can be used to code the quantized coefficients. As indicated above, different templates can be used to derive contexts for coding non-zero flags of a non-zero map based on the transform type used. Thus, in an embodiment, a template can be selected based on the transform type used to generate a transform block. As indicated above, examples of transform types include: 1D DCT applied to a row (or column) and the identity transform applied to a column (or row); 1D DST applied to a row (or column) and the identity transform applied to a column (or row); 1D DCT applied to a row (or column) and 1D DST applied to a column (or row); 2D DCT; and 2D DST. Other combinations of transforms can include transform types.

[0075] Figure 6 is a block diagram of a representation of a portion 600 of a frame (such as, Figure 3 frame 306) according to an embodiment of the present disclosure. As shown, the portion 600 of the frame includes four 64×64 blocks 610 in two rows and two columns in a matrix or Cartesian plane, which can be referred to as superblocks. The superblocks can have a larger or smaller size. Although Figure 6 explained with reference to superblocks of size 64×64, the description can be easily extended to larger (e.g., 128×128) or smaller superblock sizes.

[0076] In an example and without loss of generality, a superblock can be a basic or maximum coding unit (CU). Each superblock can include four 32×32 blocks 620. Each 32×32 block 620 can include four 16×16 blocks 630. Each 16×16 block 630 can include four 8×8 blocks 640. Each 8×8 block 640 can include four 4×4 blocks 650. Each 4×4 block 650 can include 16 pixels, which can be represented as four rows and four columns in each corresponding block in a Cartesian plane or matrix. The pixels can include information representing an image captured in a frame, such as luminance information, color information, and location information. In an example, a block such as a 16×16 pixel block as shown can include a luminance block 660, which can include luminance pixels 662; and two chrominance blocks 670 / 680, such as a U or Cb chrominance block 670 and a V or Cr chrominance block 680. The chrominance blocks 670 / 680 can include chrominance pixels 690. For example, the luminance block 660 can include 16×16 luminance pixels 662, and each chrominance block 670 / 680 can include 8×8 chrominance pixels 690, as shown. Although one block arrangement is shown, any arrangement can be used. Although Figure 6 an N×N block is shown, in some embodiments, an N×M block can be used, where N≠M. For example, a 32×64 block, a 64×32 block, a 16×32 block, a 32×16 block, or any other sized block can be used. In some embodiments, an N×2N block, a 2N×N block, or a combination thereof can be used.

[0077] In some embodiments, coding (i.e., image or video coding) can include ordered block-level coding. Ordered block-level coding can include coding the blocks of a frame in an order such as a raster scan order, where the blocks can be identified and processed starting from the block at the upper left corner of the frame or a portion of the frame, and proceeding along rows from left to right and from the top row to the bottom row, identifying each block in turn for processing. For example, the superblocks in the top row and left column of the frame can be the first blocks to be coded, and the superblock immediately to the right of the first block can be the second block to be coded. The second row from the top can be the second row to be coded, such that the superblock in the left column of the second row can be coded after the superblock in the rightmost column of the first row.

[0078] In an example, coding the blocks can include using quadtree coding, and quadtree coding can include coding smaller block units with respect to the blocks in a raster scan order. In Figure 6The lower left corner of the portion of the frame shown, e.g., a 64×64 superblock, can be coded using quadtree coding, where the upper left 32×32 block can be coded, then the upper right 32×32 block can be coded, then the lower left 32×32 block can be coded, and then the lower right 32×32 block can be coded. Each 32×32 block can be coded using quadtree coding, where the upper left 16×16 block can be coded, then the upper right 16×16 block can be coded, then the lower left 16×16 block can be coded, and then the lower right 16×16 block can be coded. Each 16×16 block can be coded using quadtree coding, where the upper left 8×8 block can be coded, then the upper right 8×8 block can be coded, then the lower left 8×8 block can be coded, and then the lower right 8×8 block can be coded. Each 8×8 block can be coded using quadtree coding, where the upper left 4×4 block can be coded, then the upper right 4×4 block can be coded, then the lower left 4×4 block can be coded, and then the lower right 4×4 block can be coded. In some embodiments, for 16×16 blocks, the 8×8 blocks can be omitted, and the 16×16 blocks can be coded using quadtree coding, where the upper left 4×4 block can be coded, and then the other 4×4 blocks in the 16×16 block can be coded in raster scan order.

[0079] In an example, video coding can include compressing the information included in an original (e.g., input, source, etc.) frame by omitting some information in the original frame from the corresponding coded frame. For example, coding can include reducing spectral redundancy, reducing spatial redundancy, reducing temporal redundancy, or a combination thereof.

[0080] In an example, reducing spectral redundancy can include using a color model based on a luminance component (Y) and two chrominance components (U and V or Cb and Cr), which can be referred to as a YUV or YCbCr color model or color space. Using the YUV color model can include using a relatively large amount of information to represent the luminance component of a portion of the frame, and using a relatively small amount of information to represent each corresponding chrominance component of the portion of the frame. For example, a portion of the frame can be represented by a high-resolution luminance component, which can include 16×16 pixel blocks, and can be represented by two lower-resolution chrominance components, each of which represents the portion of the frame as 8×8 pixel blocks. A pixel can indicate a value (e.g., a value in the range from 0 to 255) and can be stored or transmitted using, e.g., eight bits. Although the present disclosure is described with reference to the YUV color model, any color model can be used.

[0081] Reducing spatial redundancy can include transforming a block into the frequency domain as described above. For example, units of an encoder (such as Figure 4 the entropy coding stage 408) can perform a DCT using transform coefficient values based on spatial frequency.

[0082] Reducing temporal redundancy can include using the similarity between frames to encode a frame using a relatively small amount of data based on one or more reference frames, where the one or more reference frames can be previously encoded, decoded, and reconstructed frames of a video stream.

[0083] As described above, quadtree coding can be used to code a superblock. Figure 7 is a block diagram of an example 700 of a quadtree representation of a block according to an embodiment of the present disclosure. Example 700 includes a block 702. As described above, block 702 can be referred to as a superblock or a CTB. Example 700 illustrates the partitioning of block 702. However, block 702 can be partitioned differently, such as by an encoder (e.g., Figure 4 the encoder 400).

[0084] Example 700 illustrates that block 702 is partitioned into four blocks, namely blocks 702-1, 702-2, 702-3, and 702-4. Block 702-2 is further partitioned into blocks 702-5, 702-6, 702-7, and 702-8. Thus, if, for example, the size of block 702 is N×N (e.g., 128×128), then blocks 702-1, 702-2, 702-3, and 702-4 each have a size of N / 2×N / 2 (e.g., 64×64), and blocks 702-5, 702-6, 702-7, and 702-8 each have a size of N / 4×N / 4 (e.g., 32×32). If a block is partitioned, the block is partitioned into four non-overlapping square sub-blocks of equal size.

[0085] The quadtree data representation is used to describe how block 702 is partitioned into sub-blocks, such as blocks 702-1, 702-2, 702-3, 702-4, 702-5, 702-6, 702-7, and 702-8. A quadtree 703 showing the partitioning of block 702. If each node of the quadtree 703 is further split into four child nodes, the node is assigned a flag "1", and if the node is not split, the node is assigned a flag "0". The flag can be referred to as a split bit (e.g., 1) or a stop bit (e.g., 0) and is coded in the compressed bitstream. In a quadtree, a node has four child nodes or no child nodes. A node with no child nodes corresponds to a block that is not further partitioned. Each child node that splits a block corresponds to a sub-block.

[0086] In the quadtree 703, each node corresponds to a sub-block of the block 702. The corresponding sub-block is shown in parentheses. For example, the node 704-1 with the value 0 corresponds to the block 702-1.

[0087] The root node 704-0 corresponds to the block 702. Since the block 702 is divided into four sub-blocks, the value of the root node 704-0 is a splitting bit (e.g., 1). At the intermediate level, the flag indicates whether the sub-block of the block 702 is further divided into four sub-sub-blocks. In this case, the node 704-2 includes the flag "1" because the block 702-2 is divided into the blocks 702-5, 702-6, 702-7, and 702-8. Each of the nodes 704-1, 704-3, and 704-4 includes the flag "0" because the corresponding block is not divided. Since the nodes 704-5, 704-6, 704-7, and 704-8 are at the bottom layer of the quadtree, these nodes do not require the flags "0" or "1". It can be inferred from the absence of additional flags corresponding to these blocks that the blocks 702-5, 702-6, 702-7, and 702-8 are not further divided.

[0088] The quadtree data representation of the quadtree 703 can be represented by the binary data "10100", where each bit represents a node 704 of the quadtree 703. The binary data indicates the partitioning of the block 702 to the encoder and the decoder. In the case where the encoder needs to transmit the binary data to the decoder (such as Figure 5 the decoder 500), the encoder can encode the binary data in a compressed bitstream (such as Figure 4 the compressed bitstream 420).

[0089] The blocks corresponding to the leaf nodes of the quadtree 703 can be used as a basis for prediction. That is, prediction can be performed for each of the blocks 702-1, 702-5, 702-6, 702-7, 702-8, 702-3, and 702-4 (referred to herein as coded blocks). As mentioned with respect to Figure 6 the coded blocks can be luminance blocks or chrominance blocks. It should be noted that in the example, the super-block partitioning can be determined with respect to the luminance blocks. The same partitioning can be used with the chrominance blocks.

[0090] At the coded block (e.g., block 702-1, 702-5, 702-6, 702-7, 702-8, 702-3, or 702-4) level, the prediction type (e.g., intra-frame or inter-frame prediction mode) is determined. That is, the coded block is a decision point for prediction.

[0091] In some embodiments, the coding efficiency based on blocks can be improved by partitioning a current residual block into one or more transform partitions, and the one or more transform partitions can be rectangular, including square partitions for transform coding. For example, a current residual block such as block 610 can be a 64×64 block and can be executed without partitioning using a 64×64 transform. In an example, a current residual block such as block 610 can be a 32×32 block and can be executed without partitioning using a 32×32 transform.

[0092] Although not explicitly shown in Figure 6 a residual block can be transform partitioned into transform sub-blocks. For example, a 64×64 residual block can be transform partitioned into four 32×32 transform blocks, including sixteen 16×16 transform blocks, using a transform partitioning scheme including sixty-four 8×8 transform blocks or a unified transform partitioning scheme including 256 4×4 transform blocks.

[0093] In some embodiments, video coding using a dual-criterion block splitting heuristic can include identifying the transform block size of a residual block. In some embodiments, transform partition coding can include recursively determining whether to transform a current block using a current block size transform or by partitioning the current block and performing transform partition coding on each partition.

[0094] For example, Figure 6 the lower left block 610 shown in

[0095] In some embodiments, determining the transform partition of a current transform block can be based on evaluating the respective split indicator values of sibling partitions. For example, if the split indicator value of an N×N transform block indicates that the block should be split. The possible transform partitions can be four N / 2×N / 2 partitions (i.e., the first partition, a transform partition including four N / 2×N / 2 transform sub-blocks), two N / 2×N partitions (i.e., the second partition, a transform partition including two N / 2×N transform sub-blocks), or two N×N / 2 partitions (i.e., the third partition, a transform partition including two N / 2×N / 2 transform sub-blocks). First, second, and third split indicator values can be calculated for the first, second, and third transform partitions, respectively. Then, the transform partition corresponding to the best of the respective split indicator values can be selected. The split indicator value of a transform partition can be an aggregation of the split indicator values of the transform sub-blocks of the transform partition. The aggregation can be the sum, maximum value, or some other aggregation of the split indicator values of the transform sub-blocks of the partition.

[0096] In an example, a split indicator value can be determined for the current block (i.e., the current transform block), and a corresponding split indicator value can be determined for each of the possible transform partitions of the current block. If the split indicator value of the current transform block is better than the split indicator values of the sub-partitions, the current block is not split; otherwise, the current block is split according to the partition with the smallest (e.g., best) split indicator value.

[0097] For illustration, for the shown lower left 64×64 block 610, the segmentation indicator value for encoding the 64×64 block 610 using a 64×64 size transform may indicate that the lower left 64×64 block 610 should be segmented. Thus, the aggregation of the segmentation indicator values for encoding the four 32×32 sub-blocks 620 using a 32×32 transform is determined. The segmentation indicator value for encoding the upper left 32×32 sub-block 620 using a 32×32 transform may indicate that the upper left 32×32 sub-block 620 should not be segmented, and the upper left 32×32 sub-block 620 may be coded using a 32×32 transform. Similarly, the segmentation indicator value for encoding the upper right 32×32 sub-block 620 using a 32×32 transform may be less than the sum of the costs for encoding the upper right 32×32 sub-block 620 using four 16×16 transforms, and the upper right 32×32 sub-block 620 may be coded using a 32×32 transform. Similarly, the segmentation indicator value for encoding the lower left 32×32 sub-block 620 using a 32×32 transform may be less than the aggregation of the segmentation indicator values for encoding the lower left 32×32 sub-block 620 using four 16×16 transforms, and the lower left 32×32 sub-block 620 may be coded using a 32×32 transform. The segmentation indicator value for encoding the lower right 32×32 sub-block 620 using a 32×32 transform may exceed the aggregation of the segmentation indicator values for encoding the lower right 32×32 sub-block 620 using four 16×16 transforms, and the lower right 32×32 sub-block 620 may be partitioned into four 16×16 sub-blocks 630, and each 16×16 sub-block 630 may be coded using a multi-form transform partitioning code.

[0098] Figure 8 FIG. 800 is a flow chart of a technique 800 for partitioning an image block to reduce quantization artifacts according to an embodiment of the present disclosure. The technique 800 determines whether a block should be transform partitioned using a dual-criterion block segmentation heuristic. The first criterion estimates the expected entropy; and the second criterion is an estimate of the impact of the artifacts in the sub-blocks on the adjacent sub-blocks of the block. The technique 800 determines whether to transform the block to the frequency domain using a transform of the size of the block, or whether the block should be partitioned into smaller sub-blocks, and each sub-block is individually transformed using a transform of the size of the sub-block or further partitioned using the technique 800.

[0099] The image can be a single image or a frame of a video sequence. A block includes pixels, and each pixel has a corresponding pixel value. The pixel value can be a luminance Y value, a chrominance U value, a chrominance V value, or other color component values. In an example, the block can be the maximum coded block size (e.g., super block, macro block, etc.). In an example, the maximum coded block size can have a size of 32×32, 64×64, 128×128, smaller or larger. In an example, the block can be a smaller sub-block of the maximum coded block. For illustration, the block can be a 16×16, 32×16, 16×32, 8×16, 16×8, or some other sub-block of the 32×32 maximum coded block size.

[0100] Technique 800 can be implemented, for example, as a software program executable by a computing device, such as Figure 1 the computing and transmitting station 102 or the receiving station 106. The software program can include machine-readable instructions that can be stored in a memory such as memory 204 and, when executed by a processor such as CPU 202, can cause the computing device to execute technique 800.

[0101] Technique 800 can be in or implemented by an encoder, and the operations of technique 800 can be fully or partially implemented in Figure 4 the transform unit 420, the dequantization unit 450, the inverse transform unit 460, the reconstruction unit 470, other units, or any combination thereof of the encoder 400.

[0102] Technique 800 can be implemented using dedicated hardware or firmware. Multiple processors, memories, or both can be used.

[0103] At 802, technique 800 estimates the expected entropy of the block. Technique 800 can estimate the expected entropy required to have a specified quality in the block. The expected entropy of the block can be approximated as described below.

[0104] In an example, technique 800 can calculate the Laplacian of the pixels in the block and take the maximum absolute value. That is, the maximum absolute Laplacian value can be used as an estimate of the expected entropy. Thus, estimating 802 the expected entropy of the block can include determining the corresponding Laplacian of the pixels of the block and selecting the maximum value of the absolute Laplacian as the expected entropy of the block.

[0105] The Laplacian operator can indicate the high-frequency energy in a block. Statistically, the high-frequency energy in a block can be mapped to at least most of the high-frequency elements (i.e., transform coefficient positions) of the transformed block. When there is energy in the high-frequency elements, the values (i.e., transform coefficient values) at those high-frequency elements (i.e., at the transform coefficient positions) are non-zero. When the high-frequency elements of a transform (e.g., DCT transform) are non-zero, storing these transform coefficient values can be costly (in terms of bits in the compressed bitstream). For re-iteration, the Laplacian operator can signal the highest-frequency elements of the transform (e.g., DCT transform).

[0106] Figure 9 900 is an example of calculating the Laplacian operator according to an embodiment of the present disclosure. For simplicity, the calculation of the Laplacian operator is described with respect to a block 902 of size 8×8. However, the block can have any rectangular size, including square size. Block 902 can be a luminance block, a chrominance block, or some other color component block. In the example, in the case where the block is a luminance block, the pixel values of the block can be represented by 8-bit values. Thus, the pixel values of block 902 can be in the range [0, 255].

[0107] In the example, determining the respective Laplacian operator of the pixels of the block can include: for the pixels in the rows and columns of the block, subtracting a first value from the pixels, the first value being a first function of the first average of all the pixels in the row; and subtracting a second value from the pixels, the second value being a second function of the second average of all the pixels in the column. For example, for each pixel at the position (x, y) in block 902, subtracting the function of the average pixel value of column x and the function of the average pixel value of row y from the pixel value at the position (x, y). The Laplacian operator of a pixel (i.e., pixel value) can be considered as subtracting the average of the adjacent pixels of the pixel from the pixel. This can be equivalent to applying a high-pass filter to the pixel. The Laplacian operator of all the pixels of block 902 is a block (e.g., Laplacian block 918 or Laplacian block 922) that describes the rapid changes in block 902.

[0108] The row average 904 is a column vector including the average value of each row of block 902. The row average can be rounded to the nearest integer. For example, the average value 906 (i.e., 147) is the average pixel value of row 908. Thus, the average value 906 is calculated as round((85 + 65 + 116 + 196 + 138 + 139 + 183 + 250) / 8) = 147. The column average 910 is a row vector including the average value of each column of block 902. The column average can be rounded to the nearest integer. For example, the average value 912 (i.e., 166) is the average pixel value of column 914. Thus, the average value 912 is calculated as round((200 + 68 + 138 + 230 + 241 + 87 + 151 + 214) / 8) = 166.

[0109] In an example, the function can be a unitary function. A unitary function in this article means that the average value is taken as it is or multiplied by 1. Thus, the Laplacian 920 of pixel 916 is calculated in the Laplacian block 918 as round((138 - 147 - 166) = -175). In an example, the function can multiply the average value by 1 / 2. Thus, the Laplacian 924 of pixel 916 is calculated in the Laplacian block 922 as round((138 - 147 / 2 - 166 / 2) = -19). It should be noted that by subtracting half of the average value, the integral (i.e., the sum of the values in the Laplacian block 922) will actually have a zero effect. That is, the sum of the values is very close to zero. Subtracting half of the average value has already shown good results. However, other fractional values can be subtracted. The function of the average value can be different from the unitary function of the 1 / 2 function.

[0110] If the row and column averages are not removed from the pixels of the block, a non-trivial entropy weight can be set for the first row and the first column of the DCT (i.e., the transform block) because these averages can be mapped in the first row and the first column of the transform block. To reiterate, removing the horizontal and vertical integrals (i.e., row and column averages) from the pixels can allow for a more realistic entropy estimate because these integrals project onto the first row and the first column and the total entropy is lower than that of other high-frequency elements.

[0111] For further illustration, assume the image is an image of the horizon, where above the horizon is the sky and below the horizon is the lake. Thus, in a single block (e.g., a 3×32 block), there can be a single line defining the division between the sky and the lake. When this block is transformed using a 32×32 DCT, it includes 1,024 elements and is a large amount of data to store. However, by removing the average value, only the horizon (i.e., a single line) remains, and the integral transform will not use all 32×32 elements. The transform only uses 32 elements: the transform only places all the values in the first column (for a horizontal line, such as the described horizon) or the first row (for a vertical line). Thus, a significantly reduced entropy (and data to store) is obtained. Therefore, in order to reduce the entropy in the estimation of the Laplacian, the horizontal and vertical integrals can be removed.

[0112] The maximum absolute value of the Laplacian block can be used as the expected entropy.

[0113] Return to Figure 8, in an example, estimating the expected entropy of block 802 can include encoding the block using a hypothetical encoder and using the multiple bits resulting from the hypothetical encoding as the expected entropy of the block. The hypothetical encoding process is a process that performs the coding steps but does not output the bits to the compressed bitstream. Since the aim is to estimate the bit rate (or simply the rate), the hypothetical encoding process can be regarded as or referred to as a rate estimation process. The hypothetical encoding process calculates the number of bits required to encode the region. In another example, the expected entropy of a block can be estimated (e.g., approximated) using heuristics known to be positively correlated with the calculation of the final entropy of the block. For example, the Shannon entropy of the block can be calculated.

[0114] At 804, technique 800 partitions the block into sub-blocks. The size of each sub-block can be the smallest possible partition size. In an example, the smallest possible partition size of a luminance block can be 8×8. In an example, the smallest possible partition size of a luminance block can be 5×5. In an example, the smallest possible partition size of a chrominance block can be 4×4. For at least transform partitioning purposes, a block with the smallest possible partition size cannot be further partitioned.

[0115] At 806, technique 800 calculates the corresponding visual masking amount of the sub-blocks. That is, for each sub-block, technique 800 calculates the visual masking amount. That is, technique 800 estimates (e.g., approximates, correlates, etc.) how easily artifacts are visible in the reconstructed sub-block of that sub-block.

[0116] For each sub-block, technique 800 can calculate the integral of all pixels in the sub-block using formula (1)

[0117] 1 / (1 + x * x / (C * C)). (1)

[0118] In formula (1), x represents (e.g., indicates, describes, etc.) the local increment; and the constant C represents the amount of local increment expected to exist when compressing the sub-block at the current quality. "Local increment" refers to the difference between adjacent pixels. For example, the local increment can be the magnitude of the Laplacian operator. In an example, when calculating the integral of all pixels in the sub-block using formula (1), x can be the maximum absolute Laplacian operator in the sub-block. In another example, x can be calculated as the sum of all Laplacian operators of the sub-block.

[0119] The integral value can be compared with an empirically determined threshold (e.g., 5 or some other value). The empirically determined threshold can be derived using optimization techniques such as the Nedler-Mead optimization technique. A high value of the integral can indicate that not many pixels in the sub-block will participate in visual masking. Therefore, any artifacts propagated to the sub-block from adjacent sub-blocks may be visible.

[0120] In another example, the constant C can be a multiple of the butteraugli limit of a sub-block. The butteraugli limit estimates the psychovisual similarity of two images. It scores images reliably in domains with little apparent difference and computes a spatial map of the level of difference. The butteraugli limit can be used as a quality target for lossy image and video compression. Butteraugli can be used as a quality assessor (i.e., metric); however, the exact value of the butteraugli limit can be used to define the quality target for lossy image and video compression. Thus, the butteraugli limit can be an estimate of the quality difference between a source sub-block (e.g., a pre-coded sub-block) and a reconstructed sub-block (e.g., a post-decoded sub-block). Tools for computing the butteraugli limit can be obtained from https: / / opensource.google / projects / butteraugli. The butteraugli limit is an objective image quality assessment metric. It assigns a difference mean opinion score (DMOS) value to the difference between an original image and a degraded version. To reiterate, the constant C can be computed as a multiple of the quality of a sub-block. The quality of a sub-block can be defined as a multiple of the just noticeable difference in the sub-block. That is, the constant C can be a multiple of the butteraugli limit.

[0121] C is a constant that depends on the desired (e.g., selected, etc.) quality level. It is well known that the target quality level is typically specified when encoding an image. Thus, C can be related to the target quality level. It is well known that lower quality results in fewer bits in the compressed bitstream compared to higher target quality. The quality level can be or can be specified as a quantization parameter QP. Thus, the constant C can be a function of QP. In an example, C can be a linear function (e.g., multiple) of QP.

[0122] At 808, the technique 800 selects the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block. For example, if the size of the block (e.g., 32×32) is larger than the smallest possible partition size (e.g., 8×8), then the worst visual masking feature of all the sub-blocks of the block can be considered as the visual masking feature of the block. If the block has the smallest possible partition size (e.g., 8×8), then the visual masking feature is taken.

[0123] At 810, technique 800 combines the visual masking characteristics of a block with the expected entropy of the block to obtain a segmentation indicator value. That is, the entropy required to achieve a target quality can be combined with the worst-case visual masking to obtain a final segmentation decision. As described above, the masking characteristics of a block can be the worst-case visual masking of all sub-blocks of the block. In an example, a linear model can be used to combine the masking characteristics of a block with the expected entropy of the block. In an example, the segmentation indicator value can be obtained by multiplying the visual masking of the block by the desired entropy of the block.

[0124] At 812, technique 800 determines whether to segment the block based on the segmentation indicator value. In an example, the segmentation indicator value can be compared to a predefined constant determined empirically (e.g., 5). If the segmentation indicator value is less than the predefined constant, the block is segmented; otherwise, the block is not segmented.

[0125] The transform segmentation (partitioning) decision for a block is a decision that naturally results in an amount of ringing suitable for the sub-blocks of the block. Transform segmentation can be beneficial for larger transform sizes. To illustrate, assume the block is a 32×32 block. If the amount of ringing for the 32×32 block (e.g., as calculated at 806 of technique 800) is acceptable for the masking (as calculated at 810 of technique 800), the 32×32 transform is selected. On the other hand, if the amount of ringing is unacceptable, the possible partitions of the block are each individually tested using technique 800. For example, each of a first partition consisting of two 16×32 blocks and a second partition consisting of two 32×16 blocks is tested. If the corresponding amount of ringing is unacceptable in the first and second partitions, a third partition consisting of four 16×16 blocks is tested. Next, the following partitions can be tested: a fourth partition consisting of four 32×8 blocks, a fifth partition consisting of eight 16×8 blocks, a sixth partition consisting of eight 8×16 blocks, and a seventh partition consisting of sixteen 8×8 sub-blocks. That is, the block is partitioned into increasingly smaller partitions until an acceptable amount of ringing result and / or the smallest possible partition is reached. As can be appreciated, there can be many different ways to test the partitioning of a block.

[0126] The partitions of a block (or sub-block) can be tested in order of size, and the largest partition that produces an acceptable value of the segmentation indicator can be selected as the partition of the block.

[0127] Figure 11 An example of an assumed comparison of a transform partition 1100 according to a conventional technique and a transform partition 1110 according to an embodiment of the present disclosure is illustrated. Figure 11 The examples are for illustrative purposes only and do not limit the teachings of the present disclosure in any way.

[0128] In a transform partition 1100 according to the conventional art, a residual block (e.g., of size 32×32) can be partitioned into four transform sub-blocks 1102, 1104, 1106, and 1108 (e.g., each having a size of 16×16). The transform sub-block 1102 is further partitioned into sub-blocks 1102A, 1102B, 1102C, and 1102D (e.g., each having a size of 8×8). The transform sub-block 1104 can be further partitioned into sub-blocks 1104A, 1104B, 1104C, and 1104D (e.g., each having a size of 4×16). The transform sub-block 1106 is not further partitioned. The transform sub-block 1108 is further partitioned into sub-blocks 1108A and 1108B (e.g., each having a size of 16×8). The transform partition 1100 illustrates that a larger block (or sub-block) can be further divided as it is determined to contain smooth sub-blocks.

[0129] In contrast, a transform partition 1110 according to an embodiment of the present disclosure results in the block being partitioned into sub-block 1102 (e.g., having a size of 16×16), sub-block 1110 (e.g., having a size of 32×16), and sub-block 1106 (e.g., having a size of 16×16). Thus, the lossy compression two-criterion block splitting heuristic described above for determining whether a smooth sub-block should be further split can result in transform sub-blocks of a larger size.

[0130] For simplicity of illustration, the technique 800 is depicted and described as a series of steps or operations. However, the steps or operations according to the present disclosure can occur in various orders and / or simultaneously. Additionally, other steps or operations not presented and described herein can be used. Furthermore, not all of the illustrated steps or operations may be required to implement the method according to the disclosed subject matter.

[0131] The terms "example" or "exemplary" are used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "example" or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs. Instead, the use of the terms "example" or "exemplary" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X includes A or B" is intended to mean any of its natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then "X includes A or B" is satisfied in any of the foregoing instances. Additionally, as used in this application and the appended claims, the words "a" and "an" generally should be construed to mean "one or more" unless otherwise specified or clearly indicated as singular from the context. Further, the use of the terms "embodiment" or "an embodiment" or "implementation" or "an implementation" throughout the text is not intended to mean the same embodiment or implementation unless so described. As used herein, the terms "determine" and "identify" or any variant thereof include selecting, verifying, calculating, looking up, receiving, determining, establishing, obtaining, or otherwise identifying or determining using one or more of the devices shown in Figure 1 the manner shown.

[0132] In addition, for simplicity of explanation, although the figures and description herein may include a sequence or series of operations or stages, the elements of the methods disclosed herein can occur in various sequences and / or concurrently. Additionally, the elements of the methods disclosed herein can occur with other elements not explicitly presented and described herein. Further, one or more elements of the methods described herein can be omitted from an implementation of the methods according to the disclosed subject matter.

[0133] Embodiments of the transmitting station 102 and / or the receiving station 106 (and algorithms, methods, instructions, etc. stored thereon and / or executed thereby) can be implemented in hardware, software, or any combination thereof. The hardware can include, for example, a computer, an intellectual property (IP) core, an application specific integrated circuit (ASIC), a programmable logic array, an optical processor, a programmable logic controller, microcode, a microcontroller, a server, a microprocessor, a digital signal processor, or any other suitable circuitry. In the claims, the term "processor" should be understood to include any of the foregoing hardware individually or in combination. The terms "signal" and "data" are used interchangeably. Additionally, portions of the transmitting station 102 and the receiving station 106 need not be implemented in the same manner.

[0134] In addition, in one embodiment, for example, the transmitting station 102 and the receiving station 106 can be implemented using a computer program which, when executed, implements any one of the corresponding methods, algorithms, and / or instructions described herein. Additionally or alternatively, for example, a dedicated computer / processor can be utilized, the dedicated computer / processor being capable of including dedicated hardware for implementing any one of the methods, algorithms, or instructions described herein.

[0135] The transmitting station 102 and the receiving station 106 can be implemented, for example, on a computer in a real-time video system. Alternatively, the transmitting station 102 can be implemented on a server, while the receiving station 106 can be implemented on a device separate from the server (such as a handheld communication device). In such a case, the transmitting station 102 can use an encoder 400 to encode the content into an encoded video signal and transmit the encoded video signal to the communication device. The communication device can then, in turn, use a decoder 500 to decode the encoded video signal. Alternatively, the communication device can decode content locally stored on the communication device (e.g., content not transmitted by the transmitting station 102). Other suitable embodiments of the transmitting station 102 and the receiving station 106 are available. For example, the receiving station 106 can be a generally fixed personal computer instead of a portable communication device, and / or a device including the encoder 400 can also include the decoder 500.

[0136] Furthermore, all or part of the embodiments can take the form of a computer program product accessible from, for example, a tangible computer-usable or computer-readable medium. The computer-usable or computer-readable medium can be any device that can tangibly contain, store, transmit, or convey a program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media are also available. The foregoing embodiments have been described so as to enable easy understanding of the present application and not by way of limitation. On the contrary, the present application covers various modifications and equivalent arrangements included within the scope of the appended claims, which scope should be given the broadest interpretation permitted by law to cover all such modifications and equivalent arrangements.

Claims

1. A method for partitioning transform blocks of an image to reduce quantization artifacts, the method comprising: estimating the expected entropy of a block of the image; partitioning the block into sub-blocks; calculating a corresponding visual masking amount for the sub-blocks, wherein the visual masking amount for a corresponding sub-block is: calculated based on the texture present in the corresponding sub-block; and indicates the visibility of artifacts in the corresponding sub-block; selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; combining the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value, wherein the combination includes combining the visual masking feature of the block with the expected entropy using a linear model, or multiplying the visual masking feature with the expected entropy of the block; and determining whether to partition the transform block corresponding to the block based on the segmentation indicator value.

2. The method according to claim 1, wherein, The block has a size of 32×32, and wherein the minimum possible partition size is 8×8.

3. The method according to claim 1, wherein, Estimating the expected entropy of the block includes: determining the corresponding Laplacian operator of the pixels of the block; and selecting the absolute maximum value of the corresponding Laplacian operator as the expected entropy of the block.

4. The method according to claim 3, wherein, Determining the corresponding Laplacian operator of the pixels of the block includes: for the pixels in the rows and columns of the block, subtracting a first value from the pixels and subtracting a second value from the pixels, the first value being a first function of the first average of all the pixels in the row, and the second value being a second function of the second average of all the pixels in the column.

5. The method according to claim 4, wherein The first value is half of the first average, and wherein the second value is half of the second average.

6. The method according to claim 1, wherein Estimating the expected entropy of the block includes: encoding the block using a hypothesis encoder, wherein the hypothesis encoder encodes the block to calculate a plurality of bits required to encode the block; and using the plurality of bits generated from the hypothesis encoder as the expected entropy of the block.

7. The method according to claim 1, wherein, Calculating the corresponding visual masking amount for the sub-blocks includes: calculating the corresponding visual masking amount for a given sub-block using the formula 1 / (1 + x*x / (C*C)), wherein x is a local increment, and wherein C is a constant related to the quality level.

8. The method according to claim 7, wherein x is the maximum absolute Laplacian operator in the sub-block.

9. The method according to claim 7, wherein x is the sum of the corresponding Laplacian operators of the pixels of the sub-block.

10. An apparatus for partitioning transform blocks of an image to reduce quantization artifacts, comprising: a memory; and a processor configured to execute instructions stored in the memory for: estimating the expected entropy of a block of the image; partitioning the block into sub-blocks; calculating a corresponding visual masking amount for the sub-blocks, wherein the corresponding visual masking amount for a corresponding sub-block is calculated based on the texture present in the corresponding sub-block and indicates the visibility of artifacts in the corresponding sub-block; selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; Combining the visual masking feature of the block with the expected entropy of the block using a linear model or by multiplying the visual masking feature of the block by the expected entropy of the block to obtain a segmentation indicator value; and Determining whether to partition the transform block corresponding to the block based on the segmentation indicator value.

11. The apparatus according to claim 10, wherein, The block has a size of 32×32, and wherein, the size of each sub-block is 5×5, and the size of each sub-block is used as the visual masking sub-block size.

12. The device according to claim 10, wherein, Estimating the expected entropy of the block includes: Determining the corresponding Laplacian operator of the pixels of the block; and Selecting the absolute maximum value of the corresponding Laplacian operator as the expected entropy of the block.

13. The device according to claim 12, wherein, Determining the corresponding Laplacian operator of the pixels of the block includes: For the pixels in the rows and columns of the block, subtracting a first value from the pixels and subtracting a second value from the pixels, where the first value is a first function of the first average of all the pixels in the row, and the second value is a second function of the second average of all the pixels in the column.

14. The device according to claim 13, wherein, The first value is half of the first average, and wherein, the second value is half of the second average.

15. The apparatus according to claim 10, wherein, Estimating the expected entropy of the block includes: Encoding the block using a hypothesis encoder, where the hypothesis encoder encodes the block to calculate a plurality of bits required to encode the block; and Using the plurality of bits generated from the hypothesis encoder as the expected entropy of the block.

16. The apparatus according to claim 10, wherein, Calculating the corresponding visual masking amount of the sub-block includes: Calculating the corresponding visual masking amount of a given sub-block using the formula 1 / (1 + x*x / (C*C)), where x indicates a local increment, and where C is a constant related to the quality level.

17. The apparatus according to claim 16, wherein, x is the maximum absolute Laplacian operator in the sub-block.

18. The device according to claim 16, wherein, x is the sum of the corresponding Laplacian operators of the pixels of the sub-block.

19. A non-transitory computer-readable storage medium including executable instructions that, when executed by a processor, facilitate the execution of operations, the operations including: Partitioning a transform block of an image to reduce quantization artifacts, including: Estimating the expected entropy of the blocks of the image; Partitioning the block into sub-blocks; Calculating the corresponding visual masking amount of the sub-blocks, where the corresponding visual masking amount of the corresponding sub-block is calculated based on the texture present in the corresponding sub-block and indicates the visibility of artifacts in the corresponding sub-block; Selecting the highest visual masking value among the corresponding visual masking amounts of the sub-blocks as the visual masking feature of the block; Combining the visual masking feature of the block with the expected entropy of the block to obtain a segmentation indicator value, where the combination includes using a linear model or multiplying the visual masking feature by the expected entropy of the block; and Determining whether to partition the transform block corresponding to the block based on the segmentation indicator value.