Adaptive quantization and deadband modulation

By using adaptive quantization and block/band adaptive quantization methods, combined with symbolic representation and noise addition, the image coding process is optimized, solving the image quality problem that existing image compression methods struggle to achieve high compression ratios and low computational complexity, thus realizing efficient image coding.

CN114556930BActive Publication Date: 2026-04-14GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-14
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing lossy image compression methods struggle to achieve high compression ratios while maintaining image quality, and traditional adaptive quantization methods have high computational complexity during encoding, making them difficult to apply directly to existing decoders.

Method used

By adaptively selecting quantization levels and scaling factors, combining block and band adaptive quantization, identifying and adjusting the maximum coefficients to ensure image quality, applying symbolic representation and noise addition to optimize the encoding process, and generating an encoded version.

Benefits of technology

It achieves a significant increase in compression ratio while maintaining image quality and reduces the computational complexity of the encoding process, enabling existing decoders to directly decode encoded images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114556930B_ABST
    Figure CN114556930B_ABST
Patent Text Reader

Abstract

Methods are provided for improving the quality and compression factor of compressed images. These methods include determining band-specific quantization levels on a block-by-block level. This results in adaptive dead-zones, allowing certain blocks to be represented by fewer non-zero elements while other blocks are represented by more non-zero elements. Thus, the quality of the encoded image is improved while maintaining or improving the compression ratio. The adaptive quantization levels are determined by comparing the post-quantization energy level to a threshold energy criterion for each band within the block. In cases where the energy threshold criterion is not met via these methods, additional methods can be applied to improve the image quality. The methods described herein allow for adapting the effective compression ratio of an image on a block-by-block, frequency-sensitive basis in order to more effectively allocate the encoded image bits that will have the greatest impact on the image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Digital images can be compressed to provide advantages such as reduced storage and / or transmission costs. Various lossy and lossless methods exist for image compression. Lossy image compression methods result in a compressed version of the input image that cannot be used to accurately recreate it. However, such lossy compression methods allow for the generation of output images that appear sufficiently similar to the input image to human perception, making them acceptable in at least some contexts. Some lossy image compression techniques allow for an increased compression ratio in exchange for this similarity, thus allowing for smaller compressed image file sizes in return for a reduction in the image quality of the output compressed image. Summary of the Invention

[0002] Various methods exist for compressing images. These methods allow for a reduction in image size while preserving the image's subjective appearance after the compression (or encoding) process. "Lossy" compression methods result in encoded images that cannot be used to perfectly reconstruct the source image. However, such lossy methods can be used to generate encoded image representations that can be used to reconstruct versions of the source image that most viewers cannot immediately distinguish from it, or versions of the source image that are otherwise acceptable in at least some contexts according to at least some criteria. Such lossy methods can provide significantly higher compression ratios compared to lossless compression methods.

[0003] Many types of lossy compression methods achieve these compression ratios by performing one or more quantization steps on the pixels themselves or on coefficients derived from them (e.g., coefficients corresponding to different spatial frequency components of the pixel). Such quantization steps allow the bit depth of the coded version of the image information (e.g., pixels, or coefficients derived from them) to be reduced, thereby reducing the size of the coded version of the image. Additionally, the image information can be scaled before quantization, such that many scaled image coefficients become "0" after quantization. If this proportion of "0" is high enough, the size of the coded version of the image can be further reduced by utilizing this high proportion of "0" in the scaled and quantized version of the source image using run-length coding or other methods.

[0004] A scaling factor can be selected to scale the spatial frequency coefficients of the source image to improve the quality and / or compression ratio of the image encoded using that scaling factor. The threshold used to quantize the coefficients scaled by this scaling factor set can also be adjusted to improve the compression ratio and / or improve the quality of the encoded image. For example, adjusting the quantization threshold used to determine whether a particular coefficient is quantized to a "0" value allows for direct adjustment of the number of "0"s, and thus the compressibility of the image. Various different quantization levels can be evaluated (e.g., regarding the quality of the encoded image, regarding the residual image energy in one or more spatial frequency bands), and the lowest quantization level that satisfies the image quality threshold can be used to encode the image, or a subset of the coefficients of the encoded image. Selecting the quantization level in this way can advantageously produce the most "0" values ​​after quantization (in order to minimize the size of the encoded image) while still satisfying image quality criteria. Adaptively setting the quantization level during encoding in this way can also advantageously avoid adaptive computational steps involved in the decoder used to decode the encoded image.

[0005] Additional or alternative methods can be applied to improve the compression ratio of an image while preserving or improving the quality of the encoded image. These methods can be implemented when the aforementioned adaptive quantization methods fail (e.g., when no evaluated quantization level results in meeting image quality criteria), or they can be implemented in other contexts. For example, the source image can be encoded using only a single quantization level, but if that single quantization level fails to meet image quality criteria, one or more supplementary methods may be applied. A first supplementary method includes: for a subset of image coefficients within a specific spatial frequency band, determining which coefficients are the largest, and setting a local quantization level such that the largest coefficients are quantized to non-zero values. A second supplementary method includes selecting a different set of scaling factors to scale the image content before quantization. A third supplementary method includes using symbols representing values ​​of 1 / 2, 1 / 4, or some other subunit value to represent one or more coefficients. A fourth supplementary method includes adding noise (e.g., blue noise) to the image coefficients before quantization, or adjusting the amount of noise added to the image coefficients before quantization.

[0006] This disclosure relates to a method for encoding an image, the method comprising: (i) generating a set of coefficients indicating image content of pixel blocks at a plurality of spatial frequencies based on pixel blocks of the image, wherein each coefficient in the set of coefficients is for a corresponding spatial frequency among the plurality of spatial frequencies; (ii) scaling the set of coefficients according to a first set of scaling factors to generate a first scaled set of coefficients; (iii) performing an evaluation of a plurality of quantization levels; (iv) generating a scaled and quantized version of the set of coefficients based on the evaluation of the plurality of quantization levels; and (v) generating an encoded version of the image based on the scaled and quantized version of the set of coefficients. Performing the evaluation comprises: for each corresponding quantization level among the plurality of quantization levels: (a) quantizing a subset of the first scaled set of coefficients according to the corresponding quantization level to generate a quantized subset of the first scaled set of coefficients, wherein the subset of the first scaled set of coefficients is for spatial frequencies within a first spatial frequency band; and (b) determining a quantized energy of the quantized subset of the first scaled set of coefficients. The evaluation may further comprise: (c) comparing the quantized energy with a threshold energy criterion.

[0007] The evaluation of the plurality of quantization levels can be performed iteratively from the maximum quantization level to the minimum quantization level. The maximum quantization level can be greater than 0.675. The maximum quantization level can be greater than 0.575. The minimum quantization level can be less than 0.325. Comparing the quantized energy with the threshold energy criterion for a specific quantization level can include determining that the quantized energy satisfies the threshold energy criterion. Generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels can include quantizing a subset of a first set of scaling coefficients for spatial frequencies within the first spatial frequency band using the specific quantization level. Performing the evaluation of the plurality of quantization levels can include determining that none of the plurality of quantization levels satisfies the threshold energy criterion. Generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels can include, in response to determining that none of the plurality of quantization levels satisfies the threshold energy criterion: identifying the maximum scaling coefficient within the subset of the first set of scaling coefficients for spatial frequencies within the first spatial frequency band; determining a quantization level with a value less than the identified maximum scaling coefficient; and using the determined quantization level to quantize the subset of the first set of scaling coefficients for spatial frequencies within the first spatial frequency band. The method may further include: determining a first set of scaling factors based on pixel blocks of the image. Performing an evaluation of the plurality of quantization levels may include determining that none of the plurality of quantization levels satisfies the threshold energy criterion. Generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels may include, in response to determining that none of the plurality of quantization levels satisfies the threshold energy criterion: determining a second set of scaling factors, wherein at least one scaling factor in the second set of scaling factors is magnitude lower than a corresponding scaling factor in the first set of scaling factors; performing the evaluation of the plurality of quantization levels using the second set of scaling factors; and generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels using the second set of scaling factors. Performing the evaluation of the plurality of quantization levels may include determining that none of the plurality of quantization levels satisfies the threshold energy criterion. Generating an coded version of the image based on the scaled and quantized version of the coefficient set may include, in response to determining that none of the plurality of quantization levels satisfies the threshold energy criterion, representing at least one coefficient in a subset of the first set of scaling coefficients for spatial frequencies within the first spatial band using a symbol representing a subunit value. The symbol representing the subunit value can represent one of a half value, a quarter value, or an eighth value. Determining that none of the plurality of quantization levels satisfies the threshold energy criterion can include determining that each of the plurality of quantization levels results in a quantized energy of zero.The method may further include: determining the energy of a first subset of a set of coefficients indicating image content, wherein the first subset of the set of coefficients indicating image content indicates a corresponding spatial frequency within the first spatial frequency band. Comparing the quantized energy with a threshold energy criterion may include comparing the ratio between the quantized energy and the energy of the first subset of the coefficient set with a threshold energy. The first spatial frequency band may include low spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction. The first spatial frequency band may include low spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction. The first spatial frequency band may include high spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction. The first spatial frequency band may include high spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction. The set of coefficients indicating image content may be discrete cosine transform coefficients or coefficients of some other transform (e.g., overcomplete transform).

[0008] Another aspect of this disclosure relates to a method for encoding an image, the method comprising: (i) generating a set of coefficients indicating image content of transformed pixel blocks at a plurality of spatial frequencies based on pixel blocks of the image, wherein each coefficient in the set of coefficients is used for a corresponding spatial frequency among the plurality of spatial frequencies; (ii) scaling a subset of the set of coefficients according to a first set of scaling factors to generate a first scaled subset of coefficients, wherein the subset of the set of coefficients is used for a corresponding spatial frequency within a first spatial frequency band; (iii) quantizing the first scaled subset of coefficients to generate a first scaled quantized set of coefficients; (iv) determining a quantized energy of the first scaled subset of quantized coefficients; (v) determining that the quantized energy does not satisfy a threshold energy criterion; and (vi) in response to determining that the quantized energy does not satisfy the threshold energy criterion, applying at least one process from a set of processes to generate an encoded version of the image. The set of processes includes: a first process comprising: (a) identifying a maximum scaling factor within the first subset of scaling coefficients; (b) determining a quantization level less than the identified maximum scaling factor; and (c) quantizing the first subset of scaling coefficients using the determined quantization level. The set of processes includes a second process, which includes: (a) determining a second set of scaling factors, wherein at least one scaling factor in the second set of scaling factors is magnitude lower than a corresponding scaling factor in the first set of scaling factors; (b) scaling a subset of the coefficient set according to the second set of scaling factors to generate a second subset of scaling coefficients, wherein the subset of the coefficient set indicates a corresponding spatial frequency within the first spatial frequency band; and (c) quantizing the second subset of scaling coefficients to generate a second scaled quantized set of coefficients. The set of processes includes a third process, which includes using a notation representing a subunit value in the coded version of the image to represent at least one subset of the coefficient set indicating a corresponding spatial frequency within the first spatial frequency band. The set of processes includes a fourth process, which includes adding blue noise to the first subset of scaling coefficients before quantizing the first subset of scaling coefficients. The set of processes includes a fifth process, which includes adjusting the amount of blue noise added to the first subset of scaling coefficients before quantizing the first subset of scaling coefficients.

[0009] The symbol representing the subunit value can represent one of a half value, a quarter value, or an eighth value. Determining that the quantized energy does not meet the threshold energy criterion can include determining that the quantized energy is zero. The first spatial frequency band can include low spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction. The first spatial frequency band can include low spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction. The first spatial frequency band can include high spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction. The first spatial frequency band can include high spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction. The set of coefficients indicating image content can be discrete cosine transform coefficients or coefficients indicating some other transform (e.g., overcomplete transform). Blue noise added to the first subset of scaling coefficients can include band-limited blue noise limited to spatial frequencies within the first spatial frequency band.

[0010] It will be understood that the features described in the context of the first aspect can be implemented in the context of the second aspect. These, as well as other aspects, advantages, and alternatives, will become apparent to those skilled in the art by reading the following detailed description with appropriate reference to the accompanying drawings. Furthermore, it should be understood that the descriptions provided in the summary section and elsewhere in this document are intended to illustrate the claimed subject matter by way of example rather than limitation. Attached Figure Description

[0011] Figure 1A An example image is shown.

[0012] Figure 1B The diagram shows... Figure 1A An example of a portion of the image is based on frequency decomposition.

[0013] Figure 1C An example quantification table is illustrated.

[0014] Figure 1D The diagram shows... Figure 1B Frequency-based decomposition after scaling and quantization.

[0015] Figure 2 The illustration shows the example quantization level and example scaling factor values.

[0016] Figure 3A The illustration shows an example of frequency-based decomposition of a portion of an image after scaling and quantization.

[0017] Figure 3B The illustration shows an example of frequency-based decomposition of a portion of an image after scaling and quantization.

[0018] Figure 3C The illustration shows an example of frequency-based decomposition of a portion of an image after scaling and quantization.

[0019] Figure 4 This is a simplified block diagram showing some of the components of the example system.

[0020] Figure 5 This is a flowchart of a method according to an example embodiment.

[0021] Figure 6 This is a flowchart of a method according to an example embodiment. Detailed Implementation

[0022] This document describes examples of methods and systems. It should be understood that the terms “exemplary,” “example,” and “illustrative” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as “exemplary,” “example,” or “illustrative” is not necessarily to be construed as more preferred or advantageous than other embodiments or features. Furthermore, the exemplary embodiments described herein are not intended to be limiting. It will be readily understood that certain aspects of the disclosed systems and methods can be arranged and combined in a wide variety of different configurations.

[0023] It should be understood that the following embodiments and other embodiments described herein are provided for illustrative purposes and are not intended to be limiting.

[0024] Example image encoding and compression

[0025] In various applications, encoding images or other information to reduce their size can be beneficial. As a result of this encoding, less storage space and / or bandwidth can be used to store, transmit, copy, or otherwise manipulate or use the images or other information. Encoding (or compression) can be lossless or lossy. Lossless compression reduces the size of information in a way that allows it to be accurately restored later to its pre-compressed state. Lossy compression does not do this. Instead, lossy compression allows for a trade-off between the possible level of compression and the “quality” from which the image or other information can be later recovered from the compressed version.

[0026] This trade-off can be achieved based on the intended use of the compressed information. For example, when compressing an image, the compression method can take into account the characteristics of human vision (e.g., the eye's sensitivity to increases in luminance relative to chroma), allowing the compression process to discard or distort information from the image in ways less noticeable to the human eye. For instance, the encoding method can take into account the human eye's sensitivity to increases in luminance relative to chroma by downsampling the chroma information, by reducing the bit depth at which the chroma information is stored relative to the luminance information, and / or by encoding the chroma information relative to the luminance information using different, lower-quality quantization tables or other parameters. In another example, the higher spatial frequency content of the image can be quantized, rounded, or otherwise degraded during encoding to a greater extent than the lower spatial frequency content. Therefore, the size of the compressed image can be reduced while maintaining the overall level of apparent image quality.

[0027] Image encoding can be partially achieved by first transforming the image into different color representations. For example, an image with a red-green-blue (RGB) representation can be converted to a luminance-chrominance (YUV) representation. Alternatively or additionally, encoding may involve downsampling the source image (e.g., downsampling the chrominance channels), applying linear or nonlinear filters, quantizing / rounding the pixel values ​​of the image, or performing some other operation on the image "in image space" before applying any transformation of the image data from the image's two-dimensional pixel space to the spatial frequency space or some other space. This preprocessed image data can then be transformed to another domain, for example, the spatial frequency domain, where further compression may occur.

[0028] Such alternative spaces can be chosen such that the representation of an image in the alternative space is "sparse." That is, the representation in the alternative space may include a small subset of representative coefficients that contain "most" of the image (e.g., have substantially non-zero values) relative to the total energy in the image, while a larger, remaining subset of representative coefficients has values ​​at or near zero and thus represents a small portion of the image content of the source image. Therefore, the remaining subset can be discarded, thereby reducing the overall size of the encoded image while preserving most of the perceptible content and / or energy of the source image. Such a process may include quantizing or otherwise rounding (e.g., rounding down) the coefficients, for example, after a scaling process to emphasize those coefficients found to be more "important" to human visual perception (e.g., lower spatial frequencies). Such alternative spaces may include a spatial frequency space (e.g., represented by discrete cosine transform coefficients of the image spatial data), a kernel-based space, or some other transform space. Images may be represented in an overcomplete manner in the alternative space (e.g., with more coefficients than would be required for a strictly informative representation of the image).

[0029] Figure 1AThe illustration shows a specific color channel (e.g., luminance channel, chroma channel) of an example image 100 or such an image that can be encoded. Image 100 consists of multiple pixels (whose samples are...). Figure 1A The image 100 is composed of small squares (illustrated in the image). To encode (e.g., compress) image 100 (e.g., according to JPEG, JPEG XL, or some other image compression format), non-overlapping sets of pixels (e.g., ...) can be used. Figure 1A The example set 115 shown (e.g., using discrete cosine transform) is transformed individually into the corresponding set of coefficients in the transform domain. Performing this transformation on a restricted subset of the image rather than the entire image (e.g., generating discrete cosine transform coefficients from the entire image at once) can provide benefits regarding memory usage, encoder generalization (e.g., across images of different sizes), encoder optimization, or other considerations related to image encoding. As shown, the non-overlapping set is an 8x8 pixel patch, but other non-overlapping sets of different shapes and sizes can also be used.

[0030] The pixel set 115 illustrated in image 100 can be transformed into a set of coefficients representing the contents of the pixel set 115 at corresponding spatial frequencies. For example, the coefficients can be discrete cosine transform coefficients determined over the range of horizontal and vertical spatial frequencies. Figure 1B The diagram illustrates a set of coefficients 120. Each coefficient represents the content of pixel set 115 at a corresponding spatial frequency in both the vertical and horizontal directions. For example, the top-left coefficient ("-415.38") represents the DC content of pixel set 115. In another example, the fourth coefficient from the right in the top row ("56.12") represents the content of pixel set 115 that varies horizontally at a middle spatial frequency but not vertically (i.e., DC relative to the vertical direction).

[0031] Therefore, subsets of coefficients 120 can be defined based on whether each coefficient 120 is used for spatial frequencies within a specified spatial frequency band. For example, a first subset 125a of coefficients 120 is for spatial frequencies within a first spatial frequency band, which includes low spatial frequencies in the horizontal direction (e.g., the lower half of the horizontal spatial frequency range represented in coefficients 120) and low spatial frequencies in the vertical direction (e.g., the lower half of the vertical spatial frequency range represented in coefficients 120). In another example, a second subset 125b of coefficients 120 is for spatial frequencies within a second spatial frequency band, which includes high spatial frequencies in the horizontal direction (e.g., the upper half of the horizontal spatial frequency range represented in coefficients 120) and low spatial frequencies in the vertical direction. In yet another example, a third subset 125c of coefficients 120 is for spatial frequencies within a third spatial frequency band, which includes high spatial frequencies in the vertical direction (e.g., the upper half of the vertical spatial frequency range represented in coefficients 120) and low spatial frequencies in the horizontal direction. In another example, the fourth subset 125d of coefficient 120 is used for spatial frequencies within a fourth spatial frequency band, which includes high spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction.

[0032] To compress these coefficients, they can be rounded (e.g., rounded down). This allows for a reduction in the bit depth used to store the coefficient values. Additionally, coefficients rounded down to zero can be omitted to avoid being explicitly stored in the resulting coded image (e.g., by employing run-length encoding). To increase the level of compression, a set of scaling factors can be applied to scale the coefficients 120 before rounding down (or “quantizing”) the scaled coefficients. Thus, the scaling factor indicates the degree of scaling to be applied to one or more of the coefficients 120. A single scaling factor can be applied to all coefficients. Alternatively, scaling factors from a quantization table can be applied individually to the respective coefficients. Factors in such a quantization table can be specified based on information about human subjective visual perception to emphasize those coefficients found to be more “important” to human visual perception (e.g., lower spatial frequencies) by applying smaller scaling factors to these coefficients (thus preserving more information present in the coefficients by quantizing them according to a finer scale). Conversely, less important coefficients can be emphasized by applying a larger scaling factor (thus preserving less information in the coefficients by quantifying them according to a coarser scale and / or by increasing the likelihood of coefficients that are completely omitted by rounding to zero).

[0033] A single set of scaling factors (e.g., a single quantization table) can be applied to scale the coefficients of each block (e.g., 115) of image 100. Alternatively, a set of scaling factors can be determined and used to scale each block of image 100 individually. This can include selecting a set of scaling factors from several possible predetermined sets of scaling factors for each block. In another example, a block-level personalized scaling factor can be determined for each block and used to prescale the coefficients of the block before applying a common set of scaling factors, and / or to scale the common set of scaling factors before scaling the coefficients of the block using the common set of scaling factors. When image 100 is encoded, information indicating the scaling factors determined for each block can be included in the encoded image to facilitate decoding of the encoded image. This can include indicating the value of the block-level personalized scaling factor (e.g., in the range of 0 to 255), providing an index or other identification number of a scaling factor set selected from several different predetermined sets of scaling factors, or some other information indicating the set of scaling factors sufficient to determine which will be used to decode each block of the encoded image.

[0034] Scaling factors can be determined based on pixel blocks and / or the energy, entropy, or other determined properties of their determined coefficient blocks (e.g., identification of a selected set of predetermined scaling factors, or block-level personalized scaling factors). Additionally or alternatively, scaling factors can be determined based on information about the location of specific blocks within the image 100 as a whole. For example, scaling factors for blocks in image 100 can be determined based on their proximity to edges, faces, or some other feature of interest.

[0035] Figure 1C An example quantization table 130 is shown, including a scaling factor used for scaling factor 120. Such a quantization table can be pre-specified and / or determined on a block-by-block basis by software used to generate the encoded image (e.g., software running on a camera, software as part of an image processing suite). (e.g., selecting quantization table 130 from multiple possible quantization tables to determine a block-level personalized scaling factor to pre-scale a default quantization table to generate quantization table 130). A particular encoded image may include a copy of such a quantization table and / or information representing it (e.g., a block-level personalized scaling factor, an ID value of the quantization table selected individually for each block of the image) for decoding the encoded image. The decoder can then multiply the image content coefficients in the encoded image by the corresponding elements of quantization table 130 to "upscale" the quantization coefficients so that they can be transformed (e.g., via discrete cosine transform) into pixel values ​​(e.g., luminance values, chrominance values) of the decoded image.

[0036] Figure 1DThis diagram shows a set of quantized coefficients 140 that have been scaled to a certain degree according to quantization table 130 (i.e., according to the corresponding scaling factors within quantization table 130) and then quantized (rounded down). In this example, most of the coefficients 140 have zero values. As a result, the set of coefficients 140 can be efficiently stored using run-length encoding. The run-length encoded coefficients can then be further compressed (e.g., using lossless Huffman encoding) before being included in the final encoded image. The scaled, quantized coefficients 140 have been quantized using a standard set of quantization thresholds, for example, where scaling values ​​between -0.5 and 0.5 are quantized to "0", scaling values ​​between -1.5 and -0.5 are quantized to "-1", scaling values ​​between 0.5 and 1.5 are quantized to "1", etc. However, different quantization thresholds can be applied across all blocks of the image or different quantization thresholds can be determined individually for each block. For example, the width and location of the “dead zone” (i.e., the width and location of the range of values ​​to be quantized as “0”) can be adaptively selected to improve the compression ratio of the encoded version of image 100 (e.g., by widening the “dead zone” to result in more coefficients 140 being “0”) and / or improve the quality of the encoded version of image 100 (e.g., by narrowing the “dead zone” to result in fewer coefficients 140 being “0”).

[0037] Adaptive quantization based on residual in-band energy

[0038] exist Figure 1D In the example illustrated, a default quantization threshold of 0.5 is used to generate scaled quantized coefficients 140 from quantization table 130 and coefficients 120. Therefore, any of the coefficients 120 that have values ​​between -0.5 and 0.5 after being scaled according to the corresponding scaling factor of quantization table 130 is represented by the value 0 in the scaled quantized coefficients 140. The prevalence of such zero values ​​in the scaled quantized coefficients 140 allows for very high compression ratios when representing the scaled quantized coefficients 140 (e.g., by using run-length encoding).

[0039] The number of zeros (and the compression ratio of the resulting coded image) can be further increased by setting the quantization level used to determine such zero values ​​to a level higher than 0.5. This can be done on a per-block or per-block-per-band basis, or adaptively across blocks, bands, and / or other partitions within the image. For example, the quantization level can be adaptively determined on a per-block or per-block-per-band basis to increase the compression ratio, where it will have minimal impact on the quality of the resulting coded image (e.g., where the energy difference of the resulting coded image is less than a reduced threshold level). Conversely, the quantization level can be reduced for those image portions (e.g., blocks, bands within blocks) where the improvement in image quality is sufficient to justify a corresponding reduction in the compression ratio. This image coding method using adaptive quantization levels also has the benefit of being able to decode images encoded according to this method using existing downstream decoding software (e.g., without updating the decoding software).

[0040] This adaptive quantization coding method may include: for each part of the image to be encoded, evaluating multiple quantization levels, and then quantizing that part of the image using one of the evaluated quantization levels (e.g., the lowest evaluated quantization level that satisfies the image quality criteria). The quantized part is then used to generate a coded version of the input image. This part can be an entire block of the image, such that the process is repeated for each block. In another example, the part can be a subset of coefficients corresponding to a different frequency band within each block of the image. For example, a corresponding quantization level can be determined for each subset of coefficients 125a, 125b, 125c, and 125d of coefficients 120 in block 115. Performing adaptive quantization at such a sub-block level can provide the benefit of allowing the use of lower quantization values ​​to quantize frequency bands that typically represent less image energy (e.g., high-frequency bands like 215d). This allows additional non-zero elements to be “assigned” to those frequency bands by using lower quantization levels to quantize the coefficients corresponding to those bands, where the effect of such assignment is more likely to have a greater positive impact on the overall coded image quality.

[0041] Figure 2 The illustration shows multiple quantization levels 220 that can be evaluated by this adaptive quantization method. Each of the multiple quantization levels 220 can be evaluated to determine whether it meets image quality criteria when used for quantization scaling of the set of image coefficients. Figure 2 The illustration shows the values ​​of the example scaled image coefficient 210 above multiple quantization levels 220.

[0042] It can be from the maximum quantization level (e.g., such as Figure 2 The 0.7 level described in the text) to the minimum quantization level (e.g., as... Figure 2Multiple quantization levels (220) are iteratively evaluated (as depicted in the 0.3 level). The lowest value that satisfies the image quality criterion can then be selected and used to quantize the scaling factor 210 to be used in generating the coded version of the input image. Iteratively evaluating the quantization values ​​from the maximum to the minimum also allows for efficient use of computation time / effort. For example, evaluation can stop at the first quantization value that fails to satisfy the image quality criterion, and the coded image can then be generated using a previous, higher quantization value that satisfies the image quality criterion.

[0043] In the example, quantization level 220 can be evaluated iteratively in this manner until quantization level l1 no longer meets the image quality criteria. A previous quantization level can be used, resulting in only one of the scaling factors 210 (the rightmost factor) being represented by a non-zero value in the encoded version of the image. This represents a potential increase in compression ratio compared to using the default quantization level 0.5, which would result in three of the scaling factors 210 being represented by non-zero values ​​in the encoded version of the image.

[0044] In another example, quantization level 220 can be evaluated iteratively in this manner until level l2 fails to meet the image quality criterion. A previous quantization level can be used, resulting in the four scaling factors 210 (the four rightmost factors) being represented by non-zero values ​​in the encoded version of the image. This represents a potential reduction in compression ratio compared to using the default quantization level 0.5, which would result in three of the scaling factors 210 being represented by non-zero values ​​in the encoded version of the image. However, this could also represent an improvement in the quality of the encoded image, since the default quantization level 0.5 would result in quantization failing relative to the image quality criterion. Therefore, the adaptive quantization method described in this paper can result in non-zero elements being “assigned” where they can have the greatest benefit, thereby improving the quality of the encoded image and / or increasing the compression ratio of the encoded image.

[0045] like Figure 2 As shown, quantization level 220 spans a range from 0.3 to 0.7. These boundaries can provide benefits regarding compression ratio and image quality in at least some contexts. However, a set of quantization levels spanning other ranges can provide benefits. For example, the maximum quantization level can be a value greater than 0.675 or greater than 0.575. In some examples, the minimum quantization level can be a value less than 0.325. Additionally, Figure 2 The intervals and positions of the quantization levels shown are intended as non-limiting examples of quantization level values ​​that can be evaluated according to the methods described herein. Alternative numbers, intervals, ranges, or other properties of such multiple quantization values ​​may be used. The individual quantization values ​​may be spaced uniformly or non-uniformly across the value range.

[0046] Evaluating a specific quantization value relative to an image quality criterion can include various processes. In some examples, evaluating a specific quantization value may include quantizing a subset of scaling factors (e.g., a subset of scaling factors from a specific image patch for spatial frequencies within a specific spatial band) using the specific quantization value. The quantization value can then be evaluated relative to an image quality criterion. This evaluation may include (e.g., by summing the squares of the quantization values) determining the quantized energy of the quantization value. The determined quantized energy can then be compared to a threshold energy criterion. This threshold energy criterion can be relative, such as relative to the energy in the scaling factors before quantization. For example, the ratio between the quantized energy and the energy of the scaling factors before quantization, and the determined ratio, can be compared to a threshold. Such a threshold energy criterion can be absolute; for example, failure to meet a threshold energy criterion may include having a quantized energy of zero.

[0047] In some examples, none of the multiple quantization levels meets the image quality criteria. In such examples, a default quantization level (e.g., a level of 0.5, the minimum quantization level to be evaluated, or some other default level) can be used to quantize the relevant scaling factor of the input image. Additionally or alternatively, one or more of the image quality improvement methods described below can be applied.

[0048] Complementary adaptive coding based on residual in-band energy

[0049] As described above, various methods can be employed to adaptively set the quantization levels used for encoding the source image on a block-by-block and / or band-by-band basis. These methods can allow for reduced coded image size and / or improved coded image quality. Additionally or alternatively, other methods can be used to improve coded image quality in a block-selective and / or band-selective manner, thereby reducing the corresponding increase in coded image size. One or more of these supplementary methods can be applied when the aforementioned iterative evaluation methods fail, for example, when it is determined that none of the evaluated quantization levels satisfy the threshold energy criterion or some other image quality criterion. Alternatively, one or more of these supplementary methods can be applied in some other context. For example, a default quantization level (e.g., a quantization level of 0.5) can be applied, and if this default level fails to satisfy the threshold energy criterion or other image quality criterion, one or more of these supplementary methods can be applied.

[0050] The first supplementary method involves setting the quantization level to a sufficiently low value such that, when applied, this value causes at least one of the scaling factors to be quantized to a non-zero (e.g., unit) value. This can be achieved by identifying the largest coefficient within a subset of coefficients (e.g., used to specify spatial frequencies within a spatial band) and setting the quantization level to a value smaller than the identified scaling factor (e.g., a specified fraction of the identified scaling factor's value, or a specified amount smaller than the identified scaling factor's value). The set quantization level can then be applied to quantize the scaling factors. Alternatively, the method can be achieved by identifying the largest coefficient within the subset of coefficients and then setting the quantization value of the identified scaling factor to 1. This method has the benefit of ensuring that at least one coefficient in a particular block and / or band is represented by a non-zero value and therefore corresponds to an energy greater than 0 in the band.

[0051] Figure 3A The diagram shows a set of quantization coefficients 310a that have been scaled to the appropriate degree according to quantization table 130 and then quantized. A subset of the quantization coefficients 310a used for the spatial coefficients within the fourth frequency band 125d has been quantized according to the first supplementary method. Therefore, this set of quantization coefficients 310a contains coefficients with "1" values ​​corresponding to the scaling coefficients of the highest values ​​(i.e., the coefficients in the rightmost column of the four rows from the bottom).

[0052] The second supplementary method involves modifying or changing the set of scaling factors used to scale image coefficients before quantization. As described above, some image coding methods may include determining the scaling factor set before applying it to encode a portion of the image. This may include selecting a scaling factor set from multiple possible sets (e.g., selecting a quantization table from a possible set of quantization tables), determining a pre-scaling factor to scale the image coefficients before applying the scaling factor set, or performing other processes to determine the scaling factor set on a block-by-block and / or band-by-band basis. The scaling factor set may be determined based on the energy, entropy, or other determined properties of pixel blocks and / or coefficient blocks determined from them. Additionally or alternatively, scaling factors may be determined based on information about the location of a specific block within the source image as a whole and / or within the image. For example, scaling factors for blocks of the image may be determined based on the proximity of the blocks of the image to edges, faces, or some other feature of interest, such that the region of interest and / or features within the image are represented with higher fidelity in the encoded image.

[0053] In response to the determination that another method of image scaling and / or quantization fails to meet image quality criteria when scaling and quantizing blocks or other portions of the source image using the first set of scaling factors, a second supplementary method determines a second, different set of scaling factors and uses those scaling factors to scale and quantize blocks or other portions of the source image. The second set of scaling factors differs from the first set in that at least one scaling factor in the second set is of a lower magnitude than the corresponding scaling factor in the first set. Therefore, scaling the image coefficient set using the second set of scaling factors should allow as much or more energy from the coefficient set as possible to be represented in the output set of the scaled quantized coefficients compared to using the first set of scaling factors. This method can be applied iteratively. For example, if the scaled quantized coefficients generated using the second set of scaling factors also fail to meet image quality criteria, a third set of scaling factors can be determined and applied.

[0054] Figure 3B The diagram shows a set of quantization coefficients 310b that have been scaled to a corresponding degree according to a second set of scaling factors and then quantized. The second set of scaling factors in quantization table 130 differs from the first set of scaling factors in that at least one scaling factor in the second set is significantly smaller than its corresponding scaling factor in the first set. Therefore, this set of quantization coefficients 310b represents more energy and / or image content than the set of quantization coefficients 140 generated using the first set of scaling factors 130. Specifically, at least one non-zero coefficient exists for the spatial frequencies in the fourth spatial band 125d.

[0055] The third supplementary method involves representing at least one scaling factor in the coded version of the image using symbols representing subunit (e.g., fractional) values. This can be achieved by applying one or more additional quantization levels corresponding to one or more subunit symbols to the scaling factors. The one or more symbols used to represent these scaling factors can represent values ​​of 1 / 2, 1 / 4, 1 / 8, or some other subunit value. Such symbols can be represented in the coded image by reserved values ​​of a fixed bit width (e.g., an 8-bit value of 0-254 can represent a scaled and quantized factor with values ​​from 0 to 254, while a value of 255 can represent a value of 1 / 2), escape characters and / or associated information bits, or by some other indication. This approach has the advantage of ensuring that at least one factor in a particular block and / or band is represented by a non-zero value and therefore the energy in the corresponding band is greater than 0.

[0056] Figure 3CThe diagram shows a set of quantization coefficients 310c that have been scaled to the appropriate degree according to quantization table 130 and then quantized. A subset of the quantization coefficients 310c used for the spatial coefficients within the fourth frequency band 125d has been quantized according to a third additional method. Therefore, this set of quantization coefficients 310c contains coefficients corresponding to the highest value scaling factor (i.e., the coefficients in the rightmost column of the four rows from the bottom) and will be represented in the coded image by the symbol for a "1 / 2" value.

[0057] The fourth supplementary method involves adding blue noise (or other random or pseudo-random noise) to the image coefficients before quantization, or adjusting the amplitude of the blue noise that is added to the image coefficients by default before quantization. This can be done by increasing the amount or amplitude of the jitter applied to the image before quantization, thereby increasing the amount of energy in a specific spatial band within the image. For example, an image that would have no energy in a specific spatial band after scaling and quantization without adding and / or adjusting the amount of added blue noise in that spatial band can instead be represented by some non-zero energy in that specific spatial band after image encoding. In some examples, the added blue noise (or various other noises) can be band-limited to a specific spatial band that does not meet the energy criterion.

[0058] Example System

[0059] The computational functions described herein (e.g., transforming the pixels of an image into the frequency content of the image, scaling and quantizing such frequency content, evaluating and / or selecting quantization factors for image patches, or performing other image coding functions) can be performed by one or more computing systems. Such computing systems can be integrated into or take the form of computing devices such as mobile phones, tablet computers, laptop computers, servers, home automation components, stand-alone video capture and processing devices, cloud computing networks, and / or programmable logic controllers. For illustrative purposes, Figure 4 This is a simplified block diagram showing some of the components of the example computing device 400.

[0060] By way of example and not limitation, computing device 400 may be a cellular mobile phone (e.g., a smartphone), a camera, a computer (such as a desktop computer, laptop computer, tablet computer, or handheld computer), a personal digital assistant (PDA), a wearable computing device, a server, a cloud computing system (e.g., multiple networked servers or other computing units), or some other type of device or combination of devices. It should be understood that computing device 400 may represent a physical device, a specific physical hardware platform on which software is applied, or other combinations of hardware and software configured to perform mapping, training, and / or audio processing functions.

[0061] like Figure 4 As shown, computing device 400 may include communication interface 402, user interface 404, processor 406 and data storage 408, all of which can be communicatively linked together via system bus, network or other connection mechanism 410.

[0062] Communication interface 402 can be used to allow computing device 400 to communicate with other devices, access networks, and / or transmission networks using analog or digital modulation of electrical, magnetic, electromagnetic, optical, or other signals. Therefore, communication interface 402 can facilitate circuit-switched and / or packet-switched communications, such as Common Old Telephone Service (POTS) communications and / or Internet Protocol (IP) or other packet-switched communications. For example, communication interface 402 may include a chipset and antenna arranged for wireless communication with a radio access network or access point. Furthermore, communication interface 402 may take the form of a wired interface or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI) port. Communication interface 402 may also take the form of a wireless interface or include a wireless interface, such as Wi-Fi. Global Positioning System (GPS) or wide-area radio interface (e.g., WiMAX or 3GPP Long Term Evolution (LTE)). However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols can be used on communication interface 402. Furthermore, communication interface 402 may include multiple physical communication interfaces (e.g., Wi-Fi interface, etc.). Interface and wide area wireless interface).

[0063] In some embodiments, the communication interface 402 may be used to allow the computing device 400 to communicate with other devices, remote servers, access networks, and / or transmission networks. For example, the communication interface 402 may be used to receive a request from a requester device (e.g., a cellular phone, desktop, or laptop computer) for an image (e.g., an image on a website, an image stored in a user's online image hosting / storage account, or an image used as a thumbnail to indicate the content of a video related to other content requested by the requester device) to transmit an indication of an encoded image that has been encoded according to the methods described herein, or some other information. For example, the computing device 400 may be a server, a cloud computing system, or other system configured to perform the methods described herein, and the remote system may be a cellular phone, a digital camera, or another device configured to request information (e.g., a webpage that may have thumbnails or other images embedded therein) and receive from the computing device 400 one or more encoded images that may be modified as described herein (e.g., to improve image quality, increase the compression ratio of the image, and / or reduce the encoded size of the image) or other information from the computing device 400.

[0064] User interface 404 can be used to allow computing device 400 to interact with a user, such as receiving input from the user and / or providing output to the user. Therefore, user interface 404 may include input components such as a keypad, keyboard, touch-sensitive panel or presence-sensitive panel, computer mouse, trackball, joystick, microphone, etc. User interface 404 may also include one or more output components, such as a display screen that can be combined with a presence-sensitive panel. The display screen may be based on CRT, LCD, and / or LED technology, or other technologies now known or developed in the future. User interface 404 may also be configured to generate audible output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices.

[0065] In some embodiments, the user interface 404 may include a display for presenting video or other images to a user. Additionally, the user interface 404 may include one or more buttons, switches, knobs, and / or dials to facilitate the configuration and operation of the computing device 400. Some or all of these buttons, switches, knobs, and / or dials may be implemented as touch-sensitive panels or as functions on presence-sensitive panels.

[0066] Processor 406 may include one or more general-purpose processors (e.g., microprocessors) and / or one or more special-purpose processors (e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating-point units (FPUs), network processors, or application-specific integrated circuits (ASICs)). In some instances, the special-purpose processor may be capable of image processing and neural network computations, among other applications or functions. Data storage 408 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated wholly or partially with processor 406. Data storage 408 may include removable and / or non-removable components.

[0067] Processor 406 may be able to execute program instructions 418 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 408 to perform the various functions described herein. Therefore, data storage 408 may include a non-transitory computer-readable medium having program instructions stored thereon that, when executed by computing device 400, cause computing device 400 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings.

[0068] As an example, program instructions 418 may include an operating system 422 (e.g., an operating system kernel, device driver, and / or other modules) installed on computing device 400 and one or more application programs 420 (e.g., image encoding or other image processing programs for performing any of the methods described herein).

[0069] In some examples, depending on the application, portions of the method described herein may be executed by different devices. For example, different devices in a system may have varying amounts of computational resources (e.g., memory, processor cycles) and different information bandwidths for communication between devices. Different portions of the method described herein can be allocated based on these considerations.

[0070] Example Method

[0071] Figure 5 This is a flowchart of a method 500 for encoding an image. Method 500 includes generating a set of coefficients based on pixel blocks of an image, indicating image content at a plurality of spatial frequencies, wherein each coefficient in the set of coefficients is used for a corresponding spatial frequency among the plurality of spatial frequencies (510). Method 500 additionally includes scaling the set of coefficients according to a first set of scaling factors to generate a first scaled coefficient set (520).

[0072] Method 500 further includes performing an evaluation of multiple quantization levels (530). Performing the evaluation includes, for each of the multiple quantization levels: quantizing a subset of a first set of scaling factors according to the corresponding quantization level to generate a quantized subset of the first set of scaling factors, wherein the subset of the first set of scaling factors is used for spatial frequencies within a first spatial frequency band (532); determining the quantized energy of the quantized subset of the first set of scaling factors (534); and comparing the quantized energy with a threshold energy criterion (536).

[0073] Method 500 additionally includes generating a scaled and quantized version of the coefficient set based on evaluations at multiple quantization levels (540). Method 500 also includes generating an coded version of the image based on the scaled and quantized version of the coefficient set (550).

[0074] Figure 6 This is a flowchart of method 600. The method includes generating a set of coefficients based on pixel blocks of an image, indicating image content at a plurality of spatial frequencies, wherein each coefficient in the set of coefficients is used for a corresponding spatial frequency among the plurality of spatial frequencies (610). Method 600 additionally includes scaling a subset of the set of coefficients according to a first set of scaling factors to generate a first scaled subset of coefficients, wherein the subset of the set of coefficients is used for a corresponding spatial frequency within a first spatial frequency band (620).

[0075] Method 600 further includes quantizing a first subset of scaling factors to generate a first scaled set of quantized factors (630); determining the quantized energy of the first scaled subset of quantized factors (640); determining that the quantized energy does not meet a threshold energy criterion (650); and in response to determining that the quantized energy does not meet the threshold energy criterion, applying at least one process from a set of processes to generate an encoded version of the image (660).

[0076] The set of procedures includes a first procedure comprising: (a) identifying a maximum scaling factor within a first subset of scaling factors; (b) determining a quantization level smaller than the identified maximum scaling factor; and (c) quantizing the first subset of scaling factors using the determined quantization level. The set of procedures includes a second procedure comprising: (a) determining a second set of scaling factors, wherein at least one scaling factor in the second set of scaling factors is magnitude lower than a corresponding scaling factor in the first set of scaling factors; (b) scaling a subset of the set of scaling factors according to the second set of scaling factors to generate a second subset of scaling factors, wherein the subset of scaling factors indicates a corresponding spatial frequency within a first spatial frequency band; and (c) quantizing the second subset of scaling factors to generate a second scaled quantized set of scaling factors. The set of procedures includes a third procedure comprising, in a coded version of the image, representing at least one subset of the set of scaling factors indicating a corresponding spatial frequency within the first spatial frequency band using a symbol representing a subunit value.

[0077] Methods 500 and 600, or both, may include additional elements or features.

[0078] in conclusion

[0079] The above detailed description, with reference to the accompanying drawings, illustrates various features and functions of the disclosed systems, devices, and methods. In the drawings, similar symbols generally identify similar components unless the context otherwise indicates. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments and changes may be utilized without departing from the scope of the subject matter presented herein. It will be readily understood that aspects of this disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.

[0080] Regarding any or all message flowcharts, scenarios, and flowcharts in the accompanying drawings, and as discussed herein, each step, block, and / or communication may represent the processing and / or transmission of information according to exemplary embodiments. Alternative embodiments are included within the scope of these exemplary embodiments. In these alternative embodiments, for example, functions described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order, depending on the functions involved. Furthermore, more or fewer steps, blocks, and / or functions may be used with any of the message flowcharts, scenarios, and flowcharts discussed herein, and these message flowcharts, scenarios, and flowcharts may be combined with each other in part or in whole.

[0081] A step or block representing information processing may correspond to a circuit that can be configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a step or block representing information processing may correspond to a module, segment, or portion of program code (including associated data). Program code may include one or more processor-executable instructions for implementing a specific logical function or action in the method or technique. Program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk drive, hard disk drive, or other storage medium.

[0082] Computer-readable media may also include non-transitory computer-readable media, such as computer-readable media that store data for short periods, such as register memory, processor cache, and / or random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for longer periods, such as secondary or permanent long-term storage, such as read-only memory (ROM), optical discs or magnetic disks, and / or optical disc read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media can be considered, for example, computer-readable storage media or tangible storage devices.

[0083] Furthermore, a step or block representing one or more information transfers may correspond to information transfers between software and / or hardware modules in the same physical device. However, other information transfers may occur between software and / or hardware modules in different physical devices.

[0084] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and not intended to be limiting; the true scope is indicated by the appended claims.

Claims

1. A method for encoding an image, the method comprising: Based on the pixel blocks of the image, a set of coefficients is generated indicating the image content of the pixel blocks at multiple spatial frequencies, wherein each coefficient in the set of coefficients is generated for a corresponding spatial frequency among the multiple spatial frequencies; The set of coefficients is scaled according to a first set of scaling factors to generate a first set of scaling coefficients; Perform an evaluation across multiple quantization levels, wherein performing the evaluation includes, for each corresponding quantization level among the multiple quantization levels: A subset of the first scaling factor set is quantized according to the corresponding quantization level to generate a quantized subset of the first scaling factor set, wherein the subset of the first scaling factor set is quantized for spatial frequencies within a first spatial frequency band. as well as Determine the quantized energy of the quantized subset of the first scaling factor set; A scaled and quantized version of the coefficient set is generated based on the evaluation of the multiple quantization levels; as well as An encoded version of the image is generated based on a scaled and quantized version of the coefficient set.

2. The method according to claim 1, wherein, The evaluation of the multiple quantization levels is performed iteratively from the maximum quantization level to the minimum quantization level.

3. The method according to claim 2, wherein, The maximum quantization level is greater than 0.

675.

4. The method according to claim 2, wherein, The maximum quantization level is greater than 0.

575.

5. The method according to claim 2, wherein, The minimum quantization level is less than 0.

325.

6. The method according to claim 1, wherein, Performing an evaluation of multiple quantization levels also includes comparing the quantized energy with a threshold energy criterion.

7. The method according to claim 6, wherein, Comparing the quantized energy with the threshold energy criterion for a specific quantization level includes determining that the quantized energy satisfies the threshold energy criterion, and wherein generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels includes using the specific quantization level to quantize a subset of a first set of scaling coefficients for spatial frequencies within the first spatial frequency band.

8. The method according to claim 6, wherein, Performing an evaluation of the plurality of quantization levels includes determining that none of the plurality of quantization levels satisfies the threshold energy criterion, and wherein generating a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels includes responding to determining that none of the plurality of quantization levels satisfies the threshold energy criterion: Identify the largest scaling factor within a subset of a first set of scaling factors used for spatial frequencies within the first spatial frequency band; as well as Determine the quantization level for values ​​smaller than the maximum scaling factor identified; as well as The determined quantization level is used to quantize a subset of the first set of scaling factors for spatial frequencies within the first spatial band.

9. The method according to claim 6, further comprising: The first set of scaling factors is determined based on the pixel blocks of the image; The evaluation of the plurality of quantization levels includes determining that none of the plurality of quantization levels satisfies the threshold energy criterion, and the generation of a scaled and quantized version of the coefficient set based on the evaluation of the plurality of quantization levels includes responding to determining that none of the plurality of quantization levels satisfies the threshold energy criterion: Determine a second set of scaling factors, wherein at least one scaling factor in the second set of scaling factors is smaller in magnitude than the corresponding scaling factor in the first set of scaling factors; The evaluation of the multiple quantization levels is performed using the second set of scaling factors; and Based on the evaluation of the plurality of quantization levels using the second set of scaling factors, a scaled and quantized version of the coefficient set is generated.

10. The method according to claim 6, wherein, Performing an evaluation of the plurality of quantization levels includes determining that none of the plurality of quantization levels satisfies the threshold energy criterion, and wherein generating an coded version of the image based on a scaled and quantized version of the coefficient set includes, in response to determining that none of the plurality of quantization levels satisfies the threshold energy criterion, representing at least one coefficient in a subset of a first scaling coefficient set for spatial frequencies within the first spatial frequency band using a symbol representing a subunit value.

11. The method according to claim 10, wherein, The symbol for a subunit value represents one of the following: a half value, a quarter value, or an eighth value.

12. The method according to claim 8, wherein, Determining that none of the plurality of quantization levels satisfies the threshold energy criterion includes determining that each of the plurality of quantization levels results in a quantized energy of zero.

13. The method according to claim 1, further comprising: Determine the energy of a first subset of the coefficient set indicating image content, wherein the first subset of the coefficient set indicating image content indicates a corresponding spatial frequency within the first spatial frequency band, and The comparison of the quantized energy with the threshold energy criterion includes comparing the ratio between the quantized energy and the energy of the first subset of the coefficient set with the threshold energy.

14. The method according to any one of claims 1-13, wherein, The first spatial frequency band includes low spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction.

15. The method according to any one of claims 1-13, wherein, The first spatial frequency band includes low spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction.

16. The method according to any one of claims 1-13, wherein, The first spatial frequency band includes high spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction.

17. The method according to any one of claims 1-13, wherein, The first spatial frequency band includes high spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction.

18. The method according to any one of claims 1-13, wherein, The set of coefficients indicating the content of an image is the discrete cosine transform coefficient.

19. A method for encoding an image, the method comprising: A set of coefficients is generated based on the pixel blocks of the image to indicate the image content of transformed pixel blocks at multiple spatial frequencies, wherein each coefficient in the set of coefficients is generated for a corresponding spatial frequency among the multiple spatial frequencies; A subset of the coefficient set is scaled according to a first set of scaling factors to generate a first scaled coefficient subset, wherein the subset of the coefficient set is scaled for a corresponding spatial frequency within a first spatial frequency band; Quantize the first subset of scaling factors to generate a first set of quantized scaling factors; Determine the quantized energy of the first quantized subset of scaling factors; It is determined that the quantized energy does not meet the threshold energy criterion; as well as In response to determining that the quantized energy does not meet the threshold energy criterion, at least one process from a set of processes is applied to generate an encoded version of the image, wherein the set of processes includes: The first process includes: Identify the maximum scaling factor within the first subset of scaling factors; Determine the quantization level for values ​​smaller than the identified maximum scaling factor; and The first subset of scaling factors is quantized using the determined quantization level; The second process includes: Determine a second set of scaling factors, wherein at least one scaling factor in the second set of scaling factors is smaller in magnitude than the corresponding scaling factor in the first set of scaling factors; A subset of the coefficient set is scaled according to the second scaling factor set to generate a second scaling coefficient subset, wherein the subset of the coefficient set indicates a corresponding spatial frequency within the first spatial frequency band; and The second subset of scaling factors is quantized to generate a second set of quantized scaling factors; The third process, which includes: In the coded version of the image, a symbol representing a subunit value is used to represent at least one coefficient in a subset of the set of coefficients indicating the corresponding spatial frequency within the first spatial frequency band. The fourth process includes: Add blue noise to the first scaling factor subset before quantizing the first scaling factor subset; and The fifth process includes: The amount of blue noise added to the first scaling factor subset is adjusted before quantization.

20. The method according to claim 19, wherein, The symbol for a subunit value represents one of the following: a half value, a quarter value, or an eighth value.

21. The method according to claim 19, wherein, Determining that the quantized energy does not meet the threshold energy criterion includes determining that the quantized energy is zero.

22. The method according to any one of claims 19-21, wherein, The first spatial frequency band includes low spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction.

23. The method according to any one of claims 19-21, wherein, The first spatial frequency band includes low spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction.

24. The method according to any one of claims 19-21, wherein, The first spatial frequency band includes high spatial frequencies in the horizontal direction and low spatial frequencies in the vertical direction.

25. The method according to any one of claims 19-21, wherein, The first spatial frequency band includes high spatial frequencies in the horizontal direction and high spatial frequencies in the vertical direction.

26. The method according to any one of claims 19-21, wherein, The set of coefficients indicating the content of an image is the discrete cosine transform coefficient.

27. The method according to any one of claims 19-21, wherein, The blue noise added to the first scaling factor subset includes band-limited blue noise with spatial frequencies limited to the first spatial frequency band.

28. An article of manufacture comprising a non-transitory computer-readable medium having program instructions stored thereon, the program instructions causing the computing device to perform the method according to any one of claims 1-27 when executed by a computing device.

29. A system for encoding images, comprising: Controller; as well as A non-transitory computer-readable medium storing program instructions that, when executed by the controller, cause the controller to perform the method according to any one of claims 1-27.

Citation Information

Patent Citations

  • Video Data Encoding And Decoding

    CN103096075A

  • Optimisation of a quantisation matrix for image and video coding

    EP1675405A1