Estimation Method for Optimal Quantization Parameters in Image Compression

By constructing a multi-layer neural network to utilize image complexity feature vectors, the accuracy of image compression quantization parameters under finite bandwidth is solved, and the image transmission quality and bandwidth utilization are improved.

CN115665413BActive Publication Date: 2025-07-18WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211184629.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-18
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately estimate the quantization parameters of image compression under limited bandwidth, resulting in poor image transmission quality and unable to meet the balance between bandwidth utilization and image quality at the same time.

Method used

A multi-layer neural network is used to construct image complexity feature vectors, and through training and verification, a prediction method for optimal quantization parameters is established. Using features such as information entropy, absolute transformation difference and image gradient, a multi-layer neural network is constructed for estimation of quantization parameters.

Benefits of technology

It improves image transmission quality, adapts to the complexity of different image contents, makes full use of bandwidth, and achieves better image reconstruction effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115665413B_ABST
    Figure CN115665413B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for estimating optimal quantization parameters for image compression, comprising the following steps: Step 1, select an image set, and for each image I in the image set i , set the target bit rate B0 and calculate the optimal quantization parameter QP for each image; Step 2, calculate the complexity feature vector C of each image in the image set i ; Step 3, construct a multi-layer neural network NN. For this multi-layer neural network NN, the input layer is the complexity feature vector C of the image i , and the output layer is the regression-type quantization parameter QP i . Use the image set to train the multi-layer neural network NN; Step 4: For the image I′, calculate its complexity feature vector C′ and input it into the trained multi-layer neural network NN, so as to output the optimal estimate QP′ of the quantization parameter. The present invention can estimate the optimal image compression quantization parameters under limited bandwidth, make full use of the bandwidth, and improve the quality of image transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image compression, and particularly relates to a method for estimating optimal quantization parameters for image compression. Background Art

[0002] To achieve high-throughput image transmission over a channel with limited bandwidth, the original image data needs to be lossily compressed. The lossy compression result generated by the encoder must, on the one hand, ensure that the size is within the bandwidth limit to avoid information loss; on the other hand, it should utilize the bandwidth as much as possible to ensure better transmitted image quality. To meet these two requirements, various encoders are required to have reasonable rate control algorithms.

[0003] Existing image compression algorithms such as JPEG or JPEG2000 can generally be summarized into three steps: input data undergoes time-frequency transformation, quantization, and encoding. (1) Preprocessing, which generally refers to multi-component conversion of the color map or time-frequency conversion of the data in image compression; (2) Quantization, which is to perform precision downsampling on the original image data using the limitation of the human eye's resolution accuracy, thereby reducing the amount of data and facilitating subsequent data compression; (3) Encoding, that is, performing data compression encoding according to different image coding standards, and finally generating a compressed bitstream. Among them, the operation of precision downsampling of the quantization part of the data is the main reason for the degradation of image quality. The setting of the quantization parameter QP (Quantization Parameter, QP) will directly affect the data precision after quantization, and further affect the size of the compressed data stream and the image quality. Therefore, the reasonable setting of the quantization parameter is crucial for achieving rate control.

[0004] Currently, in the case of limited bandwidth, the rate control algorithm optimized for the quantization parameter QP in high-throughput image coding should be able to ensure that the use of the bandwidth does not exceed the limit range, does not lose data, and does not affect image reconstruction; and can also utilize the bandwidth as much as possible to achieve better image quality, presenting a better visual experience and a lower distortion rate. According to the rate-distortion theory, under the limited bandwidth, the higher the complexity of the image content, the larger the quantization parameter QP to ensure that the specified bitrate is not exceeded. Conversely, when the complexity of the image content is low, the quantization parameter QP can be smaller to fully utilize the bandwidth and improve the image restoration quality.

[0005] The calculation methods for the complexity of image content include the entropy of the image, the sum of absolute transform differences (SATD), and the gradient of the image. These features are representative in a certain aspect. Due to reasons such as image type and lighting conditions, there is still room for further improvement in the accuracy of quantization parameter estimation.

[0006] The information entropy of an image is originally a concept in statistics. In image processing, it can reflect the average amount of information in the image. The information entropy of an image is a measure of the uncertainty of the image content and can also reflect the complexity of the image content to a certain extent [1]. The calculation formula is shown in Equation (1). Among them, p represents the frequency of occurrence of a certain pixel value k in the entire image, as shown in Equation (2). When the probability p of each gray value appearing is equal, the information entropy of the image reaches the maximum value; when only one gray value appears in the image, that is, p is 1, the information entropy of the image takes the minimum value of 0. In an actual image, it is reflected that when the gray values of an image fluctuate in the entire value range, the information entropy of the image is relatively large, and conversely, when the gray values are relatively single, the information entropy is relatively small. Therefore, the information entropy of an image is often considered to be able to reflect the texture complexity of the image [1]. In many image or video coding environments, information entropy plays an important role in content coding [2].

[0007] Entropy = -∑p log2 p; (1)

[0008]

[0009] The sum of absolute transform differences (SATD) can reflect the magnitude of the residual signal of an image and video frames. It is often used as a measure of coding complexity in video coding and has been proven to have a strong linear relationship with the bitrate occupied by the compressed video. Therefore, it can participate in the quantization parameter QP decision in video coding [3][4]. SATD transforms the residual signal into the frequency domain and then sums the absolute values, as shown in Equation (3), where H represents the Hadamard matrix, M is the size of the square matrix, and X is the residual signal square matrix:

[0010] SATD = ∑ M ∑ M |HXH|; (3)

[0011] The gradient of an image can reflect the positional relationship of pixel values in the image [1]. Compared with the image, it contains more details of the spatial distribution and is equivalent to a two-dimensional feature parameter. The original meaning of the gradient is that the change rate of a certain function at a certain point is the largest and the change is the fastest. When this parameter of the gradient is introduced into image processing, it becomes a parameter that reflects the protrusions and edges in the image [5]. The calculation formula of the image gradient is shown in Equation (4), where I i,j is the pixel value of the image I at the position (i, j). The calculation result is jointly determined by the difference between horizontal pixels and the difference between vertical pixels and can reflect the structural information of the two-dimensional image:

[0012]

[0013] Due to the diversity of image content, relying on a single feature parameter of image complexity or a linear weighting method, the prediction of the optimal quantization parameter under bandwidth-limited conditions still fails in some cases. Therefore, it is necessary to develop an optimal quantization parameter estimation method with higher prediction accuracy.

[0014] References:

[0015] [1] Rafael C. Gonzalez, Richard E. Woods, Steven L. Eddins. Digital Image Processing [M]. Publishing House of Electronics Industry, 2005.

[0016] [2] Tu C, Tran T D. Context-based entropy coding of block transform coefficients for image compression [J]. IEEE Transactions on Image Processing, 2002, 11(11): 1271 - 1283.

[0017] [3] W. Gao, S. Kwong, Q. Jiang, C.-K. Fong, P. H. Wong, W. Y. Yuen, Data-driven rate control for rate-distortion optimization in HEVC based on simplified effective initial QP learning, IEEE Trans. Broadcast. 65(1)(2018)94–108.

[0018] [4] M. Karczewicz and X. Wang, “Intra Frame Rate Control Based on SATD, Document: JCTVC-M0257, 13th Meeting, Incheon, KR, Apr. 2013.

[0019] [5] Chen G H, Yang C L, Xie S L. Gradient-based structural similarity for image quality assessment [C] / / 2006 International Conference on Image Processing. IEEE, 2006: 2929 - 2932. Summary of the Invention

[0020] The object of the present invention is to provide an estimation method for the optimal quantization parameter of image compression in view of the deficiencies of the prior art. This method can estimate the optimal quantization parameter of image compression under limited bandwidth, make full use of the bandwidth, and improve the quality of image transmission.

[0021] To solve the above technical problems, the present invention adopts the following technical solutions:

[0022] An estimation method for the optimal quantization parameter of image compression includes the following steps:

[0023] Step 1: Select an image set, and for each image I in the image set i , set the target bit rate B0 and calculate the optimal quantization parameter QP of each image.

[0024] Step 2: Calculate the complexity feature vector C of each image in the image set i ;

[0025] Step 3: Construct a multi-layer neural network NN. For this multi-layer neural network NN, the input layer is the complexity feature vector C of the image i , and the output layer is the regression quantization parameter QP i , use the image set to train the multi-layer neural network NN, and verify and predict the trained multi-layer neural network NN, and select the multi-layer neural network NN with the highest test accuracy;

[0026] Step 4: For the image I′, calculate its complexity feature vector C′, and input it into the multi-layer neural network NN obtained in Step 3, so as to obtain the optimal estimated quantization parameter QP′ of the output.

[0027] Further, the specific method in Step 1 is:

[0028] For each image I i , adjust the quantization parameter QP of the image encoder; when the size S of the image compression output file is greater than B0, continuously reduce the quantization parameter QP until S does not exceed B0 and is closest to B0. At this time, the quantization parameter QP is the optimal quantization parameter QP corresponding to the target bit rate B0 of the image I i i .

[0029] Further, in Step 2, the complexity feature vector C of the image i is one or more of the information entropy Entropy of the image, the sum of absolute transform differences SATD, and the gradient of the image.

[0030] Further, the multi-layer neural network NN includes an input layer, two or more hidden layers, and an output layer. The activation function is ReLu, and the regression function is Linear.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows: By establishing image data with different content complexities, the present invention obtains the optimal quantization parameters for each picture through multiple encodings, calculates the image feature vectors, and constructs a data set of labeled feature vectors and optimal quantization parameters. Further, a multi-layer perceptron algorithm is trained using supervised learning to establish an optimal quantization parameter prediction method under limited bandwidth, which can adapt to the diversity of image content, give quantization parameters adapted to the image complexity, make full use of the bandwidth, and improve the quality of image transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of filtering the pictures in the image set in the embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of the optimal quantization parameter model structure of the multi-layer neural network NN in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0036] The present invention will be further described below with reference to specific embodiments, but it is not limited to the present invention.

[0037] The present invention provides a method for estimating the optimal quantization parameters of image compression, and the method includes the following steps:

[0038] Step 1, select an image set, and for each image I in the image set i , set the target bit rate B0 and calculate the optimal quantization parameter QP for each image;

[0039] Select an image set. In this embodiment, the image set includes the Inria aerial image data set and the Ground image data set. As an embodiment, a specific quantization parameter estimation method is described.

[0040] To increase the richness of image complexity, the dataset is augmented. In this embodiment, the original images in the dataset are subjected to mean filtering with windows of 3×3, 5×5, 9×9, and 11×11, as shown in Figure 1 , and all the filtered images and the original images together form an image set. The formula for mean filtering is:

[0041]

[0042] where, for windows of different sizes 3×3, 5×5, 9×9, and 11×11, N takes values of 3, 5, 9, and 11 respectively, and k takes values of 1, 2, 4, and 5 respectively.

[0043] After the image set is established, for each image I in the image set i , the optimal quantization parameter QP for each image is calculated by setting the target bitrate B0;

[0044] In this embodiment, assuming the target bitrate is B0 = 16 KB, the optimal quantization parameter QP for each image is calculated. For each image I in the above-established image set i , the quantization parameter QP of the image encoder is adjusted. When the size S of the image compression output file > 16 KB, the quantization parameter QP is continuously decreased at a set amplitude each time until S does not exceed 16 KB and is closest to 16 KB. At this time, the quantization parameter QP is the optimal quantization parameter QP corresponding to the target bitrate 16 KB for this image I i ; i .

[0045] Step 2, calculate the complexity feature vector C of each image in the image set i ;

[0046] In this step, the complexity feature vector C of the image i can select one or more parameters such as the information entropy (Entropy), sum of absolute transform differences (SATD), and gradient of the image. The calculation method is shown in formulas (1)-(4) in the background technology.

[0047] Step 3, construct a multi-layer neural network NN. For this multi-layer neural network NN, the input layer is the complexity feature vector C of the image i , and the output layer is the regression-type quantization parameter QP i ; The structure of the multi-layer neural network NN established in this embodiment is shown in Table 1, and the schematic diagram is as shown in Figure 2 ;

[0048] Table 1 Structure of multi-layer neural network NN

[0049]

[0050] The input layer of the multi-layer neural network NN is the image I i The image complexity feature vector C of i , and the output layer is the regression quantization parameter QP i , the activation function is ReLu, and the regression function is Linear. The image set is divided into a training set, a test set, and a validation set. The multi-layer neural network NN is trained through the training set to obtain its optimized parameters, and then the parameters of the trained multi-layer neural network NN are verified through the validation set and tested through the test set. During the test, two network structures of the multi-layer neural network NN are tested: structure (a) and structure (b). Among them, structure (a) contains 2 hidden layers with the number of nodes being 5 and 10 respectively; structure (b) contains 3 hidden layers with the number of nodes being 5, 10, and 5 respectively, and finally the optimal estimate QP is output.

[0051] To illustrate the beneficial effects of the present invention, the output structure of the multi-layer neural network NN is verified. In this embodiment, the bitrate accuracy BRA is defined as an evaluation of the rationality of the selection of the quantization parameter QP, and its calculation formula is shown in Equation (1). Among them, TBR represents the target bitrate, and represents the actual bitrate obtained after encoding. In the formula, when ABR and TBR are closer, BRA is closer to 1, indicating that the quantization parameter QP is set more accurately and the multi-layer neural network NN has better effects. Therefore, when testing the multi-layer neural network NN, the closer BRA is to 1, the better the multi-layer neural network NN. Table 2 shows the BRA test results of 7 different input multi-layer neural network NN models as shown in Table 1.

[0052]

[0053] Table 2 BRA test results of the multi-layer neural network NN model

[0054] (a) The number of nodes in each layer of the network is (n, 5, 10, 1)

[0055] (b) The number of nodes in each layer of the network is (n, 5, 10, 5, 1)

[0056] (Note: The bold data in the table are the optimal results)

[0057]

[0058] It can be seen from the table that when using 3 features as inputs at the same time, the bitrate accuracy reaches 85.79%. The multi-layer neural network NN with the highest test accuracy is selected as the final multi-layer neural network NN for prediction.

[0059] Step 4: For the image I′, calculate the image complexity feature vector C′, input it into the multi-layer neural network NN obtained in step 3, and output the optimal estimate QP′ of the quantization parameter.

[0060] Calculate the image complexity feature vector C′ of the image I′, and input it into two structures of the tested multi-layer neural network NN respectively to obtain the optimal estimated quantization parameter QP′ of the output.

[0061] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention accordingly. For those skilled in the art, it should be able to realize that the solutions obtained by equivalent substitution and obvious changes made by using the content of the specification of the present invention should all be included in the protection scope of the present invention.

Claims

1. A method for estimating an optimal quantization parameter for image compression, characterized in that It includes the following steps: Step 1, select an image set, and for each image I in the image set i , set the target bitrate B0 and calculate the optimal quantization parameter QP for each image; Step 2, calculate the complexity feature vector C of each image in the image set i ; Step 3, construct a multi-layer neural network NN. For this multi-layer neural network NN, the input layer is the image complexity feature vector C i , and the output layer is the regression quantization parameter QP i , use the image set to train the multi-layer neural network NN, verify and predict the trained multi-layer neural network NN, and select the multi-layer neural network NN with the highest test accuracy; Step 4: For the image I′, calculate its image complexity feature vector C′, and input it into the multi-layer neural network NN obtained in Step 3, so as to obtain the optimal estimated quantization parameter QP′ of the output.

2. The estimation method of the optimal quantization parameter for image compression according to claim 1, characterized in that, The specific method in Step 1 is as follows: For each image I i , adjust the quantization parameter QP of the image encoder; when the size S of the image compression output file > B0, continuously decrease the quantization parameter QP until S does not exceed B0 and is closest to B0. At this time, the quantization parameter QP is the optimal quantization parameter QP corresponding to the target bit rate B0 of this image I i i .​ 3. The estimation method of the optimal quantization parameter for image compression according to claim 1, characterized in that In step 2, the image complexity feature vector C i is one or more of the information entropy Entropy of the image, the sum of absolute transform differences SATD, and the gradient of the image.

4. The estimation method of the optimal quantization parameter for image compression according to claim 1, characterized in that The multi-layer neural network NN includes an input layer, two or more hidden layers, and an output layer. The activation function is ReLu, and the regression function is Linear.

Citation Information

Patent Citations

  • Image compression method and image compression system based on ARM multi-core heterogeneous processor

    CN113242433A

  • Image compression method based on deep self-attention transformation network

    CN114494472A