Perceptual image compression model evaluation method and system based on multi-index fusion

Through the perceptual image compression model evaluation method of multi-index fusion, a comprehensive score is generated using a neural network adaptive fusion module, which solves the problem of inconsistent image compression model evaluation in the existing technology and achieves more accurate subjective perceptual evaluation.

CN120602655APending Publication Date: 2025-09-05PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510777207.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing image compression methods are difficult to fully reflect the subjective perception of the human eye. The evaluation results of a single perception indicator in different application scenarios are inconsistent, resulting in the inconsistency between objective evaluation and human visual experience.

Method used

A multi-index fusion method is adopted. Through a three-layer fully connected neural network adaptive fusion module, GELU and Sigmoid functions are used to generate weights. Combined with the mean square error loss function optimization, a comprehensive score is generated to achieve the unification and weighted summation of multiple perceptual image quality evaluation indicators.

Benefits of technology

The subjective perception relevance of the image compression model has been improved, and the generated comprehensive scoring results are more comprehensive and accurate, which solves the problem of inconsistent ranking of traditional indicators and improves the scientificity and consistency of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602655A_ABST
    Figure CN120602655A_ABST
Patent Text Reader

Abstract

The invention discloses a perception image compression model evaluation method based on multi-index fusion, and belongs to the technical field of digital image compression. In order to solve the problem that an existing image quality evaluation method is difficult to comprehensively reflect subjective perception of human eyes, after multiple evaluation indexes are subjected to normalization processing, adaptive fusion is carried out through a three-layer full-connection neural network, and a comprehensive score is output based on a weighted summation mode. According to the method, the image compression quality can be measured more comprehensively and accurately, and the consistency of the model and human eye perception is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image compression, and in particular to a perceptual image compression model evaluation method and system based on adaptive fusion of multiple quality evaluation indicators. Background Art

[0002] The primary goal of image compression technology is to reduce the amount of image data, thereby lowering storage space requirements and transmission bandwidth, while preserving the original image quality as much as possible. Traditional image compression methods often rely on predefined encoding rules and manually designed modules to achieve compression by removing redundant information. These methods focus on mathematical fidelity calculations. However, these traditional metrics cannot fully reflect the human eye's subjective perception of image quality. Therefore, in practical applications, there may be a discrepancy between objective metrics and human visual experience.

[0003] In recent years, with the rapid development of deep learning technology, image quality assessment metrics tailored to human vision have become a research hotspot. These metrics not only consider mathematical errors in images but also emphasize the perceptual consistency of image content and structure, thus more closely resonating with human subjective evaluations. Examples include the Perceptual Patch Image Similarity Score (LPIPS), the Distance and Texture Similarity Score (DISTS), and the Perceptual Image Error Score (PieAPP).

[0004] While each of these perceptual metrics has its own distinct focus and demonstrates good subjective quality evaluation performance in various application scenarios, in practice, performance evaluation results for the same set of perceptual image compression models often differ due to differences in their calculation methods and focus. This inconsistency between metrics poses a challenge to objectively assessing image reconstruction quality and makes it difficult for a single metric to comprehensively and objectively reflect the compression model's overall performance in balancing fidelity and realism. Summary of the Invention

[0005] To address the problem that existing image quality evaluation methods are difficult to fully reflect the subjective perception of the human eye, the present invention provides a perceptual image compression model evaluation method and system based on the adaptive fusion of multiple quality evaluation indicators, aiming to achieve a more comprehensive and accurate measurement of image quality by adaptively combining the advantages of multiple indicators.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A perceptual image compression model evaluation method based on multi-index fusion includes the following steps:

[0008] 1) Collect the original image and the corresponding compressed reconstructed image obtained by the perceptual image compression model, calculate multiple image quality evaluation indicators for each pair of images, and perform normalization;

[0009] 2) The normalized evaluation indicators are used as input to an adaptive fusion module consisting of a three-layer fully connected neural network. The module includes an input layer, a hidden layer, and an output layer. The GELU activation function is used between the hidden layers, and the output layer generates the weights corresponding to each evaluation indicator.

[0010] 3) Perform weighted summation based on the normalized value and weight of each evaluation indicator to obtain the comprehensive score of the perceptual image compression model;

[0011] Among them, the mean square error loss function is used to train the adaptive fusion module to minimize the error between the weighted summation comprehensive score and the subjective score.

[0012] Furthermore, the image quality evaluation index in step 1) includes at least two of the following: Perceptual Image Block Similarity (LPIPS), Structure and Texture Similarity (DISTS), and Perceptual Image Error Score (PieAPP).

[0013] Furthermore, the mapping function of the normalization process in step 1) is a linear normalization function.

[0014] Furthermore, in step 1), for the evaluation index with low value priority, the result of 1 minus the normalized value is used to represent the final value of the evaluation index.

[0015] Furthermore, the hidden layer of the adaptive fusion module in step 2) includes two fully connected layers, and the GELU activation function is used between the first fully connected layer and the second fully connected layer; the output layer outputs the corresponding weights of each evaluation indicator through the Sigmoid function.

[0016] Furthermore, in step 2), for the evaluation indicators with low priority, their corresponding weights are determined by subtracting the output value of the neural network from 1.

[0017] Furthermore, during the training of the adaptive fusion module, the AdamW optimizer is used to optimize the model parameters.

[0018] A perceptual image compression model evaluation system based on multi-index fusion, including:

[0019] An input module, configured to receive an original image and a corresponding compressed reconstructed image;

[0020] An evaluation module, configured to evaluate the quality of the image pair based on multiple image quality evaluation indices and normalize the results of each evaluation indices;

[0021] The adaptive fusion module receives the normalized evaluation index results and inputs them into a three-layer fully connected neural network, which includes a hidden layer using the GELU activation function and an output layer with output weights. The network is used to calculate the adaptive weights of each indicator and perform weighted summation to obtain the comprehensive score of the perceptual image compression model.

[0022] The output module is used to output the scores of each individual indicator and the weighted comprehensive score results.

[0023] The beneficial effects achieved by the present invention are as follows:

[0024] 1. The present invention normalizes multiple image quality evaluation indicators, unifies the numerical dimensions of each indicator, makes different indicators comparable, and solves the problem of inconsistent ranking standards among traditional multiple indicators.

[0025] 2. The present invention constructs an adaptive fusion module based on a three-layer fully connected neural network, mines the nonlinear relationship between indicators through the GELU activation function, and uses the Sigmoid function to generate weights, thereby realizing adaptive adjustment of the importance of each indicator in different image scenarios.

[0026] 3. The present invention generates a comprehensive score through weighted summation. During the training process, the mean square error loss function is used to optimize the fusion network, so that the output results are highly consistent with the subjective scores of the human eye, thereby improving the subjective perception relevance of the evaluation results.

[0027] 4. The test results of the present invention on the standard PIPAL dataset show that its consistency indicators such as SROCC, PLCC, and KROCC are all better than the existing mainstream single indicators, verifying the advantages of the method in terms of stability and evaluation accuracy.

[0028] 5. The system design proposed in this invention integrates indicator calculation, adaptive scoring and result display functions, builds an integrated evaluation platform, supports the coordinated operation of input and output modules and training modules, and facilitates practical deployment and promotion and application.

[0029] 6. This invention provides an objective, accurate and comprehensive quality evaluation tool for perceptual image compression models, which has good applicability and scalability, and can provide theoretical support and technical basis for subsequent image compression algorithm optimization and the formulation of image quality evaluation standards. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is an overall flow chart of the perceptual image compression model evaluation method based on multi-index fusion of the present invention. DETAILED DESCRIPTION

[0031] In order to make the various technical features and advantages or technical effects of the above technical solutions of the present invention more obvious and easy to understand, the following embodiments are described in detail with reference to the accompanying drawings.

[0032] This embodiment provides a perceptual image compression model evaluation method based on multi-index fusion. The overall process is as follows: Figure 1 The basic principle of this method is: first, normalize multiple image quality evaluation indicators, then use a three-layer fully connected neural network to achieve adaptive fusion between indicators, and finally output a weighted summation comprehensive score to measure the perceived quality of the compressed image. The specific processing steps of this method are as follows:

[0033] Step 1: Data collection and preprocessing

[0034] The original image is collected and processed through the perceptual image compression model to obtain the corresponding compressed reconstructed image. The data set composed of these two images is used as the basic data for model evaluation. For each pair of original image and compressed image, multiple perceptual quality evaluation indicators are calculated separately, including LPIPS, DISTS, PieAPP, etc. In order to eliminate the differences in the numerical range of each indicator, the normalization method is used to map the value of each indicator to the interval [0,1]. Among them, for the indicator of "the smaller the value, the better the quality", the "1-normalization value" method is used for reverse processing to ensure that all indicator values ​​are "the larger the value, the better the quality". The normalization process can be expressed as:

[0035]

[0036] Among them, m i represents the original value of the i-th evaluation index, f norm (·) represents the normalized mapping function.

[0037] Step 2: Adaptive fusion module construction

[0038] Construct an adaptive fusion module consisting of a three-layer fully connected neural network to perform nonlinear fusion of multiple normalized indicators. The structure of this module is as follows:

[0039] 1) Input layer: Receives the normalized evaluation index and forms an input vector:

[0040]

[0041] 2) Hidden layer: There are two fully connected layers, and the GELU activation function is used between the two layers to fully explore the nonlinear relationship between various indicators.

[0042] 3) Output layer: After the full connection transformation, the output layer uses the Sigmoid function to map the network output to the [0, 1] interval, which serves as the fusion weight of each evaluation indicator:

[0043] W=Sigmpod(FC2(GELU(FC1(X))))

[0044] Among them, FC1 and FC2 are the first and second fully connected layers respectively, W=[w1,w2,…,w n ] is the weight value corresponding to the indicator. For the evaluation indicator with low value priority, its corresponding final weight is calculated by “1-w i ” to make adjustments.

[0045] Step 3: Model training and loss function setting

[0046] According to the normalized value of each evaluation index And the corresponding weight w i Perform weighted summation to obtain the comprehensive score Y of the perceptual image compression model:

[0047]

[0048] The mean square error (MSE) loss function is used as the loss function:

[0049]

[0050] Among them, N is the total number of samples, Y i is the comprehensive score output by the weighted fusion module of the i-th sample, is the corresponding subjective rating label.

[0051] The adaptive fusion module is trained using the MSE loss function to optimize the parameters of the adaptive fusion module. The goal is to make the comprehensive score Y and the subjective score label Y target The error between them is the smallest. The expression for minimizing the loss function is as follows:

[0052]

[0053] During the training process, the AdamW optimizer was used with the following parameters: the learning rate was 1×10 -4 , the weight decay coefficient is 0.01, and the batch size is 32.

[0054] In this embodiment, the training data uses the preset training set in the PIPAL dataset. During the training process, the network weights are continuously updated to minimize the MSE loss, and the PLCC, SRCC, and KROCC indicators are calculated on the validation set to evaluate the degree of improvement in the consistency between the model performance and human eye perception.

[0055] In addition, this embodiment also provides a perceptual image compression model evaluation system based on multi-index fusion, which is used to perform the above method steps, so that users can comprehensively evaluate the performance of perceptual image compression models in actual application scenarios. The system includes the following functional modules:

[0056] 1) Input module: used to receive the image data to be evaluated input by the user, including the compressed reconstructed image folder path and its corresponding original image folder path, and use it as the input of the subsequent evaluation process.

[0057] 2) Evaluation module: This module is responsible for calling the preset evaluation index calculation program and calculating multiple image quality evaluation indicators in sequence based on the normalized processing flow, including but not limited to LPIPS, DISTS, PieAPP, etc., a total of ten evaluation indicators, and normalizing the results of each indicator to a unified dimension.

[0058] 3) Adaptive fusion module: Based on a pre-trained three-layer fully connected neural network model, the above normalized indicators are weightedly fused to generate a final comprehensive score result, which is used to reflect the overall perceptual quality of the perceptual image compression model.

[0059] 4) Output module: Outputs each individual evaluation indicator and the fusion score results in a visual manner, including tabular and graphical presentations, to support users in intuitively comparing and analyzing the performance of different image compression methods in multiple dimensions.

[0060] The above system can be integrated and deployed in the image quality evaluation software platform. It has good scalability and practicality and is suitable for various scenarios such as scientific research experiments and industrial applications.

[0061] Performance testing:

[0062] The proposed method was tested on the standard image quality assessment dataset PIPAL. The experimental results show that the Spearman rank correlation coefficient (SROCC), Pearson linear correlation coefficient (PLCC), and Kendall rank correlation coefficient (KROCC) of the proposed method on this dataset reached 0.7642, 0.7839, and 0.5705, respectively, significantly outperforming the performance of existing single evaluation metrics. For example, when using the PieAPP evaluation metric, the SROCC, PLCC, and KROCC were only 0.7078, 0.6956, and 0.5140, respectively.

[0063] The test results demonstrate that this method effectively addresses the discrepancies in ranking consistency found in traditional metrics. By integrating the strengths of multiple metrics, it significantly improves consistency with subjective human perception. The resulting comprehensive and objective scoring results are conducive to improving the scientific and accurate evaluation of perceptual image compression models, and provide reliable support for the development of standards for image quality assessment and the promotion and application of related technologies.

[0064] It should be noted that the number of layers, number of modules, activation function types, and specific settings of each layer involved in the above embodiments are only preferred embodiments of the present invention, which are used to illustrate the technical principles of the present invention and do not constitute a limitation of the present invention. Those skilled in the art may make equivalent replacements or functional adjustments to specific structures and parameters without departing from the spirit of the present invention, and they should still fall within the scope of protection of the present invention. The scope of protection of the present invention shall be based on the content defined in the claims.

Claims

1. A perceptual image compression model evaluation method based on multi-index fusion, comprising the following steps: 1) Collect the original image and the corresponding compressed reconstructed image obtained by the perceptual image compression model, calculate multiple image quality evaluation indicators for each pair of images, and perform normalization; 2) The normalized evaluation indicators are used as input to an adaptive fusion module consisting of a three-layer fully connected neural network. The module includes an input layer, a hidden layer, and an output layer. The GELU activation function is used between the hidden layers, and the output layer generates the weights corresponding to each evaluation indicator. 3) Perform weighted summation based on the normalized value and weight of each evaluation indicator to obtain the comprehensive score of the perceptual image compression model; Among them, the mean square error loss function is used to train the adaptive fusion module to minimize the error between the weighted summation comprehensive score and the subjective score.

2. The method according to claim 1, wherein The image quality evaluation index in step 1) includes at least two of the following: perceived image block similarity, structure and texture similarity, and perceived image error score.

3. The method according to claim 1, wherein The mapping function of the normalization process in step 1) is a linear normalization function.

4. The method according to claim 1, wherein In step 1), for the evaluation index with low value priority, the result of 1 minus the normalized value is used to represent the final value of the evaluation index.

5. The method according to claim 1, wherein The hidden layer of the adaptive fusion module in step 2) includes two fully connected layers, and the GELU activation function is used between the first fully connected layer and the second fully connected layer; the output layer outputs the corresponding weights of each evaluation indicator through the Sigmoid function.

6. The method according to claim 1, wherein In step 2), for the evaluation indicators with low priority, their corresponding weights are determined by subtracting the neural network output value from 1.

7. The method according to claim 1, wherein During the training of the adaptive fusion module, the AdamW optimizer is used to optimize the model parameters.

8. A perceptual image compression model evaluation system based on multi-index fusion, used to execute the method according to any one of claims 1 to 7, characterized in that: include: An input module, configured to receive an original image and a corresponding compressed reconstructed image; An evaluation module, configured to evaluate the quality of the image pair based on multiple image quality evaluation indices and normalize the results of each evaluation indices; The adaptive fusion module receives the normalized evaluation index results and inputs them into a three-layer fully connected neural network, which includes a hidden layer using the GELU activation function and an output layer with output weights. The network is used to calculate the adaptive weights of each indicator and perform weighted summation to obtain the comprehensive score of the perceptual image compression model. The output module is used to output the scores of each individual indicator and the weighted comprehensive score results.