Image quality evaluation method for field generalization, program, equipment and storage medium
The domain-generalized image quality evaluation method addresses the issue of performance degradation due to data distribution discrepancies by using a shared feature extractor and domain discriminator to integrate domain-specific knowledge, enhancing model generalization.
Patent Information
- Application Number
- CN202510351050.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-24
AI Technical Summary
When evaluating image quality in deep learning methods, the large difference in the distribution of training data and test data leads to performance degradation. Simple data set mixing cannot effectively retain the domain-specific quality discriminant knowledge of each data set, which affects the generalization of the model.
The image quality evaluation model with shared feature extractor and multi-regressor structure is adopted to train the model through quality-sensitive triple loss, quality prediction loss and monotonic loss, and the domain discriminator is used to weight the fusion regressor prediction results during testing to build domain generalization capabilities.
The generalization performance of the image quality evaluation model is improved, the quality discrimination knowledge of each data set is effectively integrated, and the prediction accuracy of the model on unknown target data sets is improved.
Smart Images

Figure CN120318659A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image quality evaluation, and particularly relates to a domain generalization-based image quality evaluation method, program, device, and storage medium. Background Art
[0002] Deep learning methods have shown superior performance in the field of image quality evaluation. However, when there is a large distribution difference between the training data and the test data, the performance of deep learning-based image quality evaluation methods will significantly decline. Therefore, improving the generalization ability of deep learning models is of great significance for enhancing the practicality of deep learning-based image quality evaluation methods.
[0003] Mixing multiple image quality evaluation data sets to augment the data distribution is an effective means to improve the generalization of deep learning-based image quality evaluation models. However, due to different image distortion types, image contents, and different emotional tendencies of annotators, different image quality evaluation data sets often have large distribution differences. Simple data set mixing may reduce the performance of the model due to large data distribution differences and cannot effectively retain the quality discrimination knowledge unique to different data set domains. Summary of the Invention
[0004] The purpose of the present invention is to provide a domain generalization-based image quality evaluation method, program, device, and storage medium.
[0005] A domain generalization-based image quality evaluation method includes the following steps:
[0006] Step 1: Obtain M data sets with image quality score labels, set domain labels for each data set, and construct the M data sets with domain labels into a training set;
[0007] Step 2: Train a domain generalization-based image quality evaluation model using the training set;
[0008] The domain generalization-based image quality evaluation model includes a shared feature extractor, a regressor, and a domain discriminator; the number of regressors is M, corresponding to the M data sets;
[0009] When training the domain generalization-based image quality evaluation model, calculate the quality-sensitive triplet loss, quality prediction loss, monotonicity loss, and domain discrimination loss, and construct a global loss through weighting;
[0010] Step 3: Input the image data sample to be evaluated into the trained domain-generalized image quality evaluation model. After the shared feature extractor extracts the image data quality features, input the image data quality features into the domain discriminator and M regressors respectively. Use the result output by the domain discriminator as the weight to perform weighted fusion on the quality scores output by the M regressors, obtain the quality score of the image data sample to be evaluated, and complete the evaluation of the image quality.
[0011] Further, there are M data sets in the training set D in step 1, D = {D1, D2,..., D k ,..., D M}, and there are N k samples in the k-th data set D k . And are the image data and quality score of the i-th sample in the k-th data set respectively, and d k is the domain label of the k-th data set.
[0012] Further, in step 2, the image data k of the i-th sample in the training stage data set D is input into the shared feature extractor, and the shared feature extractor extracts the image data quality feature and the regressor R k corresponding to the data set D k maps the image data quality feature to the quality score The domain discriminator D c predicts the domain label of the sample according to the image data quality feature and obtains the probability set of each domain label corresponding to the sample represents the probability that the predicted domain label of the i-th sample is d k .
[0013] Further, the calculation method of the quality-sensitive triplet loss in step 2 is as follows:
[0014] Input M data sets into the domain-generalized image quality evaluation model. For each sample use its own image data as the anchor image, and the superscript "*" indicates not distinguishing the data set where the sample comes from; randomly select samples from the M data sets and satisfy ε is a preset quality difference threshold, and use as the positive sample of the anchor image; randomly select samples from the M data sets and satisfy Will As negative samples of anchor images;
[0015] In order to bring the distance between the anchor image and the positive sample closer and push the distance between the anchor image and the negative sample farther in the feature space, the loss of the quality-sensitive triplet is calculated as:
[0016]
[0017] Wherein, N is the total number of samples in the M data sets; δ is the minimum value of the distance difference between the preset positive and negative samples and the anchor image in the feature space.
[0018] Furthermore, the calculation method of the quality prediction loss in step 2 is:
[0019]
[0020] in, represents the root mean square error;
[0021] The monotonicity loss is calculated as follows:
[0022]
[0023] in, Denotes the dataset D k The corresponding regressor R k The monotonicity loss of Err (i,j),k For the dataset D k middle and The monotonicity error,
[0024] The calculation method of the domain discrimination loss is:
[0025]
[0026] in, represents the cross entropy loss;
[0027] The quality-sensitive triplet loss, quality prediction loss, monotonicity loss, and domain discrimination loss are weighted to construct the global loss L:
[0028] L=αL pre +βL mono +γL triplet +λL dom
[0029] Among them, α, β, γ, and λ are weight hyperparameters for balancing various losses.
[0030] Further, in step 3, the image data sample x to be evaluated t is input into the trained domain generalization image quality evaluation model. After the shared feature extractor extracts the image data quality feature G(x t ), the image data quality feature G(x t ) is respectively input into the domain discriminator and M regressors. The result output by the domain discriminator is used as the weight to weight and fuse the quality scores output by the M regressors, so as to obtain the quality score of the image data sample to be evaluated
[0031]
[0032] Further, after constructing the training set in step 1, data augmentation is performed on the image data of the training data set. The augmentation means include random rotation, random cropping, random mirroring, and scaling. In step 2, the shared feature extractor is composed of the feature extraction layer of ResNet18 pre-trained on ImageNet stacked with two fully connected layers with ReLu as the activation function, denoted as FC-ReLu-FC-ReLu. The structure of the regressor is two linear layers, denoted as FC-FC. The structure of the domain discriminator is denoted as FC-ReLu-FC-ReLu-FC-SoftMax.
[0033] A computer device / system includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above-mentioned domain generalization image quality evaluation method.
[0034] A computer-readable storage medium stores a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned domain generalization image quality evaluation method are implemented.
[0035] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned domain generalization image quality evaluation method are implemented.
[0036] The beneficial effects of the present invention are as follows:
[0037] The present invention learns the image quality features of multiple datasets through a shared feature extractor and constructs a quality-sensitive feature space through a quality-sensitive triplet loss. The unique multi-regressor structure effectively preserves the quality discrimination knowledge unique to each domain. In the test phase, the present invention assigns a higher weight to the predicted value of the regressor corresponding to the training dataset most similar to the unknown target dataset based on the discrimination result of the domain discriminator, effectively integrating the prediction results of each regressor and improving the generalization performance. The present invention effectively solves the problem that the domain generalization image quality evaluation method with multi-dataset mixing cannot effectively integrate the quality discrimination knowledge of each dataset, and improves the generalization of the deep learning-based image quality evaluation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of a domain generalization image quality evaluation method in the present invention.
[0039] Figure 2 It is a schematic diagram of the domain generalization image quality evaluation method provided by an embodiment of the present invention.
[0040] Figure 3 It is a schematic diagram of a quality-sensitive triplet provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The present invention will be further described below with reference to the accompanying drawings.
[0042] The present invention designs a domain generalization image quality evaluation model with multiple regressor branches sharing a feature extractor, and effectively learns the quality discrimination knowledge of the mixed dataset through a quality-sensitive triplet loss, as well as the quality prediction loss and monotonicity loss of each regressor branch. The domain discriminator trained by the domain discrimination loss will assign a higher weight to the predicted result of the regressor corresponding to the most similar training dataset when facing an unknown target dataset, effectively integrating the prediction results of each regressor and improving the generalization performance of the model.
[0043] First, obtain M datasets with image quality score labels, set domain labels for each dataset, and construct the M datasets with domain labels as the training set;
[0044] Mix M image quality evaluation datasets with known quality scores to construct the training dataset D = {D1, D2,..., D k ,..., D M} of the domain generalization image quality evaluation method, where the k-th dataset is and are respectively the image data and quality score of the i-th sample of the k-th dataset, and d k is the domain label of the k-th dataset, and Nk is the number of samples in the k-th dataset D k in the dataset.
[0045] Before training, data augmentation can be performed on the image data of the training dataset, and the augmentation methods include but are not limited to operations such as random rotation, random cropping, random mirroring, and scaling.
[0046] In one embodiment of the present invention, the image data is scaled proportionally based on the short side, the short side is scaled to 512 pixels, the long side is scaled according to the original image ratio, and an image block of 512×512 pixels is randomly cropped, and the training data for the domain generalization model is formed by random horizontal mirroring. A total of 4 datasets are mixed in this embodiment.
[0047] Build a domain generalization image quality evaluation model, which consists of a shared feature extractor, M regressors, and a domain discriminator; the M regressors correspond to the M datasets one by one;
[0048] The shared feature extractor is used to learn the quality discrimination features contained in each dataset. The regressors corresponding to each dataset that make up the training dataset are used to retain the quality discrimination knowledge unique to each dataset. The domain discriminator is used to weighted aggregate the prediction results of the regressors corresponding to each dataset according to the domain discrimination results when testing on an unknown target dataset, improving the generalization of the model.
[0049] In one embodiment of the present invention, the feature extractor uses the feature extraction layer of ResNet18 pre-trained on ImageNet and stacks two fully connected layers with ReLu as the activation function (FC-ReLu-FC-ReLu). The structure of each regressor branch is two linear layers (FC-FC), and the structure of the domain discriminator is (FC-ReLu-FC-ReLu-FC-SoftMax).
[0050] In the training stage, the image data of the i-th sample in the dataset D k in the dataset is input into the shared feature extractor, and the shared feature extractor extracts the quality features of the image data and the regressor R k corresponding to the dataset D k maps the quality features of the image data to a quality score The domain discriminator D c predicts the domain label of the sample according to the quality features of the image data to obtain the probability set of the sample corresponding to each domain label represents the probability that the domain label of the i-th sample is predicted to be d k of.
[0051] Construct quality-sensitive triplet loss, quality prediction loss, monotonicity loss, and domain discrimination loss, and integrate all losses to train a domain-generalized image quality assessment model;
[0052] Firstly, the quality discriminant triplets are screened according to the quality score differences, and the quality features are extracted by a shared feature extractor to construct a quality-sensitive triplet loss.
[0053] M datasets are input into the domain-generalized image quality assessment model, and image samples from the mixed dataset are selected according to the quality score difference. The quality discrimination triplet is selected in the training process. * indicates that the source of the dataset is not distinguished. The quality difference threshold ε is set. As anchor image Samples with a quality score difference less than or equal to ε are considered positive samples Samples with a quality score difference greater than ε are considered negative samples.
[0054] Extracting selected quality discrimination triplets with the help of shared feature extractor The quality features of the proposed method are used to construct a quality-sensitive triplet loss, which shortens the distance between the anchor point and the positive sample and extends the distance between the anchor point and the negative sample in the feature space.
[0055] The loss of quality-sensitive triples is:
[0056]
[0057] Wherein, N is the total number of samples in the M data sets; δ is the minimum value of the distance difference between the preset positive and negative samples and the anchor image in the feature space.
[0058] In one embodiment of the present invention, the quality difference threshold ε is set to 0.1, and δ is set to 1.
[0059] Secondly, the regressors corresponding to each dataset that constitutes the training dataset map the quality features into quality scores and compare them with the quality score labels to construct quality prediction loss and monotonicity loss;
[0060]
[0061] in, represents the root mean square error;
[0062] In order to ensure that the quality scores obtained by mapping the regressor for each data set are consistent with the order of the quality score labels, the monotonicity loss of the regressor for each data set is constructed as follows:
[0063]
[0064] Among them, represents the dataset D k corresponding regressor R k monotonicity loss; Err (i,j),k is the dataset D k in and monotonicity error:
[0065]
[0066] Among them, and respectively represent the quality score label of the input image sample pair and the quality score predicted by the model during training. sign(·) is the sign function. The total monotonicity loss for M regression branches is:
[0067]
[0068] The domain discriminator predicts the dataset to which the extracted quality features belong, and compares the prediction result with the domain label to construct the domain discriminant loss:
[0069]
[0070] Among them, represents the cross-entropy loss;
[0071] Integrate all losses to construct the global loss to train the domain generalization image quality evaluation model:
[0072] L = αL pre + βL mono + γL triplet + λL dom
[0073] Among them, α, β, γ, and λ are weight hyperparameters for balancing each loss.
[0074] Based on the global loss, repeatedly iterate to optimize the domain generalization image quality evaluation model.
[0075] In one embodiment of the present invention, α, β, γ, and λ are respectively set to 1, 0.3, 0.1, and 1. The SGD optimizer is used to optimize the model weights, the weight decay weightdecay is set to 0.0001, and the momentum momentum is set to 0.9. The learning rate is set to 0.005. The total number of training steps is 9000, and the learning rate linearly increases from 0 to the preset learning rate in the first 1000 training steps and decays with a cosine decay function in the remaining training steps.
[0076] After training, the image data sample x to be evaluated tInput it into the trained domain generalization image quality evaluation model. After the shared feature extractor extracts the image data quality feature G(x t )), the image data quality feature G(x t ) is respectively input into the domain discriminator and M regressors. Using the result output by the domain discriminator as the weight, the quality scores output by the M regressors are weighted and fused to obtain the quality score
[0077]
[0078] of the image data sample to be evaluated. In an embodiment of the present invention, the Spearman Rank Order Correlation Coefficient (SROCC) and the Pearson Linear Correlation Coefficient (PLCC) are calculated between the final predicted score and the actual quality score label to verify the effectiveness of the verification scheme.
[0079] The present invention proposes a domain generalization image quality evaluation method for multi-dataset mixing. The shared feature extractor learns the image quality features of multiple datasets and constructs a quality-sensitive feature space through quality-sensitive triplet loss. The unique multi-regressor structure effectively retains the quality discrimination knowledge unique to each domain. In the test stage, the discrimination result of the domain discriminator gives a higher weight to the predicted value of the regressor corresponding to the training dataset most similar to the unknown target dataset, effectively integrating the prediction results of each regressor and improving the generalization performance. This method effectively solves the problem that the domain generalization image quality evaluation method for multi-dataset mixing cannot effectively integrate the quality discrimination knowledge of each dataset, and effectively improves the generalization of the deep learning-based image quality evaluation method.
[0080] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image quality evaluation method for domain generalization, characterized in that The following steps are involved: Step 1: Obtain M datasets with image quality score labels, set a domain label for each dataset, and construct the M datasets with domain labels as training sets; Step 2: Use the training set to train a domain-generalized image quality assessment model; The domain generalized image quality assessment model includes a shared feature extractor, a regressor and a domain discriminator; the number of the regressors is M, corresponding to the M data sets; When training a domain-generalized image quality assessment model, the quality-sensitive triplet loss, quality prediction loss, monotonicity loss, and domain discrimination loss are calculated, and a global loss is constructed through weighting. Step 3: Input the image data sample to be evaluated into the trained domain-generalized image quality evaluation model. After the shared feature extractor extracts the image data quality features, the image data quality features are input into the domain discriminator and M regressors respectively. The quality scores output by the M regressors are weighted and fused using the results output by the domain discriminator as the weight to obtain the quality score of the image data sample to be evaluated, thus completing the image quality evaluation.
2. The method for evaluating image quality with domain generalization according to claim 1, characterized in that: In the said step 1, there are M data sets in the training set D, D = {D1, D2,..., D k ,..., D M}, and there are N k samples in the k-th data set D k . and are respectively the image data and the quality score of the i-th sample in the k-th data set, and d k is the domain label of the k-th data set.
3. The method for evaluating image quality with domain generalization according to claim 2, wherein: In step 2, the training data set D k The image data of the i-th sample in Input to the shared feature extractor, which extracts image data quality features With dataset D k The corresponding regressor R k Image data quality features Mapping to quality scores Domain Discriminator D c According to the image data quality characteristics Predict the domain label of the sample and get the sample Probability set corresponding to each field label Indicates that the domain label of the i-th sample is predicted to be d k probability.
4. The method for evaluating image quality with domain generalization according to claim 3, wherein: The calculation method of the quality sensitive triplet loss in step 2 is: M datasets are input into the domain generalization image quality evaluation model. For each sample regards its own image data as the anchor image, where the superscript "*" indicates that the dataset from which the sample is sourced is not distinguished; randomly select samples from the M datasets and satisfy ε is a preset quality difference threshold, and is used as the positive sample of the anchor image; randomly select samples from the M datasets and satisfy regard as the negative sample of the anchor image; In order to bring the distance between the anchor image and the positive sample closer and push the distance between the anchor image and the negative sample farther in the feature space, the loss of the quality-sensitive triplet is calculated as: Wherein, N is the total number of samples in the M data sets; δ is the minimum value of the distance difference between the preset positive and negative samples and the anchor image in the feature space.
5. The method for evaluating image quality with domain generalization according to claim 4, wherein: The calculation method of the quality prediction loss in step 2 is: Among them, represents the root mean square error; The monotonicity loss is calculated as follows: Among them, represents the dataset D k The corresponding regressor R k of the monotonicity loss, Err (i,j),k is the dataset D k in and the monotonicity error of, The calculation method of the domain discrimination loss is: Among them, represents the cross-entropy loss; The quality-sensitive triplet loss, quality prediction loss, monotonicity loss, and domain discrimination loss are weighted to construct the global loss L: L = αL pre + βL mono + γL triplet + λL dom Among them, α, β, γ, and λ are weight hyperparameters for balancing various losses.
6. The method for evaluating image quality with domain generalization according to claim 3, characterized in that: In step 3, the image data sample x to be evaluated t is input into the trained domain generalization image quality evaluation model, and the shared feature extractor extracts the image data quality feature G(x t ). After that, the image data quality feature G(x t ) is respectively input into the domain discriminator and M regressors, and the result output by the domain discriminator is used as the weight to weight and fuse the quality scores output by the M regressors, so as to obtain the quality score of the image data sample to be evaluated 7. The method for evaluating image quality with domain generalization according to claim 1, characterized in that: After constructing the training set in step 1, data enhancement is performed on the image data of the training data set, and the enhancement means include random rotation, random cropping, random mirroring, and scaling; in step 2, the shared feature extractor uses the feature extraction layer of ResNet18 pre-trained on ImageNet to superimpose two fully connected layers with ReLu as the activation function, expressed as FC-ReLu-FC-ReLu; the structure of the regressor is two linear layers, expressed as FC-FC; the structure of the domain discriminator is expressed as FC-ReLu-FC-ReLu-FC-SoftMax.
8. A computer device / apparatus / system, comprising a memory, a processor, and a computer program stored on the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
No-reference image objective quality evaluation method based on structural distortion
CN106683079A
Blind image quality evaluation method based on semantic weighted contrast learning
CN118470395A