Image quality evaluation method based on semantic, distortion and scene feature collaborative learning
By using convolutional neural networks and three-path quality perception modules in image quality evaluation, combining semantics, distortion and scene features, the accuracy and robustness problems of the existing technology in complex scenes and multiple distortion types are solved, and efficient image quality evaluation is achieved.
Patent Information
- Application Number
- CN202510317689.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
Existing image quality evaluation methods are difficult to achieve high accuracy and robustness in complex scenes and multiple distortion types, and it is difficult to simulate the perceptual characteristics of human vision systems.
A three-path quality perception module is constructed through collaborative learning of semantics, distortion and scene features, and a multi-scale feature extraction and quality evaluation regressor is used to achieve a comprehensive evaluation of image quality.
It significantly improves the consistency between the objective evaluation results of image quality and subjective perception, shows excellent accuracy and robustness, and can achieve visual quality evaluation effect close to the human eye.
Smart Images

Figure FT_1 
Figure SMS_7 
Figure QLYQS_2
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image quality assessment. Based on a convolutional neural network, a method for evaluating image quality based on the collaborative learning of semantic, distortion, and scene features is constructed. Background Art
[0002] Image Quality Assessment (IQA) is a fundamental research direction in the fields related to image processing and computer vision. Its purpose is to use a computational model to measure the image quality so that the result is consistent with the subjective quality, that is, the better the subjective quality of an image, the higher its IQA score should be. With the rapid development of digital image and transmission technologies, IQA has also become more important in the fields of image acquisition, transmission, compression, restoration, enhancement, etc. From the early image quality databases with only a small amount of image data and a single type of distortion to the current image quality databases with a large amount of image data and a rich variety of distortion types, from the combination of early feature engineering and traditional machine learning algorithms to the current end-to-end deep learning models, and from the relatively single evaluation metrics in the early stage to the diverse application-related evaluation metrics nowadays, all have witnessed the rapid development process in the field of IQA research. IQA refers to quantifying the degree of degradation of image quality, mainly including subjective quality evaluation and objective quality evaluation. Subjective quality evaluation refers to the visual quality of an image judged subjectively by an observer, and generally uses MOS or DMOS metrics to quantitatively describe it. Generally, MOS or DMOS scores are obtained through subjective psychological experiments, and an image quality evaluation database is established. According to different distortion types, experimental methods, etc., the types of IQA databases are very rich. Objective quality evaluation refers to predicting the visual quality of an image by designing an IQA algorithm, which can be divided into three categories: Full reference IQA (FR-IQA), Reduced reference IQA (RR-IQA), and No-reference IQA (NR-IQA).
[0003] The method for evaluating image quality based on the collaborative learning of semantic, distortion, and scene features can more accurately evaluate the quality of an image. This evaluation model that integrates multiple features has higher accuracy and robustness than traditional evaluation methods that extract a single feature or simple combined feature evaluation models in complex scenarios and various distortion types, and can better simulate the perception characteristics of the human visual system, thus achieving a visual quality evaluation effect closer to the human eye.
[0004] By fusing semantic, distortion, and scene features, the model can more accurately describe the feature distributions of the reference image and the distorted image and deeply analyze the difference distribution between the reference image and the distorted image. This difference analysis is crucial for understanding the degree of image quality degradation and perceiving the visual quality of the distorted image, and helps to guide and optimize image restoration and enhancement algorithms. In multimedia content management, the model can be used to automatically screen high-quality images, improving the availability of content and the user experience. In industrial inspection and quality control, the model can be used to automatically detect defects and quality problems in images, improving production efficiency and product quality. Conducting research on the image quality evaluation method based on the collaborative learning of semantic, distortion, and scene features not only has important theoretical significance but also demonstrates strong potential and wide applicability in practical applications. The invention provides new ideas and methods for the development of the field of image quality evaluation and has important innovation and practicality. Summary of the Invention
[0005] The present invention proposes an image quality evaluation method based on the collaborative learning of semantic, distortion, and scene features, as Figure 1 shown. The model is implemented by five steps including image preprocessing, multi-scale feature extraction, three-channel quality perception, quality assessment, and model training and testing, and can significantly improve the consistency between the objective evaluation results of image quality and subjective perception. The present invention shows excellent accuracy and robustness under complex scenes and various distortion types, and obtains a visual quality evaluation effect close to that of the human eye.
[0006] The present invention is realized through the following technical solutions, including the following steps:
[0007] The first step: Input image preprocessing;
[0008] The second step: Multi-scale feature extraction;
[0009] The third step: Three-channel quality perception;
[0010] The fourth step: Quality assessment;
[0011] The fifth step: Model training and testing.
[0012] The creativity of the present invention is mainly reflected in:
[0013] (1) The present invention proposes a new general framework for image quality assessment, which does not rely on specific image patch acquisition methods and feature extractors. This means that users can select different image patch acquisition methods and feature extractors according to specific application scenarios and requirements. For example, for image patch acquisition methods, random cropping, saliency-based cropping, etc. can be adopted. For feature extractors, complex deep neural network series such as the MobileNet lightweight neural network series, ResNet, DenseNet, etc. can be used. This design improves the flexibility and universality of the framework, enabling it to adapt to a variety of different image processing tasks and practical application scenarios.
[0014] (2) The present invention designs a new three-channel quality perception module. First, the distortion channel extracts semantic information from the distorted image to detect whether the distorted image can still accurately express sufficient semantics. When there is no semantic information at all, it indicates that the distorted image has no quality or extremely poor quality. Secondly, the difference channel compares the reference image with the distorted image to obtain distortion features, which are used to detect minor degradations in distorted images with certain semantic information. Finally, the reference channel extracts scene features from the reference image to assist the semantic features and distortion features of the above two channels. Obviously, the more complex the scene, the more serious the damage to the image quality caused by various distortions. The image quality is comprehensively evaluated from different dimensions through the three-channel quality perception module. Brief Description of the Drawings
[0015] Figure 1 It is a flowchart of the "Image Quality Evaluation Method Based on Cooperative Learning of Semantic, Distortion and Scene Features" designed by the present invention. Detailed Embodiments
[0016] The following details the embodiments of the present invention. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0017] Embodiment:
[0018] The first step: Input image preprocessing;
[0019] The distorted image and the reference image are respectively randomly cropped into n image patches of size h×w, which are respectively denoted as the distorted image patch set P D (where P D ={P D1 , P D2 , …, P Dn ) and the reference image patch set P R (where P R ={P R1 , P R2 , …, P RnCompared with a single image block input, multiple image block inputs can more effectively provide local fine-grained distortion information of a single image block and global coarse-grained combination information across image blocks at the same time, thereby more comprehensively evaluating image quality. For the convenience of readers' understanding, in the present invention, n=10, h=w=256.
[0020] Step 2: Multi-scale feature extraction;
[0021] First, the present invention applies a convolutional neural network that performs well in image classification tasks as a perceptual feature extractor and shares it, and transforms the distorted image block set P D and the reference image block set P R The shared feature extractor is used to extract features. For ease of understanding, the present invention takes ResNet50 as an example as a shared feature extractor (other networks such as MobileNet, DenseNet, etc. can be replaced similarly). Features of different scales are extracted through its four network layers, and then a feature vector F is obtained through a 1×1 convolution layer, a global average pooling layer, and a concatenation operation. The feature vector corresponding to the distorted image is F D , the eigenvector corresponding to the reference image is F R .
[0022] Step 3: Three-channel quality perception;
[0023] In order to more comprehensively evaluate the image quality, the present invention constructs a three-channel quality perception module, which is as follows:
[0024] First, the distortion path is constructed. This path starts from the feature vector F D Extracting semantic features from distorted images To detect whether the distorted image can still accurately express sufficient semantic information. When no semantic information can be detected, it indicates that the distorted image has no quality or very poor quality. Secondly, a difference pathway is constructed. This pathway compares F R With F D The difference (i.e., F R -F D ), extract distortion features It is used to detect slight degradation in distorted images with certain semantic information. It is worth noting that even if the distortion is large, it does not mean that the semantic information is completely lost (for example, brightness shift). Therefore, the distortion feature cannot be relied on alone, and a comprehensive evaluation needs to be combined with the semantic feature. Therefore, the present invention designs a third path: the reference path. This path starts from the feature vector F R Extract scene features The semantic features and distortion features of the above channels are used to assist. The image quality is comprehensively evaluated from different dimensions through the three-channel quality perception module.
[0025] Step 4: Quality assessment;
[0026] The quality assessment regressor in the present invention consists of two fully connected layers: mapping the feature vector concatenated by semantic features distortion features and scene features to the quality score of the corresponding distorted image.
[0027] Step 5: Model training and testing;
[0028] After constructing the image quality evaluation method based on the collaborative learning of semantic, distortion, and scene features proposed in the present invention through the above four steps, the L1 loss function is used to train the model to enable the model to better predict the image quality score. The L1 loss function is also called the mean absolute error, which refers to the average of the absolute differences between the model prediction value and the true value. Since it is less sensitive to outliers and can handle the noise in the data more robustly, the L1 loss function is often used to handle regression problems, and its specific formula is expressed as follows: where P Di is the set of image patches corresponding to the i-th distorted image, P Ri is the set of image patches of the original image corresponding to the i-th distorted image, y i is the subjective opinion score of the i-th distorted image, Net(·) is the network model of the image quality evaluation method based on the collaborative learning of semantic, distortion, and scene features proposed in the present invention, θ is the parameter of the model, and N represents that there are a total of N distorted images. After the model is trained with the L1 loss function, it is tested to obtain a robust image quality score.
Claims
1. Input image preprocessing: The distorted image and the reference image are randomly cropped into n image blocks of size h×w, respectively, denoted as the distorted image block set P D (in, P D = {P D1 ,P D2 ,…,P Dn }) and reference image block set P R (Among them, P R = {P R1 ,P R2 ,…,P Rn Compared with a single image block input, multiple image block inputs can more effectively provide local fine-grained distortion information of a single image block and global coarse-grained combination information across image blocks at the same time, thereby more comprehensively evaluating image quality. For the convenience of readers' understanding, in the present invention, n=10, h=w=256.
2. Multi-scale feature extraction: First, the present invention applies a convolutional neural network that performs well in image classification tasks as a perceptual feature extractor and shares it, and transforms the distorted image block set P D and the reference image block set P R The shared feature extractor is used to extract features. For ease of understanding, the present invention takes ResNet50 as an example as a shared feature extractor (other networks such as MobileNet, DenseNet, etc. can be replaced similarly). Features of different scales are extracted through its four network layers, and then a feature vector F is obtained through a 1×1 convolution layer, a global average pooling layer, and a concatenation operation. The feature vector corresponding to the distorted image is F D , the feature vector corresponding to the reference image is F R .
3. Three-channel quality perception: In order to evaluate the image quality more comprehensively, the present invention constructs a three-path quality perception module, which is as follows. First, the distortion path is constructed. This path starts from the feature vector F D Extracting semantic features from distorted images To detect whether the distorted image can still accurately express sufficient semantic information. When no semantic information can be detected, it indicates that the distorted image has no quality or very poor quality. Secondly, a difference pathway is constructed. This pathway compares F R With F D The difference (i.e., F R -F D ), extract distortion features It is used to detect slight degradation in distorted images with certain semantic information. It is worth noting that even if the distortion is large, it does not mean that the semantic information is completely lost (for example, brightness shift). Therefore, the distortion feature cannot be relied on alone, and a comprehensive evaluation needs to be combined with the semantic feature. Therefore, the present invention designs a third path: the reference path. This path starts from the feature vector F R Extract scene features The semantic features and distortion features of the above channels are used to assist. The image quality is comprehensively evaluated from different dimensions through the three-channel quality perception module.
4. Quality Assessment: The quality assessment regressor in this invention consists of two fully connected layers: Semantic features Distortion characteristics and scene features The concatenated feature vector is mapped to the quality score of the corresponding distorted image.
5. Model training and testing: After the image quality evaluation method based on collaborative learning of semantics, distortion and scene features proposed by the present invention is constructed through the above four steps, the L1 loss function is used to train the model so that the model can better predict the image quality score. The L1 loss function is also called the mean absolute error, which refers to the average value of the absolute difference between the model prediction value and the true value. Due to its low sensitivity to outliers and the ability to more robustly handle noise in the data, the L1 loss function is often used to deal with regression problems, and its specific formula is expressed as follows: Where P Di is the image block set corresponding to the i-th distortion image, P Ri is the image block set of the original image corresponding to the i-th distorted image, y i is the subjective opinion score of the i-th distortion image, Net(·) is the network model of the image quality evaluation method based on the collaborative learning of semantics, distortion and scene features proposed in this invention, θ is the parameter of the model, and N indicates that there are N distortion images in total. After the model is trained by the L1 loss function, it is tested to obtain a robust image quality score.