Mixed distortion image quality evaluation method fusing online large model and offline small model

A hybrid IQA method using large and small models for feature extraction and perception improves image quality evaluation accuracy and robustness, addressing inefficiencies in existing methods by mimicking human visual perception.

CN120321386APending Publication Date: 2025-07-15BEIJING UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335931.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing image quality evaluation methods are difficult to efficiently and accurately simulate human visual perception under complex scenes and mixed distortion types, and subjective evaluation is costly and objective evaluation is difficult to match human visual perception in some scenarios.

Method used

A hybrid distortion image quality evaluation method that combines online large models and offline small models is adopted. Through global-local feature extraction, mixed dimension feature coding and quality perception steps, combined with the group wisdom of the big model and the local expert experience of the small model, a new feature extraction and coding method is designed to adapt to a variety of image processing tasks and application scenarios.

Benefits of technology

The accuracy and robustness of image quality evaluation are improved under complex scenes and mixed distortion types, and the visual quality perception effect is achieved close to the human eye.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_12
    Figure SMS_12
  • Figure FDA0005321737720000022
    Figure FDA0005321737720000022
Patent Text Reader

Abstract

The invention discloses a mixed distortion image quality evaluation method fusing an online large model and an offline small model, and belongs to the field of image quality evaluation. According to the invention, on the basis of a large language model and a convolutional neural network, a mixed distortion image quality evaluation method fusing an online large model and an offline small model is constructed through three steps of global-local feature extraction, mixed dimension feature coding and quality perception. On-line feature extraction is performed through a large model, and a new point-to-surface prompting mode is designed, so that the large model can perform overall perception and quantitative evaluation on an image from six dimensions of contrast, brightness, naturalness, ambiguity, blocking effect and noise level. And local features are coded through a mixed dimension feature encoder, and domain transformation is carried out on global features, so that effective fusion of subsequent features is realized. The method shows excellent accuracy and robustness in a complex scene and a mixed distortion type, and obtains a visual quality evaluation effect close to human eyes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image quality assessment, and based on large language models and convolutional neural networks, a hybrid distortion image quality assessment method integrating an online large model and an offline small model is constructed. Background Technique

[0002] Image Quality Assessment (IQA) is a fundamental research direction in the fields of image processing and computer vision, aiming to quantify image quality through a computational model so that the evaluation results are consistent with human subjective perception. Specifically, an image with high subjective quality should obtain a higher IQA score. With the rapid development of digital image acquisition and transmission technologies, IQA plays an increasingly important role in all aspects of image processing, such as image acquisition, transmission, compression, restoration, and enhancement.

[0003] The development process of IQA reflects the continuous progress of technology. In the early days, image quality databases only contained a small amount of image data and had a single type of distortion. Nowadays, with the explosion of data volume and the diversification of distortion types, image quality databases have become richer and more complex. In terms of algorithms, early research mainly relied on the combination of feature engineering and traditional machine learning algorithms, while nowadays, end-to-end deep learning models have become the mainstream. Evaluation metrics have also evolved from single metrics to diverse application-related metrics to adapt to different application scenarios and requirements.

[0004] The core task of IQA is to quantify the degree of image quality degradation, which is mainly divided into two categories: subjective quality assessment and objective quality assessment. Subjective quality assessment relies on the subjective judgment of observers and usually obtains the Mean Opinion Score (MOS) or Differential Mean Opinion Score (DMOS) through subjective psychological experiments to quantitatively describe the visual quality of images. These subjective evaluation results are widely used to establish image quality assessment databases, providing a basis for the development and verification of objective quality assessment algorithms. Depending on different distortion types, experimental methods, etc., the types of IQA databases are rich and diverse, covering a variety of application scenarios and requirements.

[0005] Objective quality assessment predicts the visual quality of images by designing IQA algorithms. Although significant progress has been made in the field of IQA, existing methods still have some limitations. For example, subjective evaluation methods are time-consuming and costly, making it difficult to apply them on a large scale; objective evaluation methods, although efficient, are still difficult to fully match human visual perception in some complex scenarios. In addition, with the increase in image data volume and the diversification of application scenarios, higher requirements are placed on the accuracy and generalization ability of IQA algorithms. Therefore, developing more efficient, accurate, and generalized IQA methods remains an urgent problem to be solved.

[0006] A hybrid distortion image quality assessment method that combines an online large model and an offline small model proposed by the present invention can more accurately evaluate the image quality by integrating the collective wisdom of the large model and the local expert experience of the small model. Compared with traditional evaluation methods based on single feature extraction or evaluation models that simply combine features, the image quality assessment method proposed by the present invention has higher accuracy and robustness in complex scenarios and mixed distortion types, and can better simulate the perceptual characteristics of the human visual system, thus achieving a visual quality perception effect closer to the human eye. Summary of the Invention

[0007] The present invention proposes a hybrid distortion image quality assessment method that combines an online large model and an offline small model. This method is realized through three steps: global-local feature extraction, hybrid dimension feature encoding, and quality perception, which can significantly improve the consistency between the objective evaluation results of image quality and subjective perception. The present invention shows excellent accuracy and robustness in complex scenarios and mixed distortion types, and obtains a visual quality evaluation effect close to the human eye.

[0008] The present invention is realized through the following technical solutions, including the following steps:

[0009] The first step: global-local feature extraction;

[0010] The second step: hybrid dimension feature encoding;

[0011] The third step: quality perception.

[0012] The creativity of the present invention is mainly reflected in:

[0013] (1) The present invention proposes a new framework for hybrid distortion image quality assessment that integrates online large models and offline small models. This framework does not rely on specific global feature extractors and local feature extractors. This means that users can select different global and local feature extractors according to specific application scenarios and requirements. For example, for the global feature extractor, large models such as ChatGPT and Wenxin Yiyan can be used for online feature extraction. For the local feature extractor, complex deep neural network series such as the MobileNet lightweight neural network series, ResNet, and DenseNet can be used for offline feature extraction. This design improves the flexibility and universality of the framework, enabling it to adapt to a variety of different image processing tasks and practical application scenarios.

[0014] (2) Considering that global large models are trained with large-scale and multi-domain data, demonstrating extensive knowledge coverage capabilities and enabling objective and accurate assessment of images, the present invention uses large models for online feature extraction and designs a new point-to-plane prompting method to enable large models to globally perceive and quantitatively evaluate images from six dimensions: contrast, brightness, naturalness, blur, blocking effect, and noise level.

[0015] (3) A new hybrid-dimensional feature encoder is designed. First, for local features, multiple stacked MLPs are used to encode the local features of the distorted image, the local features of the reference image, and the residual local features between the distorted image and the reference image respectively. Second, for global features, 1×1 convolution is used to perform domain transformation on the global features extracted by the large model to facilitate effective fusion of subsequent features. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of the hybrid distortion image quality assessment method that integrates online large models and offline small models designed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following details the embodiments of the present invention. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0018] Embodiment:

[0019] First step: Global-local feature extraction

[0020] For global feature extraction, considering that global large models are trained with large-scale and multi-domain data, demonstrating extensive knowledge coverage capabilities and enabling objective and accurate evaluation of images, the present invention designs a global feature extractor based on a large model. Through the designed point-to-plane prompting method, the large model is utilized for online feature extraction. Specifically, first, the large model is informed of the six aspects of features to be extracted: "From a visual perspective, please score the input image in terms of contrast, brightness, naturalness, blurriness, blocking effect, and noise level. The score range is an integer between 1 and 99." Secondly, the scoring criteria are described in detail: "The specific scoring criteria are as follows: The higher the image contrast, the higher the contrast score; the brighter the image, the higher the brightness score; the more natural elements the image contains, the higher the naturalness score; the less clear and more blurred the image, the lower the blurriness score; the more compression artifacts in the image, the lower the blocking effect score; the more noise in the image, the lower the noise level score." Through the above point-to-plane prompting method, the large model is used to perform overall perception and quantitative evaluation from six dimensions of contrast, brightness, naturalness, blurriness, blocking effect, and noise level, generating the global feature G of the reference image R and the global feature G of the distorted image D .

[0021] For local feature extraction, a small model is used as the local feature extractor. For the convenience of readers' understanding, DenseNet121 is used as the local feature extractor in this article. Specifically, first, the input distorted image and reference image are randomly cropped into n image patches respectively, and these image patches are input into the local feature extractor constructed by DenseNet121 to extract the preliminary local feature L of the reference image R and the local feature L of the distorted image D . In the present invention, n = 5.

[0022] Step 2: Hybrid-dimensional feature encoding

[0023] To further process and fuse the global features and local features, the present invention constructs a hybrid-dimensional feature encoder. First, for the global features G R and G D extracted by the online large model, domain transformation is performed on them through 1×1 convolution to generate the encoded global feature of the reference image the encoded global feature of the distorted image and the residual global feature of the two Secondly, for the local features L R and L D extracted by the offline small model, the present invention stacks multiple multi-layer perceptron (MLP) modules and concatenates 1 1×1 convolution to generate the encoded local feature of the reference image the encoded local feature of the distorted image and the residual local features of both It should be noted that the above six features have the same dimension. Finally, in the present invention, the global and local features of the reference image are first added to obtain the global and local features of the distorted image are added to obtain the global and local features of the residual are added to obtain Then, a concatenation operation is performed on the above three kinds of features to obtain the final hybrid feature F that combines global features and local features.

[0024] Step 3: Quality perception

[0025] To better perceive the quality of the image, the present invention designs a quality perception module. This module consists of two fully connected layers, and realizes mapping the hybrid feature F obtained in the second step to the quality score of the distorted image:

[0026]

[0027] where, Ψ(F) is the number (i.e., dimension) of the hybrid feature F, F i is the i-th feature value in the hybrid feature F. F passes through the first fully connected layer to generate the feature vector F2, so Ψ(F2) is the number (i.e., dimension) of the feature vector F2, are the weights and biases of the first fully connected layer and the weights and biases of the second fully connected layer respectively, Υ RB is a ReLU activation layer in series with a BN layer, is a convolution operation, and score is the quality score of the distorted image measured by the quality perception module.

Claims

1. The present invention adopts the following technical solutions and implementation steps: 1) Global-local feature extraction For global feature extraction, considering that global large models are trained with large-scale and multi-domain data, demonstrating extensive knowledge coverage capabilities and enabling objective and accurate evaluation of images, the present invention designs a global feature extractor. Through the designed point-to-face prompting method, the large model is used for online feature extraction. Specifically, first, the large model is informed of the six aspects of features to be extracted: "From a visual perspective, please rate the input image in terms of contrast, brightness, naturalness, blurriness, blocking effect, and noise level. The rating range is an integer between 1 and 99." Second, the rating criteria are described in detail: "The specific rating criteria are as follows: The higher the image contrast, the higher the contrast score; the brighter the image, the higher the brightness score; the more natural elements the image contains, the higher the naturalness score; the less clear and more blurred the image, the lower the blurriness score; the more compression artifacts in the image, the lower the blocking effect score; the more noise in the image, the lower the noise level score." Through the above point-to-face prompting method, the large model conducts overall perception and quantitative evaluation from six dimensions of contrast, brightness, naturalness, blurriness, blocking effect, and noise level, generating the global feature G of the reference image R and the global feature G of the distorted image D . For local feature extraction, a small model is used as the local feature extractor. For the convenience of readers' understanding, DenseNet121 is adopted as the local feature extractor in this paper. Specifically, first, the input distorted image and reference image are randomly cropped into n image patches respectively, and these image patches are input into the local feature extractor constructed by DenseNet121 to extract the preliminary local features L of the reference image R and the local features L of the distorted image D . In the present invention, n = 5. 2) Hybrid dimensionality feature encoding To further process and fuse global features and local features, the present invention constructs a hybrid-dimensional feature encoder. First, for the global features G R and G D extracted by the online large model, domain transformation is performed on them through 1×1 convolution to generate the encoded global features of the reference image the encoded global features of the distorted image and the residual global features of the two Second, for the local features L R and L D extracted by the offline small model, the present invention stacks multiple multi-layer perceptron (MLP) modules and concatenates 1×1 convolution to generate the encoded local features of the reference image the encoded local features of the distorted image and the residual local features of the two It should be noted that the above six features have the same dimension. Finally, the present invention first adds the global and local features of the reference image to obtain adds the global and local features of the distorted image to obtain adds the residual global and local features to obtain Then, a concatenation operation is performed on the above three features to obtain the final hybrid feature F that fuses global features and local features. 3) Quality perception To better perceive the quality of the image, the present invention designs a quality perception module. This module consists of two fully connected layers and realizes mapping the hybrid feature F obtained in the second step to the quality score of the distorted image: Among them, Ψ(F) is the number (i.e., dimension) of the mixed feature F, and F i is the i-th feature value in the mixed feature F. F generates the feature vector F2 after passing through the first fully connected layer. Therefore, Ψ(F2) is the number (i.e., dimension) of the feature vector F2, are the weights and biases of the first fully connected layer and the weights and biases of the second fully connected layer respectively. Υ RB is a ReLU activation layer concatenated with 1 BN layer, is a convolution operation, and score is the quality score of the distorted image measured by the quality perception module.