Image quality evaluation method and device, electronic equipment, storage medium and computer program product

By introducing a quality enhancement module into the image quality evaluation model, extracting the underlying quality features and combining significance area sampling and adapter alignment processing, the problem of existing models being difficult to evaluate fine-grained image quality is solved, and more accurate image quality comparison and comprehensive service are achieved.

CN120259855APending Publication Date: 2025-07-04SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271909.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing image quality evaluation models are difficult to accurately compare and evaluate the quality of fine-grained images, especially images with similar content.

Method used

The quality enhancement module is introduced into the existing image quality evaluation model to extract the underlying quality features, and image quality evaluation is carried out through significance area sampling, quality feature extraction and adapter alignment processing, combining large language models.

Benefits of technology

It improves the accuracy of fine-grained image quality comparison, can consider image quality characteristics more comprehensively, provide quality comparison results and causal reasoning descriptions, and improves the function of the image quality evaluation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259855A_ABST
    Figure CN120259855A_ABST
Patent Text Reader

Abstract

The invention relates to an image quality evaluation method and device, electronic equipment, a storage medium and a computer program product. The method comprises the steps of obtaining at least two target images to be subjected to quality evaluation; extracting target text embedding corresponding to the target quality problem description through a text compiler; respectively extracting target visual embedding of the at least two target images through a visual encoder; respectively extracting target quality embedding of the at least two target images through a quality enhancement module; splicing the target text embedding and the target visual embedding and the target quality embedding of the at least two target images; and predicting a quality evaluation target answer about the at least two target images based on the target splicing embedding through the large language model. Therefore, by additionally extracting the quality features of the relatively bottom layer, the quality features of the images can be fully considered when the quality comparison is performed on the images, and the image quality can be accurately distinguished even for some images with relatively similar contents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to an image quality assessment method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] With the rapid development of intelligent terminals and social media, users have higher and higher requirements for the quality of the images captured and the images transmitted in the network. Therefore, it is also an increasingly important requirement to optimize the imaging system of intelligent terminals and the social network image transmission system based on the evaluation results of image quality. For example, the Image Signal Processor (ISP) module in the imaging system of intelligent terminals is mainly responsible for converting the sensor RAW data into RGB images. Therefore, the optimization of the ISP parameters can be guided by evaluating the image quality.

[0003] In related technologies, there are various image quality assessment methods, such as DepictQA. DepictQA is based on the LLava model, including an image encoder, a text encoder, and a large language model of the Contrastive Language-Image Pre-training (CLIP) model. It uses the image encoder of the CLIP model as an image feature extractor. However, limited by the pre-training objective of CLIP itself, the Image Encoder has good ability to obtain semantic (also known as content features or visual features) features of images, but is poor in extracting quality features of images. At this time, for some images with relatively similar contents (also known as fine-grained images), the feature similarity extracted by the Image Encoder will be relatively high, which will cause the subsequent large language model to be difficult to accurately distinguish the quality of these images with relatively similar contents, that is, it is difficult for related technologies to accurately compare and evaluate the quality of fine-grained images. Summary of the Invention

[0004] The present disclosure provides an image quality assessment method, apparatus, electronic device, storage medium, and computer program product to at least solve the problem in the above related technologies that it is difficult to accurately compare and evaluate the quality of fine-grained images.

[0005] According to a first aspect of an embodiment of the present disclosure, an image quality assessment method is provided. The image quality assessment method is implemented based on an image quality assessment model, which includes a text compiler, a visual encoder, and a large language model. The image quality assessment model further includes a quality enhancement module. The image quality assessment method includes: obtaining at least two target images to be quality-assessed, where the at least two target images are corresponding to a target quality problem description for the at least two target images; extracting a target text embedding corresponding to the target quality problem description through the text compiler; respectively extracting target visual embeddings of the at least two target images through the visual encoder; respectively extracting target quality embeddings corresponding to at least one quality index of the at least two target images through the quality enhancement module; splicing the target text embedding, the target visual embeddings of the at least two target images, and the target quality embeddings to obtain a target spliced embedding; and predicting a quality assessment target answer for the at least two target images based on the target spliced embedding through the large language model.

[0006] Optionally, the quality enhancement module includes a quality adapter; the splicing the target text embedding, the target visual embeddings of the at least two target images, and the target quality embeddings to obtain a target spliced embedding includes: performing alignment processing on the target quality embeddings of the at least two target images through the quality adapter to obtain aligned target quality embeddings aligned with the target text embedding; and splicing the target text embedding, the target visual embeddings of the at least two target images, and the aligned target quality embeddings to obtain the target spliced embedding.

[0007] Optionally, the quality enhancement module includes a quality feature extraction module; the respectively extracting target quality embeddings corresponding to at least one quality index of the at least two target images through the quality enhancement module includes: respectively extracting quality feature values corresponding to the at least one quality index of the at least two target images through the quality feature extraction module, and generating the target quality embeddings based on the extracted quality feature values corresponding to the at least one quality index.

[0008] Optionally, the quality enhancement module further includes a salient region sampling module; the image quality assessment method further includes: respectively sampling at least one salient region from the at least two target images through the salient region sampling module; where the respectively extracting quality feature values corresponding to the at least one quality index of the at least two target images through the quality feature extraction module includes: extracting quality feature values corresponding to the at least one quality index from the at least one salient region through the quality feature extraction module.

[0009] Optionally, the quality assessment target answer includes a quality comparison result for the at least two target images and a causal reasoning description for the quality comparison result.

[0010] Optionally, the at least one quality metric includes at least one of the following: chroma, contrast, saturation, brightness, noise, Gaussian blur, compression.

[0011] Optionally, the image quality assessment model is trained by the following training method: obtaining training image samples, where the training image samples include at least two training images, a quality problem description for the at least two training images, and corresponding quality assessment answer labels; extracting text embeddings corresponding to the quality problem descriptions through the text compiler; respectively extracting visual embeddings of the at least two training images through the visual encoder; respectively extracting quality embeddings corresponding to at least one quality metric of the at least two training images through the quality enhancement module; splicing the text embeddings, the visual embeddings of the at least two training images, and the quality embeddings to obtain a spliced embedding; predicting, by the large language model, a quality assessment answer for the at least two training images based on the spliced embedding; and adjusting parameters of the large language model based on the quality assessment answer labels and the quality assessment answers to train the image quality assessment model.

[0012] Optionally, the splicing the text embeddings, the visual embeddings of the at least two training images, and the quality embeddings to obtain a spliced embedding includes: performing alignment processing on the quality embeddings of the at least two training images through a quality adapter to obtain aligned quality embeddings aligned with the text embeddings; splicing the text embeddings, the visual embeddings of the at least two training images, and the aligned quality embeddings to obtain a spliced embedding; where the adjusting parameters of the large language model includes: adjusting parameters of the large language model and the quality adapter.

[0013] Optionally, the respectively extracting quality embeddings corresponding to at least one quality metric of the at least two training images through the quality enhancement module includes: respectively extracting quality feature values corresponding to the at least one quality metric of the at least two training images through a quality feature extraction module, and generating the quality embeddings based on the extracted quality feature values corresponding to the at least one quality metric.

[0014] Optionally, the training method further includes: sampling at least one salient region from each of the at least two training images through a salient region sampling module; wherein, the extracting, by the quality feature extraction module, the quality feature values corresponding to the at least one quality metric for the at least two training images respectively includes: extracting, by the quality feature extraction module, the quality feature values corresponding to the at least one quality metric from the at least one salient region.

[0015] Optionally, the quality assessment answer includes a quality comparison result for the at least two training images and a causal reasoning description for the quality comparison result.

[0016] According to a second aspect of the embodiments of the present disclosure, there is provided an image quality assessment device, which is implemented based on an image quality assessment model. The image quality assessment model includes a text compiler, a vision encoder, and a large language model. The image quality assessment model further includes a quality enhancement module. The image quality assessment device includes: a target image acquisition module configured to acquire at least two target images to be subject to quality assessment, wherein the at least two target images correspond to a target quality problem description for the at least two target images; a text embedding extraction module configured to extract a target text embedding corresponding to the target quality problem description through the text compiler; a vision embedding extraction module configured to extract target vision embeddings of the at least two target images respectively through the vision encoder; a quality embedding extraction module configured to extract target quality embeddings corresponding to at least one quality metric for the at least two target images respectively through the quality enhancement module; a splicing module configured to splice the target text embedding, the target vision embeddings of the at least two target images, and the target quality embeddings to obtain a target spliced embedding; and a prediction module configured to predict a quality assessment target answer for the at least two target images based on the target spliced embedding through the large language model.

[0017] Optionally, the quality enhancement module includes a quality adapter; the splicing module is configured to: perform alignment processing on the target quality embeddings of the at least two target images through the quality adapter to obtain aligned target quality embeddings aligned with the target text embedding; and splice the target text embedding, the target vision embeddings of the at least two target images, and the aligned target quality embeddings to obtain the target spliced embedding.

[0018] Optionally, the quality enhancement module includes a quality feature extraction module; the quality embedding extraction module is configured to: respectively extract, through the quality feature extraction module, quality feature values corresponding to the at least one quality metric of the at least two target images, and generate the target quality embedding based on the extracted quality feature values corresponding to the at least one quality metric.

[0019] Optionally, the quality enhancement module further includes a saliency region sampling module; the quality embedding extraction module is configured to: respectively sample at least one saliency region from the at least two target images through the saliency region sampling module; extract, through the quality feature extraction module, quality feature values corresponding to the at least one quality metric from the at least one saliency region.

[0020] Optionally, the quality assessment target answer includes a quality comparison result for the at least two target images and a causal reasoning description for the quality comparison result.

[0021] Optionally, the at least one quality metric includes at least one of the following items: chroma, contrast, saturation, brightness, noise, Gaussian blur, compression.

[0022] Optionally, the image quality assessment model is trained by the following training method: obtaining training image samples, where the training image samples include at least two training images, a quality problem description for the at least two training images, and corresponding quality assessment answer labels; extracting, through the text compiler, text embeddings corresponding to the quality problem descriptions; respectively extracting, through the visual encoder, visual embeddings of the at least two training images; respectively extracting, through the quality enhancement module, quality embeddings corresponding to at least one quality metric of the at least two training images; splicing the text embeddings, the visual embeddings of the at least two training images, and the quality embeddings to obtain a spliced embedding; predicting, through the large language model, a quality assessment answer for the at least two training images based on the spliced embedding; and adjusting parameters of the large language model based on the quality assessment answer labels and the quality assessment answer to train the image quality assessment model.

[0023] Optionally, the concatenating the text embedding, the visual embeddings of the at least two training images, and the quality embedding to obtain a concatenated embedding includes: aligning the quality embedding of the at least two training images through a quality adapter to obtain an aligned quality embedding aligned with the text embedding; concatenating the text embedding, the visual embeddings of the at least two training images, and the aligned quality embedding to obtain a concatenated embedding; wherein, the adjusting the parameters of the large language model includes: adjusting the parameters of the large language model and the quality adapter.

[0024] Optionally, the extracting, by the quality enhancement module, the quality embeddings corresponding to at least one quality metric from the at least two training images respectively includes: extracting, by a quality feature extraction module, the quality feature values corresponding to the at least one quality metric from the at least two training images respectively, and generating the quality embedding based on the extracted quality feature values corresponding to the at least one quality metric.

[0025] Optionally, the training method further includes: sampling at least one salient region from the at least two training images respectively through a salient region sampling module; wherein, the extracting, by the quality feature extraction module, the quality feature values corresponding to the at least one quality metric from the at least two training images respectively includes: extracting, by the quality feature extraction module, the quality feature values corresponding to the at least one quality metric from the at least one salient region.

[0026] Optionally, the quality assessment answer includes a quality comparison result for the at least two training images and a causal reasoning description for the quality comparison result.

[0027] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the image quality assessment method according to the present disclosure.

[0028] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the image quality assessment method according to the present disclosure.

[0029] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the image quality assessment method according to the present disclosure.

[0030] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In the present disclosure, based on the existing image quality assessment model, a quality enhancement module is additionally introduced to extract relatively low-level quality features, and the quality features and the original features of the model are incorporated into the evaluation considerations of the image quality assessment model together, so that when comparing the quality of pictures, the quality features of the images can be fully considered. Even when faced with some images with relatively similar content, the high and low of the image quality can be accurately distinguished, thereby improving the ability in the fine-grained image quality comparison task.

[0031] According to an exemplary embodiment of the present disclosure, by setting a saliency region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when humans compare image quality, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0032] According to an exemplary embodiment of the present disclosure, by performing alignment processing on quality embeddings based on text embeddings, various types of embeddings can be made more regular with each other, and thus the splicing efficiency can be improved.

[0033] According to an exemplary embodiment of the present disclosure, by setting that the quality assessment answer includes the quality comparison result and the causal reasoning description for the quality comparison result, the image quality assessment model can not only predict the high and low of the quality between images, but also infer the reasons for the high and low differences in image quality, that is, the function of the image quality assessment model can be made more abundant, so as to provide more comprehensive services for users.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0035] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0036] Figure 1 is a schematic diagram showing the DepictQA framework in the related art; Figure 2 is a schematic diagram showing a MLLM-based IQA framework according to an exemplary embodiment of the present disclosure; Figure 3 is an example diagram showing the sampled saliency region according to an exemplary embodiment of the present disclosure; Figure 4 is a flowchart showing the image quality assessment method according to an exemplary embodiment of the present disclosure; Figure 5is a schematic diagram showing the Eye Moments Experiments of Alfred L. Yarbus according to an exemplary embodiment of the present disclosure; Figure 6 is a structural schematic diagram showing a naive splicing method according to an exemplary embodiment of the present disclosure; Figure 7 is a schematic diagram showing feature integration based on Quality Projector and Visual projector according to an exemplary embodiment of the present disclosure; Figure 8 is a flowchart showing a training method of an image quality assessment model according to an exemplary embodiment of the present disclosure; Figure 9 is a block diagram showing an image quality assessment device according to an exemplary embodiment of the present disclosure; Figure 10 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0037] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0039] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of three parallel situations: "any one of the several items", "any combination of several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of step one and step two" means the following three parallel situations: (1) executing step one; (2) executing step two; (3) executing step one and step two.

[0040] Most traditional Image Quality Assessment (IQA) models are optimized for a learnable network with a regression task as the goal based on a dataset containing the Mean Opinion Score (MOS) of humans, so that the model can better fit the MOS score. However, such models can only make binary judgments on image quality and cannot obtain more detailed descriptions or summaries of image quality. Additionally, thanks to the scaling law of the Transformers architecture and the development of the computing power of Graphics Processing Units (GPUs), Multimodal Large Language Models (MLLMs) are emerging, and they have shown amazing performance in advanced Visual Question Answering (VQA) tasks. Researchers have tried to apply multimodal large language models to the IQA field. The general technical route of such models is as follows: First, prepare a large-scale dataset containing image quality descriptions. Then, perform efficient fine-tuning based on a general-domain multimodal large model. For example, the multimodal large model DepictQA for the IQA task mainly fine-tunes the general multimodal large model LLava using the large-scale dataset DQ-495K and the Lora efficient fine-tuning technique.

[0041] Figure 1 is a schematic diagram showing the DepictQA framework in the related art. Refer to Figure 1 , the Llava model used by DepictQA usually uses the Image Encoder of CLIP as its image feature extractor. Additionally, an adapter (Projector) based on a Multilayer Perceptron (MLP) can be connected in series after the Image Encoder. During the model fine-tuning process, the parameters of the Image Encoder can be frozen, that is, they can be not involved in the update; the Projector and the subsequent Large Language Model (LLM) can participate in the fine-tuning.

[0042] Exemplarily, the input of the DepictQA framework can be Image A, Image B, and the question for these two images: "Compare the overall quality of Image A and Image B and provide a comprehensive explanation"; the output of the DepictQA framework can be: "Although Image A is not as good as Image B in terms of noise, it is far superior to Image B in terms of blurriness. These two images show similar performance in terms of brightness distortion, color distortion, and artifacts. In terms of texture quality, the plant texture in Image A is clearly visible, while the plant texture in Image B is completely damaged and unrecognizable. This is mainly due to the advantage of blurriness in Image A, that is, it allows the retention of texture. Therefore, the quality of Image A is significantly better than that of Image B."

[0043] However, as mentioned above, for some images with relatively similar content (also known as fine-grained images), the feature similarities extracted by the Image Encoder of CLIP will be relatively high, which will make it difficult for subsequent large language models to accurately distinguish the quality of these images with relatively similar content, that is, the related technologies are difficult to accurately compare and evaluate the quality of fine-grained images.

[0044] To solve the above problems existing in the related technologies, the image quality assessment method, device, electronic device, storage medium, and computer program product provided by the present disclosure, based on the existing image quality assessment model, additionally introduce a quality enhancement module to extract relatively low-level quality features, and incorporate the quality features and the original features of the model into the evaluation considerations of the image quality assessment model, so that when comparing the quality of pictures, the quality features of the images can be fully considered, and even when facing some images with relatively similar content, the quality of the images can be more accurately distinguished, thereby improving the ability in the fine-grained image quality comparison task.

[0045] Figure 2 is a schematic diagram showing an MLLM-based IQA framework according to an exemplary embodiment of the present disclosure. Refer to Figure 2, the existing MLLM-based IQA frameworks generally can include a text compiler (Tokenize), a visual encoder (Visualencoder), and a large language model. Among them, the text compiler is mainly used to convert the content in text form into vector form, and the vector can contain multiple numerical elements; the visual encoder is mainly used to extract the semantic features of the image, and the semantic features can also be called content features or visual features; the large language model is mainly used to compare and evaluate the image quality based on the output results of the previous module. However, as mentioned above, limited by the characteristic that the visual encoder can only obtain the semantic features of the image well but is difficult to obtain the relatively low-level image quality features, the above-mentioned MLLM-based IQA frameworks in the related technologies cannot accurately compare and evaluate the quality of fine-grained images.

[0046] To solve the above problems existing in the related technologies, based on the existing MLLM-based IQA framework, the present disclosure adds a quality enhancement module, which is mainly used to extract the quality embeddings corresponding to at least one quality metric of the image. The quality embeddings can be a vector containing multiple numerical elements, and the quality embeddings can be used to characterize at least one quality characteristic of the image. In this way, a quality enhancement module can be additionally introduced to extract relatively low-level quality features, and then the quality features and the original features of the model can be incorporated into the evaluation considerations of the image quality evaluation model together, so that the quality features of the image can be fully considered when comparing the quality of pictures. Even when facing some images with relatively similar content, the high and low of the image quality can be distinguished more accurately.

[0047] According to an exemplary embodiment of the present disclosure, a quality feature extraction module can also be set in the above quality enhancement module to extract the above quality embeddings. For this quality feature extraction module, ResNet50 can be used as the encoder for quality feature extraction, and contrastive learning frameworks such as SimCLR can be used for self-supervised pre-training to enable the encoder to learn rich image quality features. In addition, a linear layer can be connected after the quality feature extraction module, and supervised regression MOS can be performed on the IQA dataset.

[0048] According to an exemplary embodiment of the present disclosure, the quality enhancement module provided by the present disclosure may further include a Salient Regions Sampling Module (SRSM). The SRSM may adopt, but is not limited to, Image Signature as a saliency detection sub-module. Specifically, first, Image Signature may be used to perform saliency detection on the image to be sampled to obtain the approximate foreground position of the image, which may then be used as the human eye movement fixation point. Then, a threshold filtering method may be used to select the highlighted regions in the image and crop them. Figure 3 is an example diagram showing the sampled salient regions according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , in the leftmost image, the three small regions enclosed by the solid line box are the detected salient regions. Moreover, in addition to outputting the detected salient regions, the Salient Regions Sampling Module will also output the complete image originally input to the Salient Regions Sampling Module. As Figure 3 shown, Figure 3 the three detected salient regions and the complete image A originally input to the Salient Regions Sampling Module are sequentially shown in the upper right part; Figure 3 the three detected salient regions and the complete image B originally input to the Salient Regions Sampling Module are sequentially shown in the lower right part.

[0049] After the Salient Regions Sampling Module samples the salient regions from the image, the sampled salient regions may be used as input and passed into the quality feature extraction module, that is, passed into the encoder, and the output of the last layer of the encoder may be used as the quality feature. In this way, by setting the Salient Regions Sampling Module, it is possible to accurately sample the regions that are most likely to be concerned when comparing the image quality of humans, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0050] According to an exemplary embodiment of the present disclosure, the quality enhancement module provided by the present disclosure may further include a quality adapter, which is mainly used to perform alignment processing on the quality embeddings output by the quality feature extraction module so that the quality embeddings after alignment are aligned with the text embeddings output by the text compiler. In this way, by performing alignment processing on the quality embeddings based on the text embeddings, it is possible to make various types of embeddings more regular with each other, thereby improving the splicing efficiency.

[0051] According to an exemplary embodiment of the present disclosure, the quality enhancement module provided by the present disclosure may further include a splicing module. The splicing module is mainly used to splice the text embedding output by the text compiler, the visual embedding output by the visual encoder, and the quality embedding output by the quality feature extraction module, so that the subsequent LLM can perform comparative evaluation of image quality based on the spliced embedding.

[0052] It should be noted that in the present disclosure, in addition to setting a splicing module in the quality enhancement module to splice various types of embeddings, the splicing module can also be set outside the quality enhancement module as an independent module to splice various embeddings. For example, the splicing module included in the existing MLLM-based IQA framework can also be used to splice and integrate various types of embeddings, and the present disclosure does not make specific limitations on this.

[0053] Figure 4 is a flowchart showing an image quality assessment method according to an exemplary embodiment of the present disclosure. The image quality assessment method can be implemented based on an image quality assessment model, and the image quality assessment model can include a text compiler, a visual encoder, a large language model, and a quality enhancement module.

[0054] Refer to Figure 4 , in step 401, at least two target images to be quality-assessed can be obtained, where the at least two target images can correspond to target quality problem descriptions for the at least two target images. Exemplarily, the user can manually input the target quality problem descriptions for the at least two target images.

[0055] Exemplarily, as Figure 2 shown, the at least two target images can be image A and image B respectively; the target quality problem descriptions for these two target images can be: "Which image has better quality? Image A or image B? Tell me why".

[0056] In step 402, the target text embedding corresponding to the above target quality problem description can be extracted through the text compiler. That is, the above target quality problem description can be input into the text compiler included in the trained image quality assessment model to obtain the target text embedding corresponding to the target quality problem description (Textual embedding), and the target text embedding can be a vector containing multiple digital elements.

[0057] Exemplarily, as Figure 2 shown, the target quality problem description: "Which image has better quality? Image A or image B? Tell me why" can be input into the text compiler, and then the target text embedding corresponding to the target quality problem description can be obtained.

[0058] In step 403, the target visual embeddings of the at least two target images can be respectively extracted by a visual encoder. That is, the at least two target images can be respectively input into the visual encoder included in the image quality assessment model to obtain the target visual embeddings of the at least two target images (Visual embedding), and the target visual embedding can be a vector containing multiple numerical elements. Exemplarily, as Figure 2 shown, image A and image B can be respectively input into the visual encoder, and then the target visual embedding of image A and the target visual embedding of image B can be obtained.

[0059] In step 404, the target quality embeddings corresponding to at least one quality metric of the at least two target images can be respectively extracted by a quality enhancement module. That is, the at least two target images can be respectively input into the quality enhancement module included in the image quality assessment model to obtain the target quality embeddings corresponding to at least one quality metric of the at least two target images (Quality embedding), and the target quality embedding can be a vector containing multiple numerical elements, and the target quality embedding can be used to characterize at least one quality characteristic of the image.

[0060] Exemplarily, as Figure 2 shown, image A and image B can be respectively input into the quality enhancement module, and then the target quality embedding of image A and the target quality embedding of image B can be obtained.

[0061] According to an exemplary embodiment of the present disclosure, the at least one quality metric may include at least one of the following items: chromaticity, contrast, saturation, brightness, noise, Gaussian blur, compression.

[0062] According to an exemplary embodiment of the present disclosure, the quality enhancement module may include a quality feature extraction module. The quality feature values corresponding to at least one quality metric of the at least two target images can be respectively extracted by the quality feature extraction module, and the target quality embedding can be generated based on the extracted quality feature values corresponding to at least one quality metric.

[0063] It should be noted that based on the human vision and eye movement mechanism, that is, the characteristic that the region of interest (RegionOf Interest, ROI) of the human eye observing an image depends on a preset prompt or question, the present disclosure also proposes a saliency region sampling method. Figure 5 is a schematic diagram showing the eye movement experiment (Eye MomentsExperiments) of Alfred L. Yarbus according to an exemplary embodiment of the present disclosure.

[0064] Referring to Figure 5, the first small image "free examination" indicates that if people are not prompted in advance about what to pay attention to, their eyes will wander randomly when observing the image; the remaining three small images indicate that if people are prompted in advance about what to pay attention to, their eyes will follow the pre-prompted tasks. For example, the second small image prompts "Remember the clothes people are wearing", then people will focus on the clothing of the people in the large image on the left when browsing; the third small image prompts "Remember the positions of people and objects", then people will focus on the positions of the people and objects in the large image on the left when browsing; the fourth small image prompts "Remember the ages of people", then people will focus on the ages of the people in the large image on the left when browsing.

[0065] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a saliency region sampling module. At least one saliency region can be sampled from the above at least two target images through this saliency region sampling module. Exemplarily, as Figure 2 shown, image A and image B can be respectively input into the saliency region sampling module, and then the saliency region sampled from image A and the saliency region sampled from image B can be obtained. Then, at least one quality feature value corresponding to at least one quality metric can be extracted from the at least one saliency region through the quality feature extraction module.

[0066] Next, a target quality embedding can be generated based on the at least one quality feature value corresponding to at least one quality metric extracted from the above at least one saliency region. Exemplarily, as Figure 2 shown, the saliency region sampled from image A and the saliency region sampled from image B can be respectively input into the quality feature extraction module, and then the target quality embedding of image A and the target quality embedding of image B can be obtained.

[0067] In this way, by setting the saliency region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when comparing the image quality of humans, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0068] Returning to Figure 4 , in step 405, the foregoing target text embedding, the target visual embeddings of at least two target images, and the target quality embedding can be concatenated to obtain a target concatenated embedding.

[0069] According to the exemplary embodiments of the present disclosure, two ways of integrating quality features can be provided to integrate the quality features extracted by the quality feature extraction module, the text features extracted by the text compiler, and the visual features extracted by the CLIPImage Encoder within the original DepictQA framework. These two ways of integrating quality features can be respectively: the Pure Concatenation method and the integration method based on Quality Projector. Specifically: 1. Pure Concatenation Since the transformers structure of the large language model is not sensitive to the input Token Embedding in the Token Length dimension, the above three features can be directly concatenated and fused by using the tensor concatenation method. Figure 6 FIG. is a schematic structural diagram showing the Pure Concatenation method according to the exemplary embodiments of the present disclosure.

[0070] Referring to Figure 6 , after the target quality problem description is input into the text compiler 601, the target text embedding can be obtained; after the target image is input into the visual encoder 602 of CLIP, the target visual embedding can be obtained. Exemplarily, the visual encoder 602 of CLIP can be, but is not limited to, ViT, ResNet, etc.; after the target image is input into the quality enhancement module 603, the target quality embedding can be obtained. Next, the splicing module 604 can directly splice the above-mentioned target text embedding, target visual embedding, and target quality embedding to obtain the target spliced embedding. Finally, the features with integrated quality features can be input into the subsequent large language model 605 to obtain the response output. In addition, although Figure 6 shows that the splicing module 604 is outside the quality enhancement module 603, this is only an example. For example, the splicing module 604 can also be a module within the quality enhancement module 603.

[0071] In this way, by integrating the quality features with the features of the original MLLM, it is possible to fully consider the features of the image in various aspects, thereby improving the model's ability in the fine-grained image quality comparison task.

[0072] 2. Integration based on Quality Projector It should be noted that although the features extracted by the MOS regression-based IQA model have quality discrimination ability, they are not aligned with natural language. Therefore, as Figure 2As shown, the present disclosure also provides a feature integration method based on QualityProjector. The quality adapter is mainly used to perform alignment processing on the target quality embedding output by the quality feature extraction module, so that the target quality embedding after alignment processing is aligned with the target text embedding output by the text compiler.

[0073] Specifically, the target quality embedding extracted by the quality feature extraction module will be input into the Quality Projector based on a multi-layer perceptron for alignment processing, and this Quality Projector can participate in the learning process of the entire model. That is, during the training process of the image quality evaluation model, in addition to adjusting the parameters of the LLM, the parameters of this Quality Projector can also be adjusted. The above "alignment processing" can refer to projecting two vectors into a higher-dimensional space and making their distances as similar as possible. Then, the output result of the Quality Projector, the visual features extracted by the CLIP ImageEncoder, and the target text embedding extracted by the text compiler can be concatenated in the Token Length dimension.

[0074] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a quality adapter (QualityProjector). The target quality embeddings of at least two target images can be aligned through this quality adapter to obtain aligned target quality embeddings aligned with the target text embedding. Then, the target text embedding, the target visual embeddings of at least two target images, and the aligned target quality embeddings can be concatenated to obtain target concatenated embeddings. Figure 7 It is a schematic diagram showing feature integration based on Quality Projector and Visual projector according to an exemplary embodiment of the present disclosure.

[0075] Referring to Figure 7 , after the target quality problem description is input into the text compiler 701, the target text embedding can be obtained; after the target image is input into the visual encoder 702 of CLIP, the target visual embedding can be obtained, and this target visual embedding can also be input into the visual adapter 703 for alignment processing so that the target visual embedding after alignment processing is aligned with the target text embedding; after the target image is input into the quality feature extraction module 7041 included in the quality enhancement module 704, the target quality embedding can be obtained, and this target quality embedding can also be input into the quality adapter 7042 included in the quality enhancement module 704 for alignment processing so that the target quality embedding after alignment processing is aligned with the target text embedding.

[0076] Next, the splicing module 705 can be used to splice the aforementioned target text embedding, the aligned target visual embedding after alignment processing, and the aligned target quality embedding after alignment processing to obtain a target spliced embedding. Finally, the feature with integrated quality features can be input into the subsequent large language model 706 to obtain a response output. It should be noted that although Figure 7 shows that the splicing module 705 is outside the quality enhancement module 704, this is only an example. For example, the splicing module 705 can also be a module in the quality enhancement module 704. In this way, by performing alignment processing on the target quality embedding based on the target text embedding, various types of embeddings can be made more regular with each other, thereby improving the splicing efficiency.

[0077] Return to reference Figure 4 , in step 406, the large language model can be used to predict a quality assessment target answer regarding at least two target images based on the target spliced embedding. Exemplarily, as Figure 2 shown, by inputting the target spliced embedding into the large language model, the quality assessment target answers for image A and image B can be obtained. For example, the quality assessment target answer can be: The quality of image A is higher than the quality of image B.

[0078] According to an exemplary embodiment of the present disclosure, the above quality assessment target answer may include a quality comparison result for at least two target images and a causal reasoning description for the quality comparison result. In this way, in the present disclosure, by setting the quality assessment answer to include the quality comparison result and the causal reasoning description for the quality comparison result, the image quality assessment model can not only predict the quality level between images, but also infer the reason for the high or low quality difference between images, that is, the function of the image quality assessment model can be made more abundant, thereby providing more comprehensive services for users.

[0079] Figure 8 is a flowchart showing a training method of an image quality assessment model according to an exemplary embodiment of the present disclosure.

[0080] Refer to Figure 8 , in step 801, training image samples can be obtained, where the training image samples may include at least two training images, a quality problem description for at least two training images, and corresponding quality assessment answer labels. Exemplarily, as Figure 2 shown, the training image samples may include two training images, namely image A and image B; the quality problem description for these two training images can be: "Which image has better quality? Image A or image B? Tell me why"; the quality assessment answer label for these two training images can be: "The quality of image A is higher than the quality of image B".

[0081] In step 802, the text embedding corresponding to the above quality problem description can be extracted through a text compiler. That is, the above quality problem description can be input into the text compiler to obtain the text embedding corresponding to the quality problem description, and the text embedding can be a vector containing multiple numerical elements. Exemplarily, as Figure 2 shown, the quality problem description: "Which image has better quality? Image A or Image B? Tell me why" can be input into the text compiler, and then the text embedding corresponding to the quality problem description can be obtained.

[0082] In step 803, the visual embeddings of at least two training images can be extracted respectively through a visual encoder. That is, the above at least two training images can be input into the visual encoder respectively to obtain the visual embeddings of at least two training images, and the visual embedding can be a vector containing multiple numerical elements. Exemplarily, as Figure 2 shown, Image A and Image B can be input into the visual encoder respectively, and then the visual embedding of Image A and the visual embedding of Image B can be obtained.

[0083] In step 804, the quality embeddings corresponding to at least one quality metric of at least two training images can be extracted respectively through a quality enhancement module. That is, the above at least two training images can be input into the quality enhancement module respectively to obtain the quality embeddings corresponding to at least one quality metric of at least two training images, and the quality embedding can be a vector containing multiple numerical elements, and the quality embedding can be used to characterize at least one quality characteristic of the image. Exemplarily, as Figure 2 shown, Image A and Image B can be input into the quality enhancement module respectively, and then the quality embedding of Image A and the quality embedding of Image B can be obtained.

[0084] According to an exemplary embodiment of the present disclosure, the above at least one quality metric may include at least one of the following items: chromaticity, contrast, saturation, brightness, noise, Gaussian blur, compression.

[0085] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a quality feature extraction module. The quality feature values corresponding to at least one quality metric of at least two training images can be extracted respectively through the quality feature extraction module, and the quality embedding can be generated based on the extracted quality feature values corresponding to at least one quality metric.

[0086] According to an exemplary embodiment of the present disclosure, the above-mentioned quality enhancement module may further include a saliency region sampling module. At least one saliency region can be sampled from the above-mentioned at least two training images through the saliency region sampling module. Then, at least one quality feature value corresponding to at least one quality metric can be extracted from the above-mentioned at least one saliency region through the aforementioned quality feature extraction module. Next, a quality embedding can be generated based on the at least one quality feature value corresponding to at least one quality metric extracted from the above-mentioned at least one saliency region.

[0087] In this way, by setting the saliency region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when humans compare image quality, reduce redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0088] In step 805, the aforementioned text embedding, the visual embeddings of at least two training images, and the quality embedding can be concatenated to obtain a concatenated embedding.

[0089] According to an exemplary embodiment of the present disclosure, the above-mentioned quality enhancement module may further include a quality adapter. The quality embeddings of at least two training images can be aligned through the quality adapter to obtain an aligned quality embedding aligned with the text embedding. Then, the text embedding, the visual embeddings of at least two training images, and the aligned quality embedding can be concatenated to obtain a concatenated embedding. Finally, the feature with integrated quality features can be input into the subsequent large language model to obtain a response output. In this way, by aligning the quality embedding based on the text embedding, the various types of embeddings can be made more regular with each other, thereby improving the concatenation efficiency.

[0090] In step 806, the large language model can be used to predict a quality assessment answer regarding at least two training images based on the above-mentioned concatenated embedding. Exemplarily, as Figure 2 shown, by inputting the concatenated embedding into the large language model, a quality assessment answer for Image A and Image B can be obtained. For example, the quality assessment answer can be: The quality of Image A is higher than the quality of Image B.

[0091] In step 807, the parameters of the large language model can be adjusted based on the quality assessment answer label and the quality assessment answer to train the image quality assessment model.

[0092] According to an exemplary embodiment of the present disclosure, the above quality assessment answer may include a quality comparison result for at least two training images and a causal reasoning description for the quality comparison result. In this way, in the present disclosure, by setting the quality assessment answer to include the quality comparison result and the causal reasoning description for the quality comparison result, the image quality assessment model can not only predict the quality level between images, but also infer the reasons for the high or low quality differences between images, that is, the function of the image quality assessment model can be made more abundant, so as to provide more comprehensive services for users.

[0093] Figure 9 FIG. 4 is a block diagram showing an image quality assessment device 900 according to an exemplary embodiment of the present disclosure. The image quality assessment device 900 may be implemented based on an image quality assessment model, and the image quality assessment model may include a text compiler, a vision encoder, a large language model, and a quality enhancement module.

[0094] Referring to Figure 9 , the image quality assessment device 900 may include a target image acquisition module 901, a text embedding extraction module 902, a vision embedding extraction module 903, a quality embedding extraction module 904, a splicing module 905, and a prediction module 906.

[0095] The target image acquisition module 901 may acquire at least two target images to be subject to quality assessment, wherein the at least two target images may correspond to a target quality problem description for the at least two target images. Exemplarily, the user may input the target quality problem description for the at least two target images manually.

[0096] The text embedding extraction module 902 may extract a target text embedding corresponding to the above target quality problem description through the text compiler. That is, the above target quality problem description may be input into the text compiler included in the trained image quality assessment model to obtain a target text embedding corresponding to the target quality problem description, and the target text embedding may be a vector including a plurality of digital elements.

[0097] The vision embedding extraction module 903 may extract target vision embeddings of the at least two target images respectively through the vision encoder. That is, the at least two target images may be respectively input into the vision encoder included in the image quality assessment model to obtain target vision embeddings of the at least two target images, and the target vision embedding may be a vector including a plurality of digital elements.

[0098] The quality embedding extraction module 904 can respectively extract the target quality embeddings corresponding to at least one quality metric of the above-mentioned at least two target images through the quality enhancement module. That is, at least two target images can be respectively input into the quality enhancement module included in the image quality assessment model to obtain the target quality embeddings corresponding to at least one quality metric of the at least two target images. The target quality embedding can be a vector containing multiple numerical elements, and the target quality embedding can be used to characterize at least one quality characteristic of the image.

[0099] According to an exemplary embodiment of the present disclosure, the above-mentioned at least one quality metric may include at least one of the following items: chromaticity, contrast, saturation, brightness, noise, Gaussian blur, compression.

[0100] According to an exemplary embodiment of the present disclosure, the above-mentioned quality enhancement module may include a quality feature extraction module. The quality embedding extraction module 904 can respectively extract the quality feature values corresponding to at least one quality metric of the above-mentioned at least two target images through the quality feature extraction module, and can generate a target quality embedding based on the extracted quality feature values corresponding to at least one quality metric.

[0101] According to an exemplary embodiment of the present disclosure, the above-mentioned quality enhancement module may further include a saliency region sampling module. The quality embedding extraction module 904 can respectively sample at least one saliency region from the above-mentioned at least two target images through the saliency region sampling module. Then, the quality embedding extraction module 904 can extract the quality feature values corresponding to at least one quality metric from the at least one saliency region through the quality feature extraction module. Next, the quality embedding extraction module 904 can generate a target quality embedding based on the quality feature values corresponding to at least one quality metric extracted from the above-mentioned at least one saliency region.

[0102] In this way, by setting the saliency region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when humans compare image quality, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0103] The splicing module 905 can splice the foregoing target text embedding, the target visual embeddings of at least two target images, and the target quality embeddings to obtain a target spliced embedding.

[0104] According to an exemplary embodiment of the present disclosure, the above-mentioned quality enhancement module may further include a quality adapter. The splicing module 905 may perform alignment processing on the target quality embeddings of at least two target images through the quality adapter to obtain aligned target quality embeddings aligned with the target text embeddings. Then, the splicing module 905 may splice the target text embeddings, the target visual embeddings of at least two target images, and the aligned target quality embeddings to obtain target spliced embeddings.

[0105] In this way, by performing alignment processing on the target quality embeddings based on the target text embeddings, various types of embeddings can be made more regular with each other, thereby improving the splicing efficiency.

[0106] The prediction module 906 may predict a quality assessment target answer regarding at least two target images based on the target spliced embeddings through a large language model.

[0107] According to an exemplary embodiment of the present disclosure, the above-mentioned quality assessment target answer may include a quality comparison result for at least two target images and a causal reasoning description for the quality comparison result. In this way, in the present disclosure, by setting the quality assessment answer to include the quality comparison result and the causal reasoning description for the quality comparison result, the image quality assessment model can not only predict the quality level between images, but also infer the reasons for the high or low quality differences between images, that is, the function of the image quality assessment model can be made more abundant, so as to provide more comprehensive services for users.

[0108] According to an exemplary embodiment of the present disclosure, the image quality assessment model in the present disclosure may be trained through the following training method: First, training image samples may be obtained, where the training image samples may include at least two training images, a quality problem description for at least two training images, and corresponding quality assessment answer labels.

[0109] Then, the text embeddings corresponding to the above-mentioned quality problem descriptions may be extracted through a text compiler. That is, the above-mentioned quality problem descriptions may be input into the text compiler to obtain text embeddings corresponding to the quality problem descriptions, and the text embeddings may be a vector including multiple digital elements.

[0110] Next, the visual embeddings of at least two training images may be extracted respectively through a visual encoder. That is, the above-mentioned at least two training images may be input into the visual encoder respectively to obtain visual embeddings of at least two training images, and the visual embeddings may be a vector including multiple digital elements.

[0111] Then, the quality embeddings corresponding to at least one quality metric of at least two training images can be separately extracted through the quality enhancement module. That is, the aforementioned at least two training images can be separately input into the quality enhancement module to obtain the quality embeddings corresponding to at least one quality metric of the at least two training images. The quality embedding can be a vector containing multiple numerical elements, and the quality embedding can be used to characterize at least one quality characteristic of the image.

[0112] Next, the aforementioned text embedding, the visual embeddings of at least two training images, and the quality embeddings can be concatenated to obtain a concatenated embedding.

[0113] Then, based on the above concatenated embedding, the large language model can be used to predict the quality assessment answers regarding at least two training images.

[0114] Next, based on the quality assessment answer labels and the quality assessment answers, the parameters of the large language model can be adjusted to train the image quality assessment model.

[0115] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a quality feature extraction module. The quality feature values corresponding to at least one quality metric of at least two training images can be separately extracted through the quality feature extraction module, and the quality embedding can be generated based on the extracted quality feature values corresponding to at least one quality metric.

[0116] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a salient region sampling module. At least one salient region can be separately sampled from the above at least two training images through the salient region sampling module. Then, the quality feature values corresponding to at least one quality metric can be extracted from the above at least one salient region through the aforementioned quality feature extraction module. Next, the quality embedding can be generated based on the quality feature values corresponding to at least one quality metric extracted from the above at least one salient region.

[0117] In this way, by setting the salient region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when humans compare image quality, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0118] According to an exemplary embodiment of the present disclosure, the above quality enhancement module may further include a quality adapter. The quality embeddings of at least two training images can be aligned through the quality adapter to obtain aligned quality embeddings aligned with the text embeddings. Then, the text embeddings, the visual embeddings of at least two training images, and the aligned quality embeddings can be concatenated to obtain concatenated embeddings. Finally, the features with integrated quality features can be input into the subsequent large language model to obtain a response output. In this way, by aligning the quality embeddings based on the text embeddings, various types of embeddings can be made more regular with each other, thereby improving the concatenation efficiency.

[0119] According to an exemplary embodiment of the present disclosure, the above quality assessment answer may include a quality comparison result for at least two training images and a causal reasoning description for the quality comparison result. In this way, in the present disclosure, by setting the quality assessment answer to include the quality comparison result and the causal reasoning description for the quality comparison result, the image quality assessment model can not only predict the quality level between images, but also infer the reasons for the high or low quality differences between images, that is, the function of the image quality assessment model can be made more abundant, so as to provide more comprehensive services for users.

[0120] Figure 10 is a block diagram showing an electronic device 1000 according to an exemplary embodiment of the present disclosure.

[0121] Referring to Figure 10 , the electronic device 1000 includes at least one memory 1001 and at least one processor 1002. Instructions are stored in the at least one memory 1001, and when the instructions are executed by the at least one processor 1002, an image quality assessment method according to an exemplary embodiment of the present disclosure is executed.

[0122] As an example, the electronic device 1000 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 1000 does not have to be a single electronic device, and may also be an aggregate of any devices or circuits capable of executing the above instructions (or instruction sets) alone or jointly. The electronic device 1000 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.

[0123] In the electronic device 1000, the processor 1002 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0124] The processor 1002 can execute instructions or code stored in the memory 1001, where the memory 1001 can also store data. The instructions and data can also be sent and received via the network interface device over a network, where the network interface device can employ any known transmission protocol.

[0125] The memory 1001 can be integrated with the processor 1002. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 1001 can include separate devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 1001 and the processor 1002 can be operatively coupled or can communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 1002 can read files stored in the memory.

[0126] In addition, the electronic device 1000 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 1000 can be connected to each other via a bus and / or a network.

[0127] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above image quality assessment method. Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0128] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program which, when executed by a processor, implements the image quality assessment method according to the present disclosure.

[0129] According to the image quality assessment method, apparatus, electronic device, storage medium, and computer program product of the present disclosure, by additionally introducing a quality enhancement module on the basis of the existing image quality assessment model, extracting relatively low-level quality features, and incorporating the quality features and the original features of the model into the evaluation considerations of the image quality assessment model together, when comparing the quality of pictures, the quality features of the images can be fully considered, and even when facing some images with relatively similar contents, the high and low of the image quality can be distinguished relatively accurately, thereby improving the ability in the fine-grained image quality comparison task.

[0130] According to an exemplary embodiment of the present disclosure, by setting a saliency region sampling module, it is possible to accurately sample the regions that are most likely to be concerned when comparing the image quality of humans, reduce the redundant sampling of irrelevant regions, and improve the sampling efficiency.

[0131] According to an exemplary embodiment of the present disclosure, by integrating the quality features with the features of the original MLLM, it is possible to fully consider the features of the image in various aspects, thereby improving the model's ability in the fine-grained image quality comparison task.

[0132] According to an exemplary embodiment of the present disclosure, by aligning the quality embedding based on the text embedding, it is possible to make various types of embeddings more regular with each other, and further improve the splicing efficiency.

[0133] According to an exemplary embodiment of the present disclosure, by setting the quality evaluation answer to include the quality comparison result and the causal reasoning description for the quality comparison result, the image quality evaluation model can not only predict the high and low quality between images, but also infer the reasons for the high and low differences in image quality, that is, the function of the image quality evaluation model can be made more abundant, so as to provide more comprehensive services for users.

[0134] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0135] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image quality assessment method, which is implemented based on an image quality assessment model. The image quality assessment model includes a text compiler, a visual encoder, and a large language model, and is characterized in that The image quality assessment model further includes a quality enhancement module, and the image quality assessment method includes: Obtain at least two target images to be quality-assessed, where the at least two target images are corresponding to a target quality problem description for the at least two target images; Extract a target text embedding corresponding to the target quality problem description through the text compiler; Extract target visual embeddings of the at least two target images respectively through the visual encoder; Extract target quality embeddings corresponding to at least one quality metric of the at least two target images respectively through the quality enhancement module; Concatenate the target text embedding, the target visual embeddings of the at least two target images, and the target quality embeddings to obtain a target concatenated embedding; Predict a quality assessment target answer for the at least two target images based on the target concatenated embedding through the large language model.

2. The image quality evaluation method according to claim 1, wherein The quality enhancement module includes a quality adapter; The step of concatenating the target text embedding, the target visual embeddings of the at least two target images, and the target quality embeddings to obtain a target concatenated embedding includes: Align the target quality embeddings of the at least two target images through the quality adapter to obtain aligned target quality embeddings aligned with the target text embedding; Concatenate the target text embedding, the target visual embeddings of the at least two target images, and the aligned target quality embeddings to obtain the target concatenated embedding.

3. The image quality assessment method according to claim 1, characterized in that The quality enhancement module includes a quality feature extraction module; The step of extracting target quality embeddings corresponding to at least one quality metric of the at least two target images respectively through the quality enhancement module includes: Extract quality feature values corresponding to the at least one quality metric of the at least two target images respectively through the quality feature extraction module, and generate the target quality embeddings based on the extracted quality feature values corresponding to the at least one quality metric.

4. The image quality assessment method according to claim 3, wherein The quality enhancement module further includes a salient region sampling module; The image quality assessment method further includes: Sample at least one salient region from the at least two target images respectively through the salient region sampling module; Among them, the step of extracting quality feature values corresponding to the at least one quality metric of the at least two target images respectively through the quality feature extraction module includes: Extract quality feature values corresponding to the at least one quality metric from the at least one salient region through the quality feature extraction module.

5. The image quality assessment method according to claim 1, wherein The quality assessment target answer includes a quality comparison result for the at least two target images and a causal reasoning description for the quality comparison result.

6. The image quality assessment method according to claim 1, wherein The at least one quality metric includes at least one of the following items: Chromaticity, contrast, saturation, brightness, noise, Gaussian blur, compression.

7. The image quality evaluation method according to any one of claims 1 to 6, characterized in that The image quality assessment model is trained through the following training method: Obtain training image samples, where the training image samples include at least two training images, a quality problem description for the at least two training images, and corresponding quality assessment answer labels; Extract the text embedding corresponding to the quality problem description through the text compiler; Extract the visual embeddings of the at least two training images respectively through the visual encoder; Extract the quality embeddings corresponding to at least one quality metric of the at least two training images respectively through the quality enhancement module; Concatenate the text embedding, the visual embeddings of the at least two training images, and the quality embeddings to obtain a concatenated embedding; Predict, through the large language model, a quality assessment answer regarding the at least two training images based on the concatenated embedding; Adjust the parameters of the large language model based on the quality assessment answer label and the quality assessment answer to train the image quality assessment model.

8. The image quality assessment method according to claim 7, wherein The step of concatenating the text embedding, the visual embeddings of the at least two training images, and the quality embeddings to obtain a concatenated embedding includes: Align the quality embeddings of the at least two training images through a quality adapter to obtain aligned quality embeddings aligned with the text embedding; Concatenate the text embedding, the visual embeddings of the at least two training images, and the aligned quality embeddings to obtain a concatenated embedding; Among them, the step of adjusting the parameters of the large language model includes: Adjust the parameters of the large language model and the quality adapter.

9. The image quality evaluation method according to claim 7, characterized in that The step of extracting the quality embeddings corresponding to at least one quality metric of the at least two training images respectively through the quality enhancement module includes: Extract the quality feature values corresponding to the at least one quality metric of the at least two training images respectively through a quality feature extraction module, and generate the quality embeddings based on the extracted quality feature values corresponding to the at least one quality metric.

10. The image quality assessment method according to claim 9, wherein, The training method further includes: Sample at least one salient region from the at least two training images respectively through a salient region sampling module; Among them, the step of extracting the quality feature values corresponding to the at least one quality metric of the at least two training images respectively through the quality feature extraction module includes: Extract the quality feature values corresponding to the at least one quality metric from the at least one salient region through the quality feature extraction module.

11. The image quality evaluation method according to claim 7, wherein The quality assessment answer includes a quality comparison result for the at least two training images and a causal reasoning description for the quality comparison result.

12. An image quality assessment device, which is implemented based on an image quality assessment model. The image quality assessment model includes a text compiler, a vision encoder, and a large language model, and is characterized in that The image quality assessment model further includes a quality enhancement module, and the image quality assessment device includes: A target image acquisition module configured to acquire at least two target images to be quality-assessed, wherein the at least two target images correspond to a target quality problem description for the at least two target images; A text embedding extraction module configured to extract a target text embedding corresponding to the target quality problem description through the text compiler; A visual embedding extraction module configured to extract target visual embeddings of the at least two target images respectively through the visual encoder; A quality embedding extraction module configured to extract target quality embeddings corresponding to at least one quality metric of the at least two target images respectively through the quality enhancement module; A splicing module, configured to splice the target text embedding, the target visual embeddings of the at least two target images, and the target quality embedding to obtain a target spliced embedding; A prediction module, configured to predict, by the large language model based on the target spliced embedding, a quality assessment target answer regarding the at least two target images.

13. An electronic device, characterized in that, Comprising: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the image quality assessment method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image quality assessment method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image quality assessment method according to any one of claims 1 to 11.