Image quality evaluation system and method based on multi-modal large model

Through an image quality evaluation system based on multimodal large model, combined with multi-scale feature abstraction and fusion of visual and text features, the problem of insufficient analysis of low-dimensional local distortion in the existing technology is solved, and unified processing and efficient description of image quality are achieved.

CN119991649APending Publication Date: 2025-05-13SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169091.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks in-depth analysis of low-dimensional local distortions in image visual question and answer tasks, and traditional image quality evaluation models cannot describe image quality loss in natural language and cannot label different regions.

Method used

Using an image quality evaluation system based on a multimodal large model, the input image and text description are converted into features through a visual encoder and a text encoder. The multi-scale feature abstractor extracts multi-scale features and merges them with the text embedding features to form fusion features. Then, the task processing module performs processing of mass division quantification, quality description and quality marking areas based on the fusion characteristics.

Benefits of technology

A unified processing framework for image quality is realized, data utilization efficiency is improved, low-level visual features are analyzed, and image quality loss can be described in natural language and different regions are marked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991649A_ABST
    Figure CN119991649A_ABST
Patent Text Reader

Abstract

The invention provides an image quality evaluation system and method based on a multi-modal large model. The system comprises an input module used for receiving an input image and text description; the visual encoder is used for converting the input image into visual feature codes; the text encoder is used for converting the text description into text embedding features; the multi-scale feature abstractor is used for extracting multi-scale features from the visual feature codes and combining the multi-scale features with the text embedded features; the task processing module is used for completing one or more of quality score quantification, quality description and quality labeling areas according to task types; and the output module is used for outputting a processing result of the task processing module. According to the method, a unified multi-modal framework is constructed, and quality score quantification, quality loss description and quality loss area labeling tasks of an image are integrated into a unified multi-modal large model, so that multi-task cooperative processing is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an image quality evaluation system and method based on a multimodal large model. Background Art

[0002] Current multimodal large-scale models perform well in image visual question answering tasks, but they mainly focus on high-dimensional semantic analysis and are insufficient in the in-depth analysis of low-dimensional local distortion and fine-grained visual features. In addition, traditional image quality assessment models mainly stay at the stage of quantifying image quality scores, and cannot properly describe image quality loss in natural language, nor can they mark different areas of image quality loss. In general, the shortcomings of existing technologies include:

[0003] Lack of a unified processing framework. Existing methods are independent of each other and cannot handle multiple quality-related visual tasks simultaneously within a unified framework.

[0004] Insufficient data integration: Data sets of different tasks have not been effectively integrated, which limits the generalization ability of the model;

[0005] Due to the shortcomings of fine-grained analysis, the existing technology needs to further improve the accuracy of quality-related low-level visual feature reasoning. Summary of the invention

[0006] In view of the defects in the prior art, the object of the present invention is to provide an image quality assessment system and method based on a multimodal large model.

[0007] According to one aspect of the present invention, there is provided an image quality assessment system based on a multimodal large model, comprising:

[0008] Input module: receives input images and text descriptions;

[0009] Visual encoder: convert the input image into visual feature encoding;

[0010] Text encoder: converts the text description into text embedding features;

[0011] Multi-scale feature abstractor: extracts multi-scale features from the visual feature encoding and merges them with the text embedding features to obtain fused features;

[0012] Task processing module: based on the fusion features and according to the task type, complete one or more of quality score quantification, quality description, and quality marking area;

[0013] Output module: outputs the processing result of the task processing module.

[0014] Preferably, the multi-scale feature abstractor extracts multi-scale features from different layers of the visual encoder, distills the multi-scale features using a fixed length, and cross-modally fuses the distilled multi-scale features with the text embedding features to obtain fused features.

[0015] Preferably, the task processing module includes a quality classification header module, a text generation header module and a segmentation decoder.

[0016] Preferably, the quality classification head module is a word vector probability classification head based on a multimodal large model;

[0017] The input of the quality classification head module is the fused features output by the multi-scale feature abstractor, and the output is a vector probability list and the corresponding quality score;

[0018] The quality classification head module uses a cross entropy loss function to optimize the discrete quality results predicted by the classification head.

[0019] Preferably, the text generation header module is a text generation header based on a multimodal macro model;

[0020] The input of the text generation head module is the fused features output by the multi-scale feature abstractor, and the output is the natural language text of the quality description corresponding to the fused features.

[0021] Preferably, the input of the segmentation decoder is the fused features output by the multi-scale feature abstractor, and the output is the corresponding segmentation mask;

[0022] The segmentation decoder uses a combination of binary cross entropy loss and DICE loss to optimize the segmentation task, ensuring the accuracy of the model in locating and segmenting distorted areas.

[0023] Preferably, the output module outputs the scoring results, text generation or marked area to the user;

[0024] The scoring result is a quality score output by the quality classification head module, which is presented in the form of discrete text-defined levels or in a specific score of 1-5;

[0025] The text generation is a natural text description output by the text generation header module, which is used to respond to specific instructions from the user;

[0026] The marked area is a segmentation mask output by the segmentation decoder, and is used to show the boundary of the quality distortion area.

[0027] According to a second aspect of the present invention, there is provided an image quality assessment method based on a multimodal large model, comprising:

[0028] Receive input image and text description;

[0029] Converting the input image into a visual feature encoding;

[0030] Converting the text description into text embedding features;

[0031] Extracting multi-scale features from the visual feature encoding and merging them with the text embedding features to obtain fused features;

[0032] Based on the fusion features, one or more tasks of quality score quantification, quality description, and quality marking area are completed according to the task type;

[0033] The processing result of the task processing is output.

[0034] According to a third aspect of the present invention, there is provided a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can be used to run the system or the method when executing the program.

[0035] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to run the system or to execute the method.

[0036] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0037] The image quality assessment system based on a multimodal large model of an embodiment of the present invention constructs a unified multimodal framework: the image quality score quantification, quality loss description and quality loss area labeling tasks are integrated into a unified multimodal large model to achieve collaborative processing of multiple tasks.

[0038] The image quality evaluation system based on a multimodal large model in an embodiment of the present invention improves data utilization efficiency: by cross-task data fusion, it fully utilizes data sets of different tasks, and uses corresponding loss functions in a targeted manner to enhance the generalization ability and robustness of the model.

[0039] The image quality assessment system based on a multimodal large model in an embodiment of the present invention strengthens the fine-grained analysis capability related to quality assessment: a multi-scale feature abstractor is introduced to improve the model's understanding of low-level visual features and the accuracy of fine-grained quality loss area annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0041] Figure 1is a framework diagram of an image quality assessment system based on a multimodal large model in one embodiment of the present invention;

[0042] Figure 2 The figure is a flowchart of an image quality assessment method based on a multimodal large model in one embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0044] like Figure 1 As shown, in one embodiment of the present invention, a system for evaluating image quality based on a multimodal large model is provided, comprising:

[0045] Input module: used to receive input images and text descriptions;

[0046] Visual encoder: converts the input image into visual feature encoding;

[0047] Text encoder: converts text description into text embedding features;

[0048] Multi-scale feature abstractor: extracts multi-scale features from visual feature encoding and merges them with text embedding features to obtain fused features;

[0049] Task processing module: Based on the fusion features and according to the type of tasker, complete one or more of the following: quality score quantification, quality description, and quality marking area;

[0050] Output module: outputs the processing results of the task processing module.

[0051] The above embodiments can achieve quantitative evaluation of image quality, describe quality loss in detail, and mark quality loss areas. This can unify the task framework in a targeted manner, integrate data sets of different tasks, improve the reasoning accuracy of the model, and significantly enhance the interpretability of quality evaluation.

[0052] The input module in the above embodiment receives images and text descriptions, such as quality evaluation, instructions, etc., which are mainly used to standardize task types such as image quality score quantification, quality loss description, and quality loss area annotation.

[0053] In some specific embodiments, CLIP-ViT-L / 14-336 can be used as a visual encoder to convert the input image into a visual feature encoding. The pre-trained large language model LLaVA-7B can be used to convert the input text into a text feature embedding. Multi-scale feature abstractor: extracts multi-scale features from different layers of the visual encoder and merges them with the text embedding. Useful information is distilled from the multi-scale features using a fixed-length query (such as 256 tokens). This module cross-modally fuses visual features and text features through a multi-head attention mechanism to ensure that the model can understand both image and text information at the same time.

[0054] In order to better handle different tasks, in a preferred embodiment, the task processing module includes a quality classification header module, a text generation header module and a segmentation decoder.

[0055] Furthermore, in a preferred embodiment, the quality classification head module is a word vector probability classification head based on a multimodal large model; the input of the quality classification head module is the fusion feature output by the multi-scale feature abstractor, and the output is a word vector probability list (representing the probability of discrete adjectives) and a quality score. Specifically, based on the quality classification output head, a probability vector of a discrete adjective (such as "good", "medium", "bad") is generated. Assume that the model output is a word probability vector p = [p 好 ,p 中 ,p 差 ],in:

[0056] p i =softmax(z i ),z i ∈R,

[0057] Here i are the logits of the classification head for category i.

[0058] The quantification of quality scores is based on the weighted average of the probabilities of discrete adjectives. Assume that the scores corresponding to "good", "average" and "poor" are s 好 ,s 中 ,s 差 , then the mass fraction q can be expressed as:

[0059]

[0060] Among them, s 好 >s 中 >s 差 , these scores can be set according to task requirements (e.g. 5, 3, 1).

[0061] The quality classification head module uses the cross entropy loss function to optimize the discrete quality results predicted by the classification head. Specifically, assuming that the training data is labeled y∈{1, 2, 3} corresponding to the adjectives "good", "medium", and "bad", the cross entropy loss is defined as:

[0062]

[0063] Among them, l|y=i| is an indicator function, which is 1 when y=i and 0 otherwise. By introducing quality-aware annotation data into the multimodal large model, the model's perception of quality classification can be enhanced, thereby optimizing the output of the word probability vector.

[0064] Furthermore, in a preferred embodiment, the text generation header module is a text generation header based on a multimodal large model; the input of the text generation header module is the fused features output by the multi-scale feature abstractor, and the output is the natural language text of the quality description corresponding to the fused features.

[0065] Further, in a preferred embodiment, the input of the segmentation decoder is the fused features output by the multi-scale feature abstractor, and the output is the corresponding segmentation mask, specifically:

[0066] The input of the segmentation decoder is the multi-scale feature representation generated by the multimodal large model, denoted as F = {F1, F2, ..., F n}, where F i is the feature tensor of the i-th scale. The multi-scale feature fuser abstracts and fuses the multi-scale features to obtain the fused and optimized feature representation F * Assuming that the fusion process is a function g(·), then:

[0067] F * =g({F1,F2,…,F n})

[0068] The segmentation decoder transforms the feature F * Decoded into a segmentation mask M, where the value of each pixel indicates whether it belongs to the distorted area. Assuming the segmentation decoder is a function h(·), then:

[0069] M=h(F * )

[0070] where M∈[0,1] H×W is the segmentation mask, H and W are the height and width of the input image, and the mask value is the probability value, indicating the confidence that the pixel belongs to the distorted area. Finally, M can be binarized by a threshold τ to obtain a binary mask M bin :

[0071]

[0072] The segmentation decoder uses a combination of binary cross entropy loss and DICE loss to optimize the segmentation task, ensuring the accuracy of the model positioning and segmentation distorted areas. Specifically:

[0073] A combination of binary cross entropy loss (BCE) and DICE loss is used in the optimization process.

[0074] Among them, the acquisition process of binary cross entropy loss (BCE) is as follows:

[0075] Assume M gt ∈{0,1} H×W is the corresponding real segmentation mask, M is the predicted segmentation mask, and the binary cross entropy loss is:

[0076]

[0077] Among them, the process of obtaining DICE loss is as follows:

[0078] The DICE loss is used to measure the overlap between the predicted mask and the true mask, which is defined as

[0079]

[0080] The final optimization target is a weighted combination of binary cross entropy loss and DICE loss, expressed as:

[0081]

[0082] Where α and β are weight hyperparameters used to balance the impact of the two losses.

[0083] Based on the same inventive concept, another embodiment of the present invention also provides an image quality assessment method based on a multimodal large model, such as Figure 2 As shown, it includes the following steps:

[0084] Step 1, receiving input image and text description;

[0085] Step 1.1: The user uploads an image and enters a text description (such as quality rating, instructions, etc.).

[0086] Step 1.2: The system determines the subsequent processing flow according to the specified task type described in the user text (such as image quality score quantification, quality loss description, and quality loss area annotation).

[0087] Step 2, convert the input image into visual feature encoding;

[0088] Step 3, convert the text description into text embedding features;

[0089] Step 4: extract multi-scale features from the visual feature encoding and merge them with the text embedding features to obtain fused features;

[0090] Step 4.1: The multi-scale feature abstractor extracts multi-scale features from different layers of the visual encoder and merges them with the text embedding.

[0091] Step 4.2: Use a multi-head attention mechanism to cross-modally fuse visual features and text features to ensure that the model can understand both image and text information.

[0092] Step 5: Based on the fusion features and according to the type of tasker, one or more of quality score quantification, quality description, and quality marking area are completed;

[0093] Step 5.1: According to the task type, call the corresponding processing module. For the quality score quantization task, use the classification head for scoring; for the quality description task, use the generation head for text generation; for the quality annotation region task, use the segmentation decoder to generate region annotations.

[0094] Step 6: Output the processing result of the task processing module.

[0095] Step 6.1: The output module outputs the scoring results, text generation or region annotation to the user.

[0096] Step 6.2: The user can further adjust the input based on the feedback from the model and repeat the above steps to achieve human-computer interactive visual quality assessment and analysis.

[0097] Each step in the above example of the present invention can refer to the specific implementation technology of the module / unit corresponding to the image quality assessment system based on the multimodal large model in the above embodiment, which will not be repeated here.

[0098] Based on the same inventive concept, in other embodiments of the present invention, a terminal is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can be used to execute the method described in the preceding claims, or to run the system described.

[0099] Based on the same inventive concept, in other embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to run the system or execute the method.

[0100] In a specific embodiment, an image quality assessment system based on a multimodal large model is used to perform image quality assessment and quality loss area labeling tasks. The execution process of each module of the system is as follows:

[0101] The user uploads an image with blur and overexposure problems to the input module and chooses to input instructions for quality score quantification and quality loss area annotation (i.e., text description entered by the user), such as "Please give the visual instruction score of this image and mark the blurry and overexposed areas."

[0102] The system first extracts the visual features of the input image through a visual encoder, and then converts the text description entered by the user into text features through a text encoder.

[0103] The multi-scale feature abstractor cross-modally fuses visual features and text features to generate multi-scale feature representations.

[0104] For the quality score quantification task, the system uses the classification head to score the image, outputs a "medium" level quality rating, and gives a specific score of 2.5 based on probability;

[0105] For the quality loss area labeling task, the system uses a segmentation decoder to generate segmentation masks of blurred and overexposed areas and displays them visually to the user.

[0106] Of course, the user can further adjust the input according to the feedback of the output, such as further requesting to mark the most blurred area, such as "please further mark the most blurred area based on the currently marked blurred area", thereby realizing human-computer interactive visual quality evaluation and analysis.

[0107] In another specific embodiment, an image quality assessment system based on a multimodal large model is used to perform the image quality description task. The execution process of each module of the system is as follows:

[0108] The user uploads an image with low-light issues to the input module and enters the instruction (a user-entered text description): "Please describe the visual quality of this image."

[0109] The system first extracts the visual features of the image through a visual encoder, and then converts the user input instructions into text features through a text encoder.

[0110] The multi-scale feature abstractor cross-modally fuses visual features and text features to generate multi-scale feature representations.

[0111] The system uses a generative head to generate a text response that fits the instruction based on the multi-scale features, describing the low-level visual features of the image, such as "This image has obvious low-light issues, resulting in dark colors and blurred details."

[0112] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be used in any combination without conflicting with each other.

Claims

1. An image quality assessment system based on a multimodal large model, characterized in that: include: Input module: receives input images and text descriptions; Visual encoder: convert the input image into visual feature encoding; Text encoder: converts the text description into text embedding features; Multi-scale feature abstractor: extracts multi-scale features from the visual feature encoding and merges them with the text embedding features to obtain fused features; Task processing module: based on the fusion features and according to the task type, complete one or more of quality score quantification, quality description, and quality marking area; Output module: outputs the processing result of the task processing module.

2. The image quality assessment system based on a multimodal large model according to claim 1, characterized in that: The multi-scale feature abstractor extracts multi-scale features from different layers of the visual encoder, distills the multi-scale features using a fixed length, and cross-modally fuses the distilled multi-scale features with the text embedding features to obtain fused features.

3. The image quality assessment system based on a multimodal large model according to claim 1, characterized in that: The task processing module includes a quality classification head module, a text generation head module and a segmentation decoder.

4. The image quality assessment system based on a multimodal large model according to claim 3, characterized in that: The quality classification head module is a word vector probability classification head based on a multimodal large model; The input of the quality classification head module is the fused features output by the multi-scale feature abstractor, and the output is a word vector probability list and the corresponding quality score; The quality classification head module uses a cross entropy loss function to optimize the discrete quality results predicted by the classification head.

5. The image quality assessment system based on a multimodal large model according to claim 3, characterized in that: The text generation header module is a text generation header based on a multimodal large model; The input of the text generation head module is the fused features output by the multi-scale feature abstractor, and the output is the natural language text of the quality description corresponding to the fused features.

6. The image quality assessment system based on a multimodal large model according to claim 3, characterized in that: The input of the segmentation decoder is the fused features output by the multi-scale feature abstractor, and the output is the corresponding segmentation mask; The segmentation decoder uses a combination of binary cross entropy loss and DICE loss to optimize the segmentation task, ensuring the accuracy of the model in locating and segmenting distorted areas.

7. The image quality assessment system based on a multimodal large model according to claim 3, characterized in that: The output module outputs the scoring results, text generation or marked areas to the user; The scoring result is a quality score output by the quality classification head module, which is presented in the form of discrete text-defined levels or in a specific score of 1-5; The text generation is a natural text description output by the text generation header module, which is used to respond to specific instructions from the user; The marked area is a segmentation mask output by the segmentation decoder, and is used to show the boundary of the quality distortion area.

8. An image quality assessment method based on a multimodal large model, characterized in that: include: Receive input image and text description; Converting the input image into a visual feature encoding; Converting the text description into text embedding features; Extracting multi-scale features from the visual feature encoding and merging them with the text embedding features to obtain fused features; Based on the fusion features, one or more tasks of quality score quantification, quality description, and quality marking area are completed according to the task type; The processing result of the task processing is output.

9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it can be used to run the system described in any one of claims 1 to 7, or to execute the method described in claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it can be used to run the system described in any one of claims 1 to 7, or to execute the method described in claim 8.