Image quality evaluation method and device, equipment and storage medium
By combining a large language model with text and visual vector sequence fusion, the subjectivity and high cost of image quality assessment are solved, and more efficient and accurate image quality assessment is achieved to meet diverse assessment needs.
Patent Information
- Application Number
- CN202510450153.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-09-05
AI Technical Summary
Existing image quality assessment methods mainly rely on manual labeling, which has the problems of highly subjective assessment results, high cost and low efficiency.
A large language model is combined with a text and visual vector sequence fusion method. The image to be evaluated and the text prompt are obtained, vectorized and then fused. The large language model is used to perform multimodal vector sequence processing to obtain the image quality assessment result.
It improves the accuracy and efficiency of image quality assessment, reduces the assessment cost, and has stronger generalization ability and higher assessment accuracy, adapting to image assessment in different types and cultural backgrounds.
Smart Images

Figure CN120599618A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to an image quality assessment method, apparatus, device, and storage medium. Background Art
[0002] With the widespread application of digital images, image quality assessment has broad application value and significance in various fields involving digital image applications, such as image recommendation algorithms and image generation effect evaluation. Image quality assessment includes the evaluation of image quality and image aesthetics. Image quality assessment focuses on the evaluation of the technical characteristics of quantified images, such as clarity, distortion, and noise. Image aesthetics assessment focuses on the evaluation of the aesthetic characteristics of quantified images.
[0003] At present, image quality assessment methods are mainly based on manual annotation and subjective evaluation. Therefore, there are problems such as highly subjective assessment results, high assessment costs, and low efficiency. Summary of the Invention
[0004] The present application provides an image quality assessment method, apparatus, device, and storage medium, which can improve the accuracy and efficiency of image quality assessment and reduce the cost of image quality assessment.
[0005] In a first aspect, an image quality assessment method is provided, comprising: obtaining an image to be assessed and a text prompt, the text prompt instructing a preset large language model to perform quality assessment on the image to be assessed; vectorizing the text prompt to obtain a text vector sequence; vectorizing the image to be assessed to obtain a visual vector sequence; fusing the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence; and processing the multimodal vector sequence through the large language model to obtain a quality assessment result of the image to be assessed in at least one quality assessment dimension.
[0006] In a second aspect, an image quality assessment device is provided, including: a first acquisition module, used to acquire an image to be assessed and a text prompt, the text prompt instructing a preset large language model to perform quality assessment on the image to be assessed; a first processing module, used to vectorize the text prompt to obtain a text vector sequence; vectorize the image to be assessed to obtain a visual vector sequence; a sequence fusion module, used to fuse the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence; and a second processing module, used to process the multimodal vector sequence through the large language model to obtain a quality assessment result of the image to be assessed in at least one quality assessment dimension.
[0007] In a third aspect, an electronic device is provided, comprising: a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in the first aspect or its various implementations.
[0008] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method according to the first aspect or its various implementations.
[0009] In a fifth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method in the first aspect or its various implementations.
[0010] In a sixth aspect, a computer program is provided, which enables a computer to execute the method in the first aspect or its various implementations.
[0011] In summary, the present application can first obtain the image to be evaluated and the text prompt, and the text prompt instructs the preset large language model to perform quality assessment on the image to be evaluated; then, the text prompt can be vectorized to obtain a text vector sequence; the image to be evaluated can be vectorized to obtain a visual vector sequence; thereafter, the text vector sequence and the visual vector sequence can be fused to obtain a multimodal vector sequence; finally, the multimodal vector sequence can be processed by the large language model to obtain the quality assessment result of the image to be evaluated under at least one quality assessment dimension. Therefore, not only can multi-dimensional and automated assessment of image quality be achieved, the accuracy and efficiency of image quality assessment can be improved, and the cost of image quality assessment and the subjectivity of assessment results can be reduced; moreover, since the large language model is pre-trained on large-scale data and has a higher contextual understanding ability, the use of the large language model for image quality assessment can better adapt to image assessments of different types, different fields and different cultural backgrounds and meet diverse assessment needs, and has stronger generalization capabilities and higher assessment accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The following is an introduction to the drawings required for describing the embodiments.
[0013] Figure 1 A flowchart of an image quality assessment method provided in an embodiment of the present application;
[0014] Figure 2 A schematic diagram of an image quality assessment method provided in an embodiment of the present application;
[0015] Figure 3 A schematic diagram of an image quality assessment device 300 provided in an embodiment of the present application;
[0016] Figure 4 4 is a schematic diagram of an electronic device 400 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The technical solution of this application will be introduced below in conjunction with the drawings in this application.
[0018] It should be noted that the information, data (including, but not limited to, data used for analysis, stored data, displayed data, etc., such as images to be evaluated, text prompts, sample images, and sample text prompts, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the images to be evaluated and the operations performed on the images to be evaluated involved in this application are all obtained with full authorization.
[0019] In one embodiment, the technical solution of the present application can be used in image quality assessment scenarios. For example, it can be applied to image quality assessment scenarios in the fields of image recommendation, image generation effect evaluation, etc., but is not limited thereto.
[0020] The evaluated image, i.e., the image to be evaluated, may be an image of any type, such as a portrait image, a landscape image, an image captured by a camera, or an image generated by a device, but is not limited thereto.
[0021] In one embodiment, the solution provided in this application can be executed by any electronic device with data processing capabilities. For example, the electronic device can be a server, specifically an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In another example, the electronic device can be a terminal device, specifically a tablet computer, a laptop computer, or a desktop computer. In another example, the electronic device can be a combination of a server and a terminal device, wherein the server and terminal device in the combination can communicate via wireless or wired means. This application does not impose any specific restrictions on the electronic device.
[0022] The following is an introduction to various embodiments of the technical solution of this application:
[0023] It should be noted that all technical solutions in this application can be combined in any way to form optional embodiments of this application, which will not be described one by one here.
[0024] Figure 1 This is a flowchart of an image quality assessment method provided in an embodiment of the present application, which can be executed by the electronic device described above. Figure 1 As shown, the method includes:
[0025] S110: Obtain an image to be evaluated and a text prompt, where the text prompt instructs a preset large language model to perform quality evaluation on the image to be evaluated;
[0026] S120: performing vectorization processing on the text prompt to obtain a text vector sequence; performing vectorization processing on the image to be evaluated to obtain a visual vector sequence;
[0027] S130: Fusing the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence;
[0028] S140: Processing the multimodal vector sequence through the large language model to obtain a quality assessment result of the image to be assessed in at least one quality assessment dimension.
[0029] In one embodiment, the large language model can be a model from the Generative Pre-trained Transformer (GPT) series of natural language processing models based on the Transformer architecture or an LLaMA series model, but is not limited thereto. Alternatively, other models other than the large language model can be used to perform S140, but is not limited thereto.
[0030] In one embodiment, in order to guide the large language model to generate descriptions related to image quality assessment, the text prompt can be set as questions related to image quality assessment, that is, text in the form of questions, specifically text in the form of dialogues, which is not limited in this application.
[0031] In addition, the text prompt can be a text in any language, such as Chinese, English, etc., but not limited thereto.
[0032] Specifically, the text prompt may include a prompt for evaluating the image under at least one quality assessment dimension. The at least one quality assessment dimension may include: an image quality assessment dimension and an image aesthetics assessment dimension. The image quality assessment dimension focuses on evaluating the technical characteristics of the quantized image, such as clarity, distortion, and noise; the image aesthetics assessment dimension focuses on evaluating the aesthetic characteristics of the quantized image.
[0033] For example, the prompt for evaluating the image to be evaluated under the image quality evaluation dimension may be:
[0034] Human: Can you evaluate the quality of the image?
[0035] GPT: The quality of the image is <level>.
[0036] The prompts for evaluating the image under the image aesthetics evaluation dimension can be:
[0037] Human: Can you evaluate the aesthetics of the image?
[0038] GPT: The aesthetics of the image is <level>.
[0039] in," <level>" is a descriptive word of different degrees obtained after evaluating the image quality or aesthetics, for example, it can be any one of {bad, poor, fair, good, excellent}. In the following embodiments, " ” for introduction.
[0040] In addition, the text prompt can also instruct the large language model to reconstruct the image to be evaluated, so that the large language model involves the underlying image features of the image to be evaluated when processing the multimodal vector sequence.
[0041] It is understandable that the pre-trained large language model has good high-level semantic information analysis capabilities, but the relevant technology often ignores the targeted optimization of low-level information analysis capabilities in the training tasks of large language models; however, visual low-level information, that is, the underlying features of the image, helps to evaluate the image quality more accurately and comprehensively, and image reconstruction often requires the use of underlying image features. Therefore, the text prompts that instruct the large language model to reconstruct the image to be evaluated can be used to guide the large language model to process the multimodal vector sequence involving the underlying image features of the image to be evaluated, so that the quality of the image to be evaluated can be more accurately and comprehensively evaluated. In addition, this application will introduce this in detail when introducing the training methods of the models and modules in subsequent embodiments. To avoid repetition, we will not go into details here.
[0042] Among them, low-level image features can be features related to local areas or details of the image, for example, features extracted from image pixels, specifically image edge features, image texture features, image color features, image corner features, image contrast or brightness features, etc. High-level image features, i.e., visual high-level information, can be related to high-level semantic information of the image, for example, the categories of objects in the image, the scenes involved, etc.
[0043] Accordingly, based on the above embodiment, the text prompt instructing the large language model to reconstruct the image to be evaluated may be as follows:
[0044] Human: Can you evaluate the quality of the image? And reconstruct the image.
[0045] GPT: The quality of the image is <level>.
[0046] Human:Can you evaluate the aesthetics of the image?
[0047] GPT: The aesthetics of the image is <level>.
[0048] in," " can identify the location where the image to be evaluated is embedded or will be generated, that is, the location where the image to be evaluated conforms to the semantics of the text prompt in the text prompt. The first" "Corresponding to the image to be evaluated, the second" " corresponds to the subsequently reconstructed image. In addition, " " can also identify the position where the visual vector sequence is embedded into the text vector sequence, which will be introduced in the subsequent embodiments.
[0049] In the above content, specific text prompts can be set to guide the large language model to perform corresponding processing on the image to be evaluated. Specifically, the large language model is guided to process according to the content indicated by the text prompts. For example, the quality of the image to be evaluated is evaluated in at least one dimension, and the underlying features of the image are taken into consideration when evaluating the image to be evaluated. This can meet the diverse and comprehensive evaluation needs of the image to be evaluated, so that the quality of the image to be evaluated can be evaluated more accurately and comprehensively.
[0050] In one embodiment, after obtaining the text prompt and the image to be evaluated, vectorization processing can be performed on the text prompt and the image to be evaluated respectively, that is, feature extraction can be performed to obtain vector sequences corresponding to the text prompt and the image to be evaluated respectively.
[0051] Exemplarily, the above-mentioned vectorization processing of the text prompt to obtain the text vector sequence includes: segmenting the text prompt to obtain a word sequence; and converting the word sequence into an embedding vector to obtain a text vector sequence.
[0052] For example, a tokenizer can be used to segment text prompts. The tokenizer can be any type of tokenizer that meets the requirements of the large language model. That is, the word sequence generated by the tokenizer meets the processing requirements set by the large language model. Word segmentation methods used include, but are not limited to, character-level segmentation, sub-word segmentation, and wordpiece segmentation.
[0053] For example, a word sequence can be converted into an embedding vector using a word embedding module. A word embedding module is a structure used to store fixed-size word embeddings that capture the grammatical and semantic relationships between words in a word sequence. It uses a lookup table to map integer indices to fixed-size dense vectors (embedding vectors), effectively converting words or word sequences into dense vector representations.
[0054] Exemplarily, the above-mentioned vectorization processing of the image to be evaluated to obtain a visual vector sequence includes: extracting visual features of the image to be evaluated using an image feature extractor to obtain a visual vector sequence. Specifically, the image feature extractor can be used to first extract visual features of the image to be evaluated to obtain an initial visual vector sequence; then, the initial visual vector sequence can be converted using a first adapter (e.g., an adapter) to obtain a visual vector sequence under the text feature dimension corresponding to the text vector sequence, so as to spatially align the text vector sequence and the visual vector sequence to ensure consistency in feature dimensions, thereby facilitating subsequent feature fusion of the text vector sequence and the visual vector sequence.
[0055] Of course, the text vector sequence can also be converted through an adapter to obtain a text vector sequence under the visual feature dimension corresponding to the visual vector sequence.
[0056] For example, the image feature extractor may be an encoder. The encoder is a part used to convert input data into a feature vector or a coded representation, which can process the input data and extract features therein. Specifically, in computer vision tasks, that is, in the above embodiment, the encoder can extract visual features in the image to be evaluated. The encoder may be an encoder in a MAE-ViT network that has been fine-tuned for the image reconstruction task on the ImageNet dataset (the image to be evaluated may be derived from the image dataset), which corresponds to the Decoder in the image reconstruction task, that is, it comes from the same network, and is more conducive to the image reconstruction task. To avoid repetition, this application will introduce the image reconstruction task in subsequent embodiments, which will not be elaborated here.
[0057] Of course, the image feature extractor can also be other encoders, for example, the image-text encoder in the Contrastive Language-Image Pre-Training (CLIP) model, and this application does not impose any restrictions on this.
[0058] For example, the image feature extractor can be a pre-trained image feature extraction network based on the Transformer structure, which can obtain a token (word segmentation) sequence representing visual features, that is, a visual vector sequence.
[0059] In one embodiment, fusing the text vector sequence and the visual vector sequence to obtain the multimodal vector sequence includes: embedding the visual vector sequence into a specific position of the text vector sequence to obtain the multimodal vector sequence, wherein the specific position corresponds to a corresponding position of the image to be evaluated in the text prompt.
[0060] For example, the specific position of the image to be evaluated in the text prompt can be determined first; then, the visual vector sequence is embedded into the specific position of the text vector sequence. The specific position can be the above Where the placeholder is located.
[0061] Exemplarily, the above-mentioned adapter can also be used to fuse the text vector sequence and the visual vector sequence. Specifically, the adapter can adopt a simple network, such as a fully connected layer (FC layer), or a more complex modal fusion network, such as a Q-former (used to build a bridge between the image encoder and the language model so that image features and text features can be effectively fused, that is, the text vector sequence and the visual vector sequence can be effectively fused). Among them, the adapter used can be determined in combination with the large language model structure and model performance.
[0062] In the above content, a visual vector sequence representing image features (for example, a visual token) and a text vector sequence representing text features of text prompts (for example, a text token) can be combined to form a multimodal vector sequence (for example, a multimodal token sequence) as the input of a large language model, thereby achieving comprehensive consideration of visual features and text features, facilitating the large language model to evaluate the image to be evaluated at a deeper level, and to achieve a deeper and more comprehensive image quality evaluation using richer semantic and contextual information in images and texts.
[0063] Furthermore, when fusing the text vector sequence and the visual vector sequence, the position of the image to be evaluated in the text prompt is taken into account, which can ensure that the fused multimodal vector sequence can more accurately express the information in the image to be evaluated and the text prompt, thereby further improving the accuracy of image quality assessment.
[0064] In one embodiment, after determining the multimodal vector sequence, the multimodal vector sequence may be processed by a large language model to obtain a quality assessment result for the image to be evaluated in at least one quality assessment dimension. The multimodal vector sequence may be input into the large language model, which may perform calculations on the multimodal vector sequence to obtain a quality assessment result.
[0065] Exemplarily, a multimodal vector sequence is processed by a large language model to obtain a text description of the image to be evaluated under at least one quality assessment dimension; then, the text description is converted to obtain a text score value or a probability distribution of the text score for the image to be evaluated under at least one quality assessment dimension; finally, the quality assessment result can be determined based on the text score value or the probability distribution of the text score.
[0066] For example, the text description can be processed by Softmax and Convert to obtain a quality assessment result. Among them, Softmax is an activation function used for multi-category classification problems. It can convert the output of the neural network, such as the above-mentioned text description, into a form representing a probability distribution; it can make the output value of each category between 0 and 1 and the sum to 1, so that the probability of each category can be output. Specifically, Softmax can be used to determine the text score probability distribution of each category corresponding to the dimension of the image to be evaluated under each quality assessment dimension. Convert is a unit conversion method used to convert a value from one unit of measurement to another unit of measurement; specifically, the value output by Softmax can be converted by Convert to obtain a text score value or a probability distribution of a text score, or the text score value or the probability distribution of a text score can be converted to obtain a quality assessment result. This application does not impose any restrictions on this.
[0067] Exemplarily, the text description can be converted into a specific score, i.e., a text rating value, according to the softmax-based conversion method. Specifically, in combination with the above embodiment, for any quality assessment dimension in at least one quality assessment dimension, a set of quality descriptors of different degrees (i.e., multiple categories corresponding to the quality assessment dimension) can be preset first, such as {bad, poor, fair, good, excellent}; then, the probability distribution of the image to be evaluated for each quality descriptor under the quality assessment dimension can be determined based on the text description, denoted as χ, and the scores corresponding to each quality descriptor: G:l i →i, for example: {bad, poor, fair, good, excellent} correspond to {1, 2, 3, 4, 5} respectively; then, the probability value pli of each quality description word can be obtained according to formula (1), and the probability value pli can be processed according to formula (2) to obtain the final score corresponding to the quality assessment dimension, that is, the text score value.
[0068]
[0069] Among them, l i is the descriptor of the preset quality descriptor, i∈{1,2,3,4,5}.
[0070] In one embodiment, the text score value and / or the probability distribution of the text score may be directly used to determine the quality assessment result.
[0071] Alternatively, before determining the quality assessment result based on the text score value or the probability distribution of the text score, the multimodal vector sequence can be processed by a large language model to obtain a reconstructed visual vector sequence; then, the reconstructed visual vector sequence can be processed to obtain the visual score value or the probability distribution of the visual score of the image to be evaluated in at least one quality assessment dimension; finally, the quality assessment result can be determined based on the text score value or the probability distribution of the text score, and the visual score value or the probability distribution of the visual score.
[0072] The reconstructed visual vector sequence involves underlying image features of the image to be evaluated. Specifically, content related to image reconstruction can be set in a text prompt (for example, a text prompt can be set to instruct the large language model to perform image reconstruction on the image to be evaluated), so that the large language model can process the multimodal vector sequence in a manner that involves underlying image features of the image to be evaluated, thereby guiding the large language model to generate a reconstructed visual vector sequence involving underlying image features of the image to be evaluated.
[0073] In addition, in addition to being used to determine the visual score value or the probability distribution of the visual score, the above-mentioned reconstructed visual vector sequence can also be used to reconstruct the image to be evaluated, so that the large language model can involve, consider, and use the underlying features of the image when evaluating the image to be evaluated. This part will be introduced in the subsequent model training part. To avoid repetition, it will not be repeated here.
[0074] For example, if the feature dimension of the visual vector sequence in the multimodal vector sequence is the text feature dimension corresponding to the text vector sequence, then before processing the reconstructed visual vector sequence, the reconstructed visual vector sequence can be converted through a second adapter (for example, Deadaptor) to obtain a reconstructed visual vector sequence under the visual feature dimension corresponding to the visual vector sequence, so that the reconstructed visual vector sequence used in subsequent image reconstruction or determination of the visual score value or the probability distribution of the visual score meets the input requirements of the corresponding model or module.
[0075] For example, by setting text prompts, the large language model can generate corresponding visual tokens, that is, reconstruct the visual vector sequence. Specifically, based on the above content, The reconstructed visual vector sequence is obtained from the output of the large language model. Then, the Deadaptor module converts the reconstructed visual vector sequence from the text feature space (i.e., the text feature dimension) to the image feature space (i.e., the image feature dimension) to obtain the image feature (VisualFeature). Then, the image quality score regression module (Score Head) can process the image features from the perspective of visual features to evaluate the image quality of the image to be evaluated, predict the image quality assessment score, and obtain the visual score value.
[0076] For example, the image quality score regression module can determine the visual score based on a single score regression scheme. Specifically, the image quality score regression module can include a multi-layer regression network of a Sigmoid function, whose output dimension is 1, and the Sigmoid function is used to control the predicted score within the range of (0,1). Alternatively, the image quality score regression module can be a multi-classification head that predicts the probability distribution of the quality assessment score and determines the final visual score value as the expected value of the probability distribution through formula (3).
[0077]
[0078] Among them, x i represents the quality assessment score, p i Represents the quality assessment score x i The probability of n is the number of categories of quality assessment scores.
[0079] Exemplarily, the weighted sum of the text score value and the visual score value can be determined first, and the weighted sum corresponding to the probability distribution of the text score and the probability distribution of the visual score can be determined. Then, at least one of the above two weighted sum results can be determined as the quality assessment result.
[0080] In the above content, we can not only utilize the higher context understanding ability and strong generalization ability of the large language model to better adapt to image evaluation of different types, different fields and different cultural backgrounds and meet diverse evaluation needs, ensuring that the generated scoring results are closer to human perception; but also take into account the visual scoring results (i.e., the visual scoring value and the probability distribution of the visual scoring), so that the text scoring results (i.e., the text scoring value and the probability distribution of the text scoring) can be corrected through the visual scoring results to further improve the evaluation accuracy.
[0081] In addition, by setting text prompts, the large language model is guided to consider the underlying features of the image during the processing process, further improving the accuracy and comprehensiveness of the evaluation.
[0082] Moreover, since the text score value, the visual score value, the probability distribution of the text score and the probability distribution of the visual score are all in numerical form, quantitative evaluation can be achieved, thereby further improving the accuracy of the evaluation results.
[0083] The following is an introduction to the training process of the models and modules involved in the technical solution of this application.
[0084] In one embodiment, the large language model and / or the modules mentioned above may be trained by the following steps:
[0085] S201: Acquire a sample image and a sample text prompt, where the sample text prompt includes a first prompt for instructing the large language model to perform quality assessment on the sample image and a second prompt for instructing the large language model to perform image reconstruction on the sample image;
[0086] S202: performing vectorization processing on the sample text prompt to obtain a sample text vector sequence; performing vectorization processing on the sample image to obtain a sample visual vector sequence;
[0087] S203: Fusing the sample text vector sequence and the sample visual vector sequence to obtain a sample multimodal vector sequence;
[0088] S204: Processing the sample multimodal vector sequence through the large language model to obtain a sample reconstructed visual vector sequence and a sample text description of the sample image in at least one quality assessment dimension;
[0089] S205: Reconstruct the sample image according to the sample reconstructed visual vector sequence to obtain a training image;
[0090] S206: Determine a sample visual score value or a probability distribution of the sample visual score of the sample image in at least one quality assessment dimension according to the sample reconstructed visual vector sequence;
[0091] S207: Determine, based on the sample text description, a sample text score value or a probability distribution of the sample text score for the sample image under at least one quality assessment dimension;
[0092] S208: Train the corresponding models or modules involved in the above steps based on at least one of the sample text vector sequence, sample visual vector sequence, sample multimodal vector sequence, sample reconstructed visual vector sequence, sample text description, training image, sample visual rating value, probability distribution of sample visual rating, sample text rating value and probability distribution of sample text rating.
[0093] Among them, the content and effects corresponding to the training process of the model and module can be referenced with the content and effects corresponding to the reasoning process of the above-mentioned model and module. For example, the sample text vector sequence, sample visual vector sequence, sample multimodal vector sequence, sample reconstructed visual vector sequence, sample text description, sample visual rating value, sample visual rating probability distribution, sample text rating value and sample text rating probability distribution corresponding to each content and determination method can all refer to the text vector sequence, visual vector sequence, multimodal vector sequence, reconstructed visual vector sequence, text description, visual rating value, visual rating probability distribution, text rating value and text rating probability distribution corresponding to the content and determination method in the above-mentioned embodiments.
[0094] It should be noted that this application does not limit the execution order of the above steps. For example, S207 can be executed first, and then S206.
[0095] In addition, each of the above steps can be executed by the same model or module, or by different models or modules. For example, the sample text prompt can be processed by a text processing module to obtain a sample text vector sequence; the sample image can be processed by an image processing module to obtain a sample visual vector sequence; the sample image can be reconstructed by an image reconstruction module based on the sample reconstructed visual vector sequence to obtain a training image; the sample visual score value or the probability distribution of the sample visual score of the sample image under at least one quality assessment dimension can be determined by a visual assessment module (which can include the above-mentioned image quality score regression module) based on the sample reconstructed visual vector sequence; the sample text score value or the probability distribution of the sample text score of the sample image under at least one quality assessment dimension can be determined by a text assessment module based on the sample text description.
[0096] The above module division is only exemplary and can be combined or split according to actual applications. For example, each of the above steps can be performed by the large language model, or can be performed by at least one module or other model other than the large language model. The at least one module can be a module in the large language model or a module or unit in other models. This application does not impose any restrictions on this.
[0097] Exemplarily, before reconstructing the sample image based on the sample reconstructed visual vector sequence, the sample reconstructed visual vector sequence can be converted using a second adapter to obtain a sample reconstructed visual vector sequence in the image feature space; then, a decoder can be used to decode the sample reconstructed visual vector sequence into a sample image, that is, an image obtained after image reconstruction of the sample image.
[0098] Specifically, the decoder in the module involved in reconstructing the sample image based on the sample reconstructed visual vector sequence and the encoder in the module involved in vectorizing the sample image correspond to the same network architecture, thereby avoiding the impact on the accuracy of image reconstruction caused by the inconsistency of the network structures corresponding to the decoder and the encoder, thereby avoiding the biased guidance of considering the underlying features of the image when performing quality assessment on the large language model, and further improving the accuracy of image quality assessment.
[0099] Alternatively, the decoder can also be other model structures with image generation capabilities, such as, but not limited to, the decoder in the MAE-ViT (Masked Autoencoder with Vision Transformer) model or the decoder in the DALL-E model. MAE-ViT is a network architecture that combines an autoencoder and a Vision Transformer. It can utilize the autoencoder for image reconstruction tasks and train the Vision Transformer model by performing masking and restoration operations on the image, effectively improving image understanding and feature extraction performance.
[0100] Exemplarily, the above S208 may include: determining at least one loss based on at least one of the sample text vector sequence, the sample visual vector sequence, the sample multimodal vector sequence, the sample reconstructed visual vector sequence, the sample text description, the training image, the sample visual rating value, the probability distribution of the sample visual rating, the sample text rating value, and the probability distribution of the sample text rating, and training the corresponding model or corresponding module involved in the above steps based on the at least one loss.
[0101] Specifically, a first loss can be determined based on the sample text vector sequence; a second loss can be determined based on the training image and the sample image; a third loss can be determined based on the sample visual rating value or the probability distribution of the sample visual rating; a fourth loss can be determined based on the difference between the sample visual rating value and the sample text rating value, or the difference between the probability distribution of the sample visual rating and the probability distribution of the sample text rating; finally, the corresponding model or corresponding module involved in the above steps is trained based on the weighted sum of at least one of the first loss, the second loss, the third loss and the fourth loss.
[0102] Specifically, each loss can be calculated according to formula (4), the target loss can be determined, and the corresponding model or module can be trained according to the target loss.
[0103]
[0104] α, β, and γ are weight coefficients used to balance the proportions of the various losses, adjusting the losses of each task to the same magnitude to maintain training stability. Each task can correspond to one loss and at least one of the above steps, as described below.
[0105] They are respectively the losses corresponding to the text generation task (corresponding to S204 or S202), the image reconstruction task (corresponding to S205), the visual evaluation task (corresponding to S206), and the modal alignment task (corresponding to determining the quality evaluation result for the sample image based on the difference between the sample visual score value and the sample text score value, or the probability distribution of the sample visual score and the probability distribution of the sample text score), which can be specifically the above-mentioned first loss, second loss, third loss, and fourth loss respectively.
[0106] Of course, the fifth loss can also be determined based on the sample text score value or the probability distribution of the sample text score; the sixth loss can be determined based on the sample visual vector sequence; and the target loss can be determined by combining the other losses mentioned above with the fifth and sixth losses. This application does not impose any restrictions on this. The determination methods of the various losses can refer to each other and will not be elaborated in this application.
[0107] against Cross Entropy Loss can be used to supervise the generation effect of sample text vector sequences.
[0108] against The mean squared error loss (MSE Loss) can be used to calculate the distance between the sample image and the training image to obtain the corresponding loss.
[0109] against MSE Loss can be used to supervise the single score, i.e., the sample visual score value, to determine the corresponding loss, or the Earth Mover's Distance Loss (EMD Loss) function can be used to supervise the distribution prediction, i.e., the probability distribution of the sample visual score, to obtain the corresponding loss.
[0110] against If a single score regression scheme is used to determine the corresponding visual score value and text score value, the loss distance between the sample visual score value and the sample text score value can be calculated using MSE Loss through formula (5) to obtain the fourth loss.
[0111]
[0112] Among them, N represents the number of samples, corresponding to the number of scoring values; s v is the sample visual score value, s t Score the sample text with a numerical value.
[0113] If a distribution prediction scheme is used to determine the corresponding probability distribution of visual ratings and the probability distribution of text ratings, the loss distance between the probability distribution of sample visual ratings and the probability distribution of sample text ratings can be calculated using EMD Loss through formula (6) to obtain the fourth loss.
[0114]
[0115] Among them, L represents the number of score distribution intervals; CDF v (i, l) represents the probability value of the visual evaluation prediction cumulative distribution function of the i-th sample in the l-th score interval, that is, the probability distribution of the sample visual score; CDF t (i, l) represents the value of the cumulative distribution function of the text evaluation prediction for the i-th sample in the l-th score interval, that is, the probability distribution of the sample text score. In addition, for the text score determined based on the sample text description, the probability values of the last number of quality descriptors, such as 5, can be selected to represent the probability of the evaluation score falling within each score interval.
[0116] It should be noted that the model or module can be trained using at least one of supervised learning, semi-supervised learning or unsupervised learning, and this application does not impose any restrictions on this. In addition, any type of loss function can be used to determine the above-mentioned losses, and this application does not impose any restrictions on this.
[0117] In the above content, the fourth loss function can be used to align the difference between the sample visual rating value and the sample text rating value, or the probability distribution of the sample visual rating and the probability distribution of the sample text rating, to achieve mutual supervision and correction between visual prediction and text prediction, so as to obtain more accurate quantitative evaluation results. The third loss function can align the sample visual rating value and the corresponding true visual rating value to improve the accuracy of the visual rating. The second loss can align the original image, i.e., the sample image, and the reconstructed image, i.e., the training image, thereby improving the accuracy of image reconstruction and guiding the model to extract, understand and retain more accurate underlying image features. The first loss can guide the model or module to generate a more accurate sample text vector sequence, which helps to generate more accurate text descriptions and ratings. Based on this, the above embodiment can train the model and / or module through multi-task training to achieve more accurate quality assessment in all aspects.
[0118] It is understandable that if only qualitative evaluation results are generated in the form of text descriptions and then indirectly converted into quantitative evaluation results, deviations are likely to occur; moreover, the sample reconstructed visual vector sequence contains rich visual features and can also predict image quality scores, but the scores determined based on visual features and text descriptions have different focuses. The former starts from the perspective of visual features and uses the rich information of the underlying image features retained by the image reconstruction task, while the latter evaluates image quality based on richer context and high-level semantic information. Therefore, the mutual supervision of the visual regression prediction branch and the text prediction branch, that is, determining the loss and training based on the corresponding visual scores and text scores, can improve the sophistication and accuracy of the quantitative evaluation results.
[0119] In one embodiment, the above-mentioned training of the corresponding models or corresponding modules involved in the above steps includes: fixing (Frozen) the parameters of the large language model or fine-tuning the large language model; adjusting some parameters or units in the corresponding modules involved in the above steps (Trainable), and not adjusting some parameters or units, that is, no training or fine-tuning is required (Train-free).
[0120] Exemplarily, the large language model may be fine-tuned (LoRA Train) using a low-rank adaptation (LoRA) method, but is not limited thereto.
[0121] Exemplarily, in combination with the above embodiments, the encoder and decoder can use a network pre-trained on the image reconstruction task, fix the parameters of the encoder and decoder during training, and only optimize the parameters of the first adapter, the second adapter, and the image quality score regression module.
[0122] Through the above content, not only can the training efficiency be improved and the training cost be reduced; moreover, by fine-tuning the large language model, the problem of model generalization caused by the use of its training model due to the relatively concentrated distribution of scores corresponding to images in related image datasets (for example, LIVE-IQA, CSIQ-IQA, TID2013, CLIVE, KonIQ, SPAQ, AVA) can be solved. Specifically, the model or module can be first trained using images in the above image dataset, and then the model or module can be fine-tuned using images other than the images in the image dataset.
[0123] Furthermore, the other images may be images of specific content categories, and the images of specific content categories may be images whose scoring accuracy is relatively obvious to humans (greater than a preset value), such as images of portraits, landscapes, text, etc., so as to solve the problem of inaccurate scoring of the images and enable the model to have evaluation capabilities that are closer to human perception.
[0124] In one embodiment, in combination with the above embodiments, Figure 2 As shown, text prompts (or text prompt words) related to image quality assessment and image reconstruction can be preset. Then, the text prompts are segmented using Tokenizer, and the word embedding module converts the segmented word sequence into a corresponding word embedding vector sequence to obtain a text token sequence "Text Tokens", i.e., a text vector sequence. At the same time, the image feature extractor (Visual Encoder) extracts the visual features of the input image, and the Adaptor converts the visual features so that the visual feature space is aligned with the text feature space, obtaining a visual token sequence "Visual tokens", i.e., a visual vector sequence. Then, the visual token sequence can be embedded into a specific position in the text token sequence to obtain a multimodal token sequence, i.e., a multimodal vector sequence, as input to the large language model. Subsequently, the large language model can generate corresponding visual tokens, i.e., a reconstructed visual vector sequence, based on the text prompts for image reconstruction and determining a visual score, and also generates a text evaluation result, i.e., a text description, describing the image quality. Alternatively, the large language model can also generate a reconstructed text vector sequence based on the multimodal vector sequence, and the text evaluation branch can determine the text description based on the reconstructed text vector sequence (or other modules or models, which are not limited in this application). Afterwards, in the text evaluation branch, the text description and specific text prompts can be combined to guide the large language model to generate words describing the image quality. Based on the confidence of the words, the probability distribution of the score or a single numerical score is determined to obtain the text score value or the probability distribution of the text score. In the visual evaluation branch, the visual tokens generated by the large language model can be used as input, and the visual tokens can be converted into the visual feature space through Deadaptor to obtain image features; then, the image can be reconstructed through the visual decoder. Because in the image reconstruction task, compared with the high-level semantic information that the large language model is good at, low-level visual features are more needed to reconstruct the details in the image, such as lighting, color, and the specific shape of the object. Therefore, the image reconstruction task can guide the model to retain and analyze more low-level information of the image, which is beneficial to image quality assessment. In addition, the above image features can be processed to predict a quality assessment score or score distribution, that is, a visual score value or a probability distribution of visual scores. Among them, the above two scores can be used to supervise and correct the text evaluation branch during the training process; during the reasoning process, the results of the text evaluation branch can be used as the main one, and the score corresponding to the visual evaluation branch can be increased or not.
[0125] in, Figure 2 The Multi-modal Tokens (label) in the training process of the model and / or module refers to the labels used to supervise the tokens generated by the large language model (including Visual Tokens (reconstructed visual vector sequence) and Text Tokens (text description) in the output sequence). It does not actually participate in the forward reasoning of the large language model and is only used in the final loss calculation. Multi-modal Tokens (input) (multimodal input sequence, that is, the multimodal vector sequence mentioned above) is used to uniformly represent the input tokens of the large language model, including Figure 2 The TextToke ns and Visual tokens marked in the input sequence in , these tokens participate in the calculation during the forward reasoning of the large language model, and their values will change.
[0126] Figure 3 A schematic diagram of an image quality assessment device 300 provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the device 300 includes: a first acquisition module 301, a first processing module 302, a sequence fusion module 303, a second processing module 304, a third processing module 305, a dimension conversion module 306, a sample acquisition module 307, a sample processing module 308, and a model training module 309.
[0127] In one embodiment, a first acquisition module 301 is used to acquire an image to be evaluated and a text prompt, where the text prompt instructs a preset large language model to perform quality evaluation on the image to be evaluated; a first processing module 302 is used to vectorize the text prompt to obtain a text vector sequence; and vectorize the image to be evaluated to obtain a visual vector sequence; a sequence fusion module 303 is used to fuse the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence; and a second processing module 304 is used to process the multimodal vector sequence through the large language model to obtain a quality evaluation result of the image to be evaluated in at least one quality evaluation dimension.
[0128] Exemplarily, the first processing module 302 is specifically configured to: segment the text prompt to obtain a word sequence; and convert the word sequence into an embedding vector to obtain a text vector sequence.
[0129] Exemplarily, the first processing module 302 is specifically configured to: extract visual features from the image to be evaluated by using an image feature extractor to obtain a visual vector sequence.
[0130] Exemplarily, the first processing module 302 is specifically used to: extract visual features of the image to be evaluated through an image feature extractor to obtain an initial visual vector sequence; and convert the initial visual vector sequence through a first adapter to obtain a visual vector sequence under the text feature dimension corresponding to the text vector sequence.
[0131] Exemplarily, the sequence fusion module 303 is specifically configured to embed the visual vector sequence into a specific position of the text vector sequence to obtain a multimodal vector sequence.
[0132] Exemplarily, the specific position corresponds to a corresponding position of the image to be evaluated in the text prompt.
[0133] Exemplarily, the second processing module 304 is specifically used to: process the multimodal vector sequence through a large language model to obtain a text description of the image to be evaluated under at least one quality assessment dimension; convert the text description to obtain a text score value or a probability distribution of the text score for the image to be evaluated under at least one quality assessment dimension; and determine the quality assessment result based on the text score value or the probability distribution of the text score.
[0134] Exemplarily, the third processing module 305 is used to: process the multimodal vector sequence through a large language model to obtain a reconstructed visual vector sequence; process the reconstructed visual vector sequence to obtain a visual rating value or a probability distribution of a visual rating for the image to be evaluated under at least one quality assessment dimension; correspondingly, the second processing module 304 is specifically used to: determine the quality assessment result based on the text rating value or the probability distribution of the text rating, and the visual rating value or the probability distribution of the visual rating.
[0135] Exemplarily, the reconstructed visual vector sequence relates to underlying image features of the image to be evaluated.
[0136] Exemplarily, the feature dimension of the visual vector sequence in the multimodal vector sequence is the text feature dimension corresponding to the text vector sequence; the dimension conversion module 306 is used to convert the reconstructed visual vector sequence through the second adapter to obtain a reconstructed visual vector sequence under the visual feature dimension corresponding to the visual vector sequence.
[0137] Exemplarily, the text prompt further instructs the large language model to perform image reconstruction on the image to be evaluated, so that the large language model involves underlying image features of the image to be evaluated when processing the multimodal vector sequence.
[0138] Exemplarily, the sample acquisition module 307 is used to acquire a sample image and a sample text prompt, wherein the sample text prompt includes a first prompt for instructing the large language model to perform quality assessment on the sample image and a second prompt for instructing the large language model to perform image reconstruction on the sample image; the sample processing module 308 is used to vectorize the sample text prompt to obtain a sample text vector sequence; vectorize the sample image to obtain a sample visual vector sequence; fuse the sample text vector sequence and the sample visual vector sequence to obtain a sample multimodal vector sequence; process the sample multimodal vector sequence through the large language model to obtain a sample reconstructed visual vector sequence and a sample text description of the sample image under at least one quality assessment dimension; and perform image reconstruction based on the sample visual vector sequence. The sample image is reconstructed according to the sequence to obtain a training image; the sample visual score value or the probability distribution of the sample visual score of the sample image under at least one quality assessment dimension is determined according to the sample reconstructed visual vector sequence; the sample text score value or the probability distribution of the sample text score of the sample image under at least one quality assessment dimension is determined according to the sample text description; the model training module 309 is used to train the corresponding model or corresponding module involved in the above steps according to at least one of the sample text vector sequence, the sample visual vector sequence, the sample multimodal vector sequence, the sample reconstructed visual vector sequence, the sample text description, the training image, the sample visual score value, the probability distribution of the sample visual score, the sample text score value and the probability distribution of the sample text score.
[0139] Exemplarily, the model training module 309 is specifically used to determine at least one loss based on at least one of the sample text vector sequence, the sample visual vector sequence, the sample multimodal vector sequence, the sample reconstructed visual vector sequence, the sample text description, the training image, the sample visual rating value, the probability distribution of the sample visual rating, the sample text rating value and the probability distribution of the sample text rating, and train the corresponding model or corresponding module involved in the above steps according to the at least one loss.
[0140] Exemplarily, the model training module 309 is specifically used to: determine a first loss based on a sample text vector sequence; determine a second loss based on a training image and a sample image; determine a third loss based on a sample visual rating value or a probability distribution of a sample visual rating; determine a fourth loss based on a difference between a sample visual rating value and a sample text rating value, or a difference between a probability distribution of a sample visual rating and a probability distribution of a sample text rating; and train the corresponding model or corresponding module involved in the above steps based on a weighted sum result of at least one of the first loss, the second loss, the third loss, and the fourth loss.
[0141] Exemplarily, the model training module 309 is specifically used to: fix the parameters of the large language model or fine-tune the large language model; and adjust some parameters or units in the corresponding modules involved in the above steps.
[0142] Exemplarily, the decoder in the module involved in reconstructing the sample image according to the sample reconstructed visual vector sequence and the encoder in the module involved in vectorizing the sample image correspond to the same network architecture.
[0143] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 3 The device 300 shown can execute the above method embodiment, and the above and other operations and / or functions of each module in the device 300 are respectively for implementing the corresponding processes in the above method. For the sake of brevity, they are not repeated here.
[0144] The above describes the device 300 of the embodiment of the present application from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0145] Figure 4 A schematic diagram of an electronic device 400 provided in an embodiment of the present application.
[0146] like Figure 4 As shown, the electronic device 400 may include:
[0147] The memory 410 and the processor 420 are configured to store computer programs and transmit the program code to the processor 420. In other words, the processor 420 can call and run the computer program from the memory 410 to implement the method in the embodiment of the present application.
[0148] For example, the processor 420 may be configured to execute the above method embodiments according to instructions in the computer program.
[0149] In some embodiments of the present application, the processor 420 may include but is not limited to:
[0150] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0151] In some embodiments of the present application, the memory 410 includes but is not limited to:
[0152] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SL DRAM), and direct RAM bus random access memory (DR RAM).
[0153] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 410 and executed by the processor 420 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0154] like Figure 4 As shown, the electronic device may further include:
[0155] The transceiver 430 may be connected to the processor 420 or the memory 410 .
[0156] The processor 420 may control the transceiver 430 to communicate with other devices. Specifically, the processor 420 may send information or data to other devices or receive information or data sent by other devices. The transceiver 430 may include a transmitter and a receiver. The transceiver 430 may further include one or more antennas.
[0157] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0158] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment.
[0159] When software is used to implement, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instruction is loaded and executed on a computer, the computer can be made to perform the corresponding flow in each method in the embodiment of the present application, generate the function that each method in the embodiment of the present application can realize in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website, computer, server, or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, server, or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0160] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the system, device or module can be electrical, mechanical or other forms.
[0162] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.< / level> < / level> < / level> < / level> < / level>
Claims
1. A method for image quality assessment, characterized in that: include: Acquire an image to be evaluated and a text prompt, wherein the text prompt instructs a preset large language model to perform quality evaluation on the image to be evaluated; Performing vectorization processing on the text prompt to obtain a text vector sequence; performing vectorization processing on the image to be evaluated to obtain a visual vector sequence; fusing the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence; The multimodal vector sequence is processed by the large language model to obtain a quality assessment result of the image to be assessed in at least one quality assessment dimension.
2. The method according to claim 1, characterized in that The vectorization processing of the text prompt to obtain a text vector sequence includes: Segmenting the text prompt to obtain a word sequence; The word sequence is converted into an embedding vector to obtain the text vector sequence.
3. The method according to claim 1, characterized in that The vectorization processing of the image to be evaluated to obtain a visual vector sequence includes: The visual feature extractor is used to extract visual features from the image to be evaluated to obtain the visual vector sequence.
4. The method according to claim 3, characterized in that The step of extracting visual features from the image to be evaluated by an image feature extractor to obtain the visual vector sequence includes: Extracting visual features from the image to be evaluated by the image feature extractor to obtain an initial visual vector sequence; The initial visual vector sequence is converted by a first adapter to obtain a visual vector sequence under a text feature dimension corresponding to the text vector sequence.
5. The method according to claim 1, wherein The fusing of the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence includes: The visual vector sequence is embedded into a specific position of the text vector sequence to obtain the multimodal vector sequence.
6. The method according to claim 5, characterized in that The specific position corresponds to a corresponding position of the image to be evaluated in the text prompt.
7. An image quality assessment device, characterized in that: include: A first acquisition module is configured to acquire an image to be evaluated and a text prompt, wherein the text prompt instructs a preset large language model to perform quality assessment on the image to be evaluated; The first processing module is configured to perform vectorization processing on the text prompt to obtain a text vector sequence; and perform vectorization processing on the image to be evaluated to obtain a visual vector sequence; A sequence fusion module, configured to fuse the text vector sequence and the visual vector sequence to obtain a multimodal vector sequence; The second processing module is configured to process the multimodal vector sequence using the large language model to obtain a quality assessment result of the image to be assessed in at least one quality assessment dimension.
8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising instructions, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Unified visual language model pre-training and adjusting method for image quality and aesthetic evaluation
CN118607611A
Attribute Recognition with Image-Conditioned Prefix Language Modeling
US20250054322A1