Model training method, image quality evaluation method and device

By conducting multi-modal large models with multi-stage training and using different sample sets to optimize each module of the model, the problem of inaccurate image quality evaluation in the prior art is solved, and the accurate evaluation of the underlying visual information of the image is achieved.

CN120047791APending Publication Date: 2025-05-27VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109001.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to evaluate the underlying visual information of the image, resulting in inaccurate evaluation of image quality.

Method used

Through the multimodal large model training method, the mapper module, vision encoder module and large language model module in the multimodal large model are trained using the first sample set, the second sample set and the third sample set to generate a fourth multimodal large model that can evaluate the visual information of the underlying image.

Benefits of technology

The accuracy and richness of multimodal large models for image quality descriptions can be improved, and the underlying visual information of the image can be more accurately evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047791A_ABST
    Figure CN120047791A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an image quality evaluation method and a device thereof. Belongs to the artificial intelligence field. The method comprises the steps that based on a first sample set, a mapper module in a first multi-modal large model is trained, a second multi-modal large model is obtained, and the first sample set comprises a first sample image and a short text used for describing the first sample image; based on a second sample set, a visual encoder module and a large language model module in the second multi-modal large model are trained, a third multi-modal large model is obtained, and the second sample set comprises a second sample image and a long text used for describing the second sample image; and based on the third sample set, training the third multi-modal large model to obtain a fourth multi-modal large model which is used for evaluating the bottom visual information of the image. The third sample set includes a third sample image, a first question text about underlying visual information of the third sample image, and a first reply text to the first question text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and specifically to a model training method, an image quality evaluation method, and their devices. Background Art

[0002] The underlying visual information of an image usually refers to some basic and local features of the image, such as features like edges, brightness, noise, contrast, color, and composition. The underlying visual information has a crucial impact on the image quality. When users take pictures and perform post-processing, images with defects in the underlying visual information may be generated, and these defects will directly or indirectly affect the image quality. Evaluating the image quality based on the underlying visual information of the image will be beneficial for the system or the user to further adjust the image or the shooting method to obtain a higher-quality image.

[0003] In the prior art, common image quality evaluation models usually can only output the quality score of the image, and cannot detect the defects in the underlying visual information of the image, making it difficult to provide valuable information. Although emerging multimodal large models can output descriptions about images, they are usually trained using open-source datasets, and the open-source datasets are mainly used for high-level semantic understanding tasks of images. Therefore, the trained models still lack an understanding of the underlying visual information of the image and cannot obtain accurate and rich descriptions of the image quality. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, an image quality evaluation method, and their devices, which can improve the accuracy and richness of the description of the image quality by the multimodal large model.

[0005] In a first aspect, the embodiments of the present application provide a model training method, which includes: training a mapper module in a first multimodal large model based on a first sample set to obtain a second multimodal large model, where the first sample set includes first sample images and short texts for describing the first sample images; training a visual encoder module and a large language model module in the second multimodal large model based on a second sample set to obtain a third multimodal large model, where the second sample set includes second sample images and long texts for describing the second sample images; training the third multimodal large model based on a third sample set to obtain a fourth multimodal large model, where the fourth multimodal large model is used to evaluate the underlying visual information of the image, and the third sample set includes third sample images, first question texts about the underlying visual information of the third sample images, and first reply texts for the first question texts.

[0006] Second aspect, an embodiment of the present application provides a method for image quality assessment, the method comprising: obtaining a to-be-tested image and a question text, where the question text is used to query the fourth multimodal large model about the image quality of the to-be-tested image, and the fourth multimodal large model is trained by using the model training method described in the first aspect; inputting the to-be-tested image and the question text into the fourth multimodal large model to obtain an image quality assessment result of the to-be-tested image output by the fourth multimodal large model, where the image quality assessment result includes a description text of the underlying visual information of the to-be-tested image.

[0007] Third aspect, an embodiment of the present application provides a model training device, the device comprising: a first training unit, configured to train a mapper module in a first multimodal large model based on a first sample set to obtain a second multimodal large model, where the first sample set includes a first sample image and a short text for describing the first sample image; a second training unit, configured to train a visual encoder module and a large language model module in the second multimodal large model based on a second sample set to obtain a third multimodal large model, where the second sample set includes a second sample image and a long text for describing the second sample image; a third training unit, configured to train the third multimodal large model based on a third sample set to obtain a fourth multimodal large model, where the fourth multimodal large model is used to evaluate the underlying visual information of an image, and the third sample set includes a third sample image, a first question text about the underlying visual information of the third sample image, and a first reply text for the first question text.

[0008] Fourth aspect, an embodiment of the present application provides an image quality assessment device, the device comprising: an obtaining unit, configured to obtain a to-be-tested image and a question text, where the question text is used to query the fourth multimodal large model about the image quality of the to-be-tested image, and the fourth multimodal large model is trained by using the model training method described in the first aspect; an image quality assessment unit, configured to input the to-be-tested image and the question text into the fourth multimodal large model to obtain an image quality assessment result of the to-be-tested image output by the fourth multimodal large model, where the image quality assessment result includes a description text of the underlying visual information of the to-be-tested image.

[0009] Fifth aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0010] Sixth aspect, an embodiment of the present application provides a readable storage medium, where a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0011] In a seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the method described in the first aspect.

[0012] In an eighth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0013] In an embodiment of the present application, the first sample set includes first sample images and short texts for describing the first sample images. The second sample set includes second sample images and long texts for describing the second sample images. The third sample set includes third sample images, first question texts about the underlying visual information of the third sample images, and first reply texts for the first question texts. First, the mapper module in the first multimodal large model is trained through the first sample set, so that the trained second multimodal large model can align image features with the text features processed by the large language model module spatially, thereby mapping the image features into the vision-language space applicable to the large language model module. Then, the visual encoder module and the large language model module in the second multimodal large model are trained through the second sample set, so that the trained third multimodal large model has more image visual knowledge, and thus can perceive and describe richer visual information. Finally, the third multimodal large model is trained through the third sample set, enabling the third multimodal large model to learn the reply strategy for questions related to image quality assessment, thereby obtaining a fourth multimodal large model capable of evaluating the underlying visual information of images. Since the first question texts about the underlying visual information of the images and their first reply texts are used for model fine-tuning during the training process of the multimodal large model, the trained multimodal large model can understand the underlying visual information of the images and have the ability to evaluate the underlying visual information, improving the accuracy and richness of the multimodal large model's description of image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flowchart of the model training method provided by an embodiment of the present application;

[0015] Figure 2A is a schematic diagram of the model structure of the multimodal large model provided by an embodiment of the present application;

[0016] Figure 2B is a schematic diagram of the model training process of the model training method provided by an embodiment of the present application;

[0017] Figure 3 is a flowchart of the image evaluation method provided by an embodiment of the present application;

[0018] Figure 4 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;

[0019] Figure 5 It is a schematic structural diagram of an image evaluation device provided by an embodiment of the present application;

[0020] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0021] Figure 7 It is a schematic hardware structure diagram of an electronic device suitable for implementing the embodiment of the present application. Specific embodiments

[0022] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0024] Next, in conjunction with the accompanying drawings, the model training methods and devices provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0025] Please refer to Figure 1 , which shows one of the flowcharts of the model training method provided by the embodiment of the present application. The model training method provided by the embodiment of the present application can be applied to an electronic device. In practice, the above-mentioned electronic device can be an electronic device such as a server.

[0026] The flow of the model training method provided by the embodiment of the present application includes the following steps:

[0027] Step 101: Based on the first sample set, train the mapper module in the first multi-modal large model to obtain a second multi-modal large model, where the first sample set includes first sample images and short texts for describing the first sample images.

[0028] In this embodiment, the first sample set may include multiple samples. Each sample may be a binary tuple, including a first sample image and a short text for describing the first sample image. The short text refers to a text with a short length, specifically, a text with a length less than a set threshold. In practice, the first sample set may adopt open-source data sets such as LLaVA-Pretrain-558K.

[0029] In this embodiment, the first multimodal large model may be a multimodal large model (Multimodal Large Language Model, MLLM) to be trained. The multimodal large model is a type of deep neural network model that can accept information of two or more modalities, such as text and images, as inputs, and can also output information of one or more modalities simultaneously.

[0030] In this embodiment, the structure of the first multimodal large model can be seen Figure 2A as shown, and may include a Vision Encoder module, a mapper module, and a Large Language Model (LLM) module.

[0031] The Vision Encoder module can receive image inputs and encode the input images to convert them into a sequence of discrete tokens, which can be called an image token sequence. A token refers to an independent meaningful unit obtained after dividing the text. These units can be words, characters, sub-words, or other smaller constituent elements, which can be specifically set according to needs. In practice, the Vision Encoder module may adopt a Vision Transformer (ViT) model, specifically, any model trained by contrastive learning of visual language on data can be used. For example, it may include, but is not limited to, models such as clip-vit-large-patch14-336, ViT-SO400M-14-SigLIP-384, etc.

[0032] The data processing process of the visual encoder module may include: First, receive an image input and scale the input image to a preset resolution. The preset resolution varies depending on the model used. For example, if the clip-vit-large-patch14-336 model is adopted, the preset resolution can be 336×336; if the ViT-SO400M-14-SigLIP-384 model is adopted, the preset resolution will be 384×384. Second, cut the input image into image patches of the same size according to the preset patch size. The size of the image patches also varies depending on the model used and is usually 14×14. Third, use a convolutional neural network to map each image patch to an image token. After the image tokens obtained by mapping all the image patches are concatenated, a discrete image token sequence can be formed. Fourth, encode the discrete image token sequence through multiple layers of vision transformer layers in the visual encoder module to obtain a new discrete image token sequence.

[0033] The mapper module can receive the image token sequence output by the visual encoder module, align the image token sequence with the text input space features of the large language model module, so as to map the image token sequence to a visual language discrete token sequence applicable to the large language model module. In practice, the mapper module uses a randomly initialized multilayer perceptron (MLP). A multilayer perceptron is a feedforward artificial neural network composed of multiple layers of neurons, which can include an input layer, one or more hidden layers, and an output layer, etc.

[0034] The large language model module can accept text input and the visual language discrete token sequence output by the mapper module, and perform inference and decoding to output a reply message for the text. In practice, the large language model is a natural language processing model based on deep learning technology, which has a large number of parameters and can process and generate natural language text. The large language model module can learn the structure and grammar of language through a large amount of text data and can perform tasks such as generation, understanding, translation, summarization, and answering questions. Specifically, an open-weight model pre-trained in any way and with language data can be adopted. For example, it can include but is not limited to Qwen2.5-7B, Llama-3.1-8B, Mistral-7B-Instruct-v0.2, etc.

[0035] The large language model module may include a tokenizer, an embedding layer, and a Transformer model. The data processing process of the large language model module may include: First, the input text is divided into discrete sequences by the tokenizer. Second, the discrete sequences are converted into a text discrete token sequence by the embedding layer. Third, according to the order of the user input, the visual language discrete token sequence output by the mapper module and the above text discrete token sequence are concatenated into a multimodal discrete token sequence. Fourth, the multimodal discrete token sequence is input into the Transformer model for decoding prediction to obtain the output text discrete token sequence. Fifth, the above text discrete token sequence is decoded by the tokenizer to obtain the text sequence output by the large language model module.

[0036] In this embodiment, the model training process may include three stages, as shown in Figure 2B In the first stage, the first sample image may be input into the visual encoder module in the first multimodal large model, and the preset prompt information may be input into the large language model module in the first multimodal large model to obtain the predicted text output by the first multimodal large model. The preset prompt information may be used to instruct the first multimodal large model to describe the input first sample image. For example, "Please briefly describe this image in words, etc.". Then, based on the predicted text output by the first multimodal large model and the short text used to describe the first sample image, the loss value of the first multimodal large model is calculated through the loss function. After that, the parameters of the visual encoder module and the large language model module are fixed, and based on this loss value, the parameters of the mapper module are updated using the backpropagation algorithm. After iteratively executing the above process multiple times, a second multimodal large model can be obtained.

[0037] During the training process of this stage, an optimizer, a cosine learning rate warm-up scheduler, and a loss function for next word prediction can be used for training. Among them, the cosine learning rate warm-up scheduler can dynamically adjust the training learning rate according to the training process. For example, the initial learning rate is low, then it increases slowly, and after reaching the maximum value, it slowly decreases to the minimum. The next word prediction loss function can adopt the cross-entropy loss function, which is used to measure the difference between the word probability distribution predicted by the model and the probability distribution of the real word. Its input is the probability distribution of the word sequence predicted by the model and the one-hot encoding vector corresponding to the real word sequence.

[0038] The mapper module in the first multi-modal large model is trained with a first sample set including a first sample image and a short text for describing the first sample image, so that the trained second multi-modal large model can spatially align the image features with the text features processed by the large language model module, thereby mapping the image features into the vision-language space applicable to the large language model module.

[0039] Step 102: Based on the second sample set, the visual encoder module and the large language model module in the second multi-modal large model are trained to obtain a third multi-modal large model. The second sample set includes a second sample image and a long text for describing the second sample image.

[0040] In this embodiment, the second sample set may include multiple samples. Each sample can be a binary tuple, which may include a second sample image and a long text for describing the second sample image. The long text refers to a text with a relatively long length, specifically, a text with a length greater than a set threshold, and this set threshold may be different from the set threshold of the short text. In practice, the first sample set can adopt open-source data sets such as LLaVA-ReCap-CC3M.

[0041] In this embodiment, after obtaining the second multi-modal large model, the second-stage training can be carried out. The second sample image is input into the visual encoder module in the second multi-modal large model, and a preset prompt message is input into the large language model module in the second multi-modal large model to obtain the predicted text output by the second multi-modal large model. The preset prompt message can be used to instruct the first multi-modal large model to describe the input first sample image. For example, "Please use words to describe this image in detail, etc.". Then, based on the predicted text output by the second multi-modal large model and the long text for describing the second sample image, the loss value of the second multi-modal large model is calculated through a loss function. After that, the parameters of the mapper module are fixed, and based on this loss value, the parameters of the visual encoder module and the large language model module are updated using the backpropagation algorithm. After iteratively executing the above process multiple times, the third multi-modal large model can be obtained. It should be noted that the training method and training settings in this stage are basically the same as those in the previous stage, and will not be elaborated here.

[0042] The visual encoder module and the large language model module in the second multi-modal large model are trained with a second sample set including a second sample image and a short text for describing the second sample image, so that the trained third multi-modal large model has more image visual knowledge, and thus can perceive and describe richer visual information.

[0043] Step 103: Based on the third sample set, train the third multi-modal large model to obtain a fourth multi-modal large model, which is used to evaluate the underlying visual information of images. The third sample set includes third sample images, first question texts regarding the underlying visual information of the third sample images, and first response texts for the first question texts.

[0044] In this embodiment, the third sample set may include multiple samples. Each sample can be a triple, which may include a third sample image, a first question text regarding the underlying visual information of the third sample image, and a first response text for the first question text.

[0045] In this embodiment, the underlying layer of an image is the visual layer, and the underlying visual information of the image is the image features of the visual layer, usually some basic and local features. Specifically, it may include but is not limited to blur, exposure, light, contrast, color, noise, sharpness, artifacts, focus, composition, visual style, emotion, overall impression, etc. The first question text can be related to the underlying visual information. For example, it can be "How about the color and composition of this image?" or "How to adjust the brightness of this image?"

[0046] In this embodiment, after obtaining the third multi-modal large model, the third-stage training can be carried out. Specifically, the third sample image can be input into the visual encoder module in the third multi-modal large model, and the first question text corresponding to the third sample image can be input into the large language model module in the third multi-modal large model to obtain the predicted text output by the third multi-modal large model. Then, based on the predicted text output by the third multi-modal large model and the first response text corresponding to the third sample image, the loss value of the third multi-modal large model is calculated through a loss function. After that, based on this loss value, the parameters of the visual encoder module, the transformer, and the large language model module are updated using the backpropagation algorithm. After iteratively executing the above process multiple times, the fourth multi-modal large model can be obtained.

[0047] The training purpose of this stage is to enable the third multi-modal large model to gradually respond according to the input third sample image and the first question text and output a response text. After the training of this stage, the obtained fourth multi-modal large model can output the expected response text according to the question text input by the user. Since the first question text regarding the underlying visual information of the image and its first response text are used for model fine-tuning during the training of the multi-modal large model, the trained multi-modal large model can understand the underlying visual information of the image and have the ability to evaluate the underlying visual information, improving the accuracy and richness of the multi-modal large model's description of the image quality.

[0048] In some alternative implementation manners of this embodiment, step 103 above may further include: training the third multi-modal large model based on the third sample set and the fifth sample set to obtain a fourth multi-modal large model. The fifth sample set includes fifth sample images, second question texts about the fifth sample images, and second reply texts for the second question texts. In practice, the fifth sample set may include, but is not limited to, open-source data sets such as LLaVA-v1.5-mix665k.

[0049] By simultaneously using the third sample set and the fifth sample set for model fine-tuning, the trained fourth multi-modal large model can correctly respond to various open questions, which may include, but are not limited to, questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0050] In some alternative implementation manners of this embodiment, based on the third sample set and the fifth sample set, the following steps may be adopted to train the third multi-modal large model to obtain a fourth multi-modal large model:

[0051] First step, merge the third sample set and the fifth sample set to obtain a sixth sample set.

[0052] Second step, iteratively execute the following training steps: extract a target sample image from the sixth sample set; use the question text corresponding to the target sample image as the target question text, use the reply text for the target question text as the target reply text, input the target sample image and the target question text into the third multi-modal large model to obtain the text to be tested output by the third multi-modal large model; determine the loss value of the third multi-modal large model based on the text to be tested and the target reply text; train based on the third multi-modal large model. The target sample image may be one or more sample images randomly extracted from the sixth sample set.

[0053] Fourth step, when the target condition is met, end the iteration to obtain the fourth multi-modal large model. The target condition can be set as needed. For example, it may be that the number of iterations is equal to a preset number threshold, the training duration is equal to a preset duration threshold, the loss value is lower than the set threshold and tends to converge, etc.

[0054] Through the above method, the parameters of the third multi-modal large model can be fine-tuned so that it can correctly respond to various open questions, which may include, but are not limited to, questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0055] The method provided by the above embodiments of the present application first trains the mapper module in the first multi-modal large model through the first sample set, so that the trained second multi-modal large model can align the image features with the text features processed by the large language model module spatially, thereby mapping the image features into the vision-language space applicable to the large language model module. Then, the visual encoder module and the large language model module in the second multi-modal large model are trained through the second sample set, so that the trained third multi-modal large model has more image visual knowledge, and thus can perceive and describe richer visual information. Finally, the third multi-modal large model is trained through the third sample set, enabling the third multi-modal large model to learn the response strategy for questions related to image quality assessment, thereby obtaining a fourth multi-modal large model capable of evaluating the underlying visual information of the image. Since the first question text and its first response text regarding the underlying visual information of the image are used for model fine-tuning during the training of the multi-modal large model, the trained multi-modal large model can understand the underlying visual information of the image and have the ability to evaluate the underlying visual information, improving the accuracy and richness of the multi-modal large model's description of the image quality.

[0056] In some alternative embodiments, before step 103 is executed, a generation step of the third sample set may also be performed, including:

[0057] Step S11, obtaining a keyword list, where the keyword list includes keywords for describing the visual attributes of the image. The keyword list can be pre-created by technicians.

[0058] Optionally, the keyword list can be created through the following steps: First, obtain the keywords for describing various visual attributes of the image. Then, expand the keywords for various visual attributes to obtain sub-lists of keywords for various visual attributes. Finally, based on the sub-lists of keywords for various visual attributes, generate the keyword list. Here, the sub-lists of keywords can be merged to obtain the keyword list. Thus, a rich keyword list can be obtained, which helps the model learn more expressions.

[0059] Among them, the types of visual attributes may include, but are not limited to: blur, exposure, light, contrast, color, noise, clarity / sharpness, artifact, focus, composition, visual style, sentiment, overall impression, etc. The keywords used to describe the above various visual attributes may be English words. For example, they may include: blur (blur), exposure (exposure), light (light), contrast (contrast), color (color), noise (noise), clarity / sharpness (clarity / sharpness), artifact (artifact), focus (focus), composition (composition), visual style (visual style), sentiment (sentiment), overall impression (overall impression), etc. After expanding the keywords for the above various visual attributes, there may be multiple keywords for each type of visual attribute, and they can be stored in the form of a list to obtain multiple keyword sub-lists. The way of expanding the keywords can be preset as needed, and no specific limitation is made here. The keywords in each keyword sub-list can also be English words.

[0060] As an example, the keyword sub-list corresponding to the blur type of visual attribute may include, but is not limited to, keywords such as "lens blur, glass blur, blurriness, obnubilate, blur, barely legible, barely recognizable, blear, barely visible, blurry, zoom blur, jitter blur, undecipherable, motion blur, gaussian blur, blurred", etc. The above keywords are all synonyms or related words of the keyword "blur", and no further explanation will be given here.

[0061] The keyword list corresponding to the visual attributes of the exposure category may include, but is not limited to, keywords such as "expose, overexposure, vignette, low exposure, overexposure, poorly expose, well expose, multiexposure, under exposure, underexposure, under expose, underexpose, overexpose, over exposed, overexposed, high exposure, exposure, long exposure, over expose, time lapse". Each of the above keywords is a synonym or related word of the keyword "expose", and will not be explained one by one here.

[0062] The keyword list corresponding to the visual attributes of the light category may include, but is not limited to, keywords such as "backlight, darkest, glare, lowlighte, high key, pitch black, non white, light on white, brightest, lowlight, backlighte, low light, undimmed, low lighting, contre jour, back light, use of light, spotlight". Each of the above keywords is a synonym or related word of the keyword "light", and will not be explained one by one here.

[0063] The keyword sub - list of visual attributes related to contrast may include, but is not limited to, keywords such as "medium contrast, high - tone contrast, low - tone contrast, high contrast, low contrast, contrast ratio, contrast adjustment, contrast settings, increase contrast, decrease contrast, adjust contrast, contrast enhancement, contrast filter, contrast effect, contrast control, contrast level, contrast slider, color contrast, image contrast, contrast sensitivity, dynamic contrast", etc. Each of the above keywords is a synonym or related word of the keyword "contrast", and will not be explained one by one here.

[0064] The keyword sub - list of visual attributes of the color category may include, but is not limited to, keywords such as "vibrant, monotonous, harmonize, color diffusion, white and black photograph, vivid, monochrome, complementary color, hdr, complementary colour, monochromatic, vividness, warmcolor, dull, vividity, rich colour, color tone, saturate, harmonious, black andwhite filter, dynamic range, harmonical, harmonise, warm colour, white and blackphoto, cold colour, color balance, saturation, black and white photograph, harmonic, cold color, black and white photo, harmony, color combination, richcolor, chromatic, white and black filter, colour balance, white balance, chroma". Each of the above keywords is a synonym or related word of the keyword "color", and will not be explained one by one here.

[0065] The keyword sub - list of visual attributes of the noise category may include, but is not limited to, keywords such as "salt noise, pepper noise, spatially correlate noise, quantization noise, gaussian noise, impulse noise, noise corruption, speckle noise, poisson noise, image grain, compression noise, film grain, noise point, mosquito noise". Each of the above keywords is a synonym or related word of the keyword "noise", and will not be explained one by one here.

[0066] The keyword sub - list of visual attributes of clarity / sharpness may include, but is not limited to, keywords such as "high resolution, medium sharpness, clearness, clearly visible, clarity, low resolution, high sharpness, clean photo, low sharpness, loss of detail", etc. Each of the above keywords is a synonym or related word of the keyword "clarity / sharpness", and no further explanation will be given here.

[0067] The keyword sub - list of visual attributes of artifacts may include, but is not limited to, keywords such as "gibb phenomenon, blockiness, edge artifact, inter frame distortion, macroblocke, color bleeding, temporal artifact, photographic, pixelation, pixelate, distortion", etc. Each of the above keywords is a synonym or related word of the keyword "artifact", and no further explanation will be given here.

[0068] The keyword sub - list of visual attributes of focus may include, but is not limited to, keywords such as "defocused, defocus, focus, in focus, depth and focus, shallow dof, bokeh, focal, out of focus, soft focus, shallow depth of field, depth of field, out of focused, main focus", etc. Each of the above keywords is a synonym or related word of the keyword "focus", and no further explanation will be given here.

[0069] The keyword sub - list of visual attributes of composition may include, but is not limited to, keywords such as "composition, lead line, main subject, rule of third, horizon line, emphasis, off center, negative space, chaotic, symmetrical, wide angle shoot, main object, golden ratio, symmetry, vertical line", etc. Each of the above keywords is a synonym or related word of the keyword "composition", and no further explanation will be given here.

[0070] The keyword sub - list of visual attributes of the visual style category may include, but is not limited to, keywords such as "natural photography, abstract photo, animation, natural photo, computer generate, abstract image, abstract photography, real picture, real image, abstract picture, realism, naturalimage, surrealism, real photography, real photo, natural picture, ai generate, visual style", etc. Each of the above keywords is a related word of the keyword "visual style", and will not be explained one by one here.

[0071] The keyword sub - list of visual attributes of the emotional category may include, but is not limited to, keywords such as "visually appeal, visualsentiment, visual appeal", etc. Each of the above keywords is a related word of the keyword "sentiment", and will not be explained one by one here.

[0072] The keyword sub - list of overall - impression - type visual attributes may include but is not limited to "aesthetic of this photo, aesthetic, quality of the image, quality of image, aesthetic of the photograph, high quality, aesthetic of the picture, aesthetic of image, wonderful image, poor image, aesthetic of this image, horrible image, poor quality, aesthetic of picture, quality of this picture, aesthetic of the photo, low quality, good quality, quality of photo, nice image, fantastic image, quality of the picture, aesthetic of this photograph, terrible image, image quality, excellent image, excellent picture, quality of this photograph, bad picture, wonderful picture, quality of the photograph, good picture, amazing image, image aesthetic, aesthetic of this picture, awful picture, quality of this image, beautiful picture, aesthetic of photo, nice picture, awful image, aesthetic of the image, fantastic picture, quality of the photo, amazing picture, ugly picture, poor picture, quality of this photo, quality of picture, aesthetic of photograph, acceptable quality, quality of photograph, aesthetic quality, ugly image,Keywords such as "bad quality", "bad image", "great picture", "terrible picture", "horrible picture", "beautiful image", "good image", "great image". Each of the above keywords is a related word of the keyword "semantic-relative", and will not be explained one by one here.,

[0073] Step S12: Extract the target long text containing at least one keyword from the keyword list from the fourth sample set. The fourth sample set includes the fourth sample image and the second long text used to describe the fourth sample image. The target long text can be the second long text containing at least one keyword from the keyword list. The fourth sample set and the above second sample set can be the same sample set or different sample sets, which is not limited here.,

[0074] Step S13: Use the fourth sample image corresponding to the target long text as the third sample image, and generate a prompt message based on the third sample image, the target long text, and the keywords in the target long text.,

[0075] Among them, the prompt message can include a system prompt word and a dialogue prompt word.,

[0076] The system prompt words may include general output requirements for unimodal large models. The unimodal large model can be an open-source large language model. The system prompt words can be preset by technicians according to needs. For example, it can be "You are an AI visual assistant who is an expert in image quality and aesthetics, and you are seeing a single image. The image may be a photograph, a picture on screen, or a painting. What you see are provided with a context, including information about the same image you are looking at. You are mainly interested in low-level visual information, which may include blur, noise, exposure, contrast, clarity / sharpness, artifact, color, lighting, focus, composition of image, visual style of image, visual sentiment of image, perceptual quality or aesthetic of picture, etc. Response as you are seeing the image.", and its meaning is "You are an AI visual assistant, an expert in image quality and aesthetics, and you are looking at a single image. The image can be a photograph, a picture on the screen, or a painting. What you see has a context, including information about the same image you are viewing. You are mainly interested in low-level visual information, which may include blur, noise, exposure, contrast, clarity / sharpness, artifacts, color, lighting, focus, composition of the image, visual style of the image, visual sentiment of the image, perceptual quality or aesthetic of the picture, etc. Respond as you see the image."

[0077] The dialogue prompt words can be used to instruct the unimodal large model to generate Q&A texts about the underlying visual information of the image according to the input content. For example, it can be "I will provide you with a context about an image and keywords that may be related to low-level visual information in the context. You can get information about the image from the context. Design question-and-answer pairs related to low-level visual information between you and a person asking about this photo. The question can be of the type What, How, Which, Does, Is, Where, etc., and you need to randomly choose one from them. The answer should be in a tone that a visual AI assistant is seeing the image and answering the question. Only include questions that have definite answers: (1) one can see the content in the image that the question asks about and can answer confidently; (2) one can determine confidently from the image that it is not in the image. Do not ask any questions that cannot be answered confidently. The answers must be absolutely accurate and consistent with the original context.When you respond, please only output JSON in a list with the key question and answer, no other words are needed. Do not imagine and give irrelevant, groundless, or unsure responses regarding the given context. Keyword: "[[KEYWORD]]" Context: "[[CONTEXT]]". Its meaning is "I will provide you with a context about an image and keywords that may be related to low-level visual information in the context. You can obtain information about the image from the context. Design Q&A pairs related to the low-level visual information between you and the person asking about the photo. The questions can be what, how, who, whether, where, etc., and one needs to be randomly selected from them. The answers should be in the tone of a visual AI assistant seeing the image and answering the question. Only include questions with clear answers: (1) being able to see the content asked in the question in the image and being able to answer confidently; (2) being able to confidently judge from the image that it is not in the image. Do not ask any questions that cannot be answered confidently. The answers must be absolutely accurate and consistent with the original context. When you answer, please only output a JSON list with the key Q&A, no other words are needed. Do not imagine and give responses that are irrelevant, groundless, or unsure regarding the given context. Keyword: "[[KEYWORD]]" Context: "[[CONTEXT]]". Among them, "\n\nKeyword" is used to identify the keywords in the target long text, "\n\nContext" is used to identify the target long text, [[KEYWORD]] is the keyword in the target long text, and [[CONTEXT]] is the target long text.

[0078] In step S14, the prompt information is input into the unimodal large model to obtain the first question text about the low-level visual information of the third sample image and the first response text for the first question text. Here, the unimodal large model can understand and analyze the input prompt information and output the first question text about the low-level visual information of the third sample image and its corresponding first response text.

[0079] In step S15, the third sample image, the first question text, and the first response text are used as a sample and stored in the third sample set.

[0080] By explicitly pointing out the keywords related to the underlying visual information of the image in the prompt message, the relevance between the Q&A text generated by the unimodal large model and the underlying visual information of the image is greatly enhanced. In addition, by leveraging the existing sample set and the unimodal large model to generate a third sample set, high-quality samples can be efficiently obtained, reducing the difficulty of sample acquisition and labor costs.

[0081] Please refer to Figure 3 , which shows one of the flowcharts of the image quality evaluation method provided by the embodiments of the present application. The image quality evaluation method provided by the embodiments of the present application can be applied to electronic devices. In practice, the above-mentioned electronic devices can be servers, smartphones, tablets, laptop computers, wearable devices, and other electronic devices.

[0082] The flow of the image quality evaluation method provided by the embodiments of the present application includes the following steps:

[0083] Step 301, obtain the image to be measured and the question text, where the question text is used to ask the fourth multimodal large model about the image quality of the image to be measured.

[0084] In this embodiment, the fourth multimodal large model is trained using the model training method in the embodiment, which will not be elaborated here. The question text can be input by the user. For example, it can be "Please conduct a detailed analysis and description of the underlying visual information such as image comparison and noise, and obtain the image quality evaluation result for this image." etc.

[0085] Step 302, input the image to be measured and the question text into the fourth multimodal large model, and obtain the image quality evaluation result of the image to be measured output by the fourth multimodal large model. The image quality evaluation result includes a description of the underlying visual information of the image to be measured.

[0086] Since the target multimodal model trained using the Figure 1 corresponding embodiment can describe the image quality more accurately and richly, the accuracy of the image quality evaluation result can be improved through this target multimodal model.

[0087] It should be noted that the image quality evaluation method of this embodiment can be used to test the target multimodal model trained in the above embodiments, and then the target multimodal model can be continuously optimized according to the test results. The image quality evaluation method of this embodiment can also be the actual application method of the target multimodal model trained in the above embodiments. Using the target multimodal model trained in the above embodiments, image quality evaluation can be performed to obtain accurate image quality evaluation results.

[0088] It should be noted that the fourth multimodal large model can be applied to various Q&A scenarios, not limited to the image quality evaluation scenario.

[0089] It should be noted that for the model training method provided in the embodiments of the present application, the execution subject may be a model training device. In the embodiments of the present application, taking the model training device executing the model training method as an example, the model training device provided in the embodiments of the present application is described.

[0090] As Figure 4 shown, the model training device 400 in this embodiment includes: a first training unit 401, configured to train a mapper module in a first multi-modal large model based on a first sample set to obtain a second multi-modal large model, where the first sample set includes first sample images and short texts for describing the first sample images; a second training unit 402, configured to train a visual encoder module and a large language model module in the second multi-modal large model based on a second sample set to obtain a third multi-modal large model, where the second sample set includes second sample images and first long texts for describing the second sample images; a third training unit 403, configured to train the third multi-modal large model based on a third sample set to obtain a fourth multi-modal large model, where the fourth multi-modal large model is used to evaluate the underlying visual information of an image, and the third sample set includes third sample images, first question texts about the underlying visual information of the third sample images, and first reply texts for the first question texts.

[0091] In some optional implementation manners of this embodiment, the device further includes a first generation unit, configured to: obtain a keyword list, where the keyword list includes keywords for describing the visual attributes of an image; extract, from a fourth sample set, target long texts that include at least one keyword in the keyword list, where the fourth sample set includes fourth sample images and second long texts for describing the fourth sample images; use the fourth sample images corresponding to the target long texts as third sample images, and generate prompt information based on the third sample images, the target long texts, and the keywords in the target long texts; input the prompt information into a single-modal large model to obtain first question texts about the underlying visual information of the third sample images and first reply texts for the first question texts; and store the third sample images, the first question texts, and the first reply texts as a sample in the third sample set. By explicitly indicating the keywords related to the underlying visual information of the image in the prompt information, the relevance between the Q&A texts generated by the single-modal large model and the underlying visual information of the image is greatly enhanced. In addition, by leveraging the existing sample set and the single-modal large model to generate the third sample set, high-quality samples can be efficiently obtained, reducing the difficulty of sample acquisition and the labor cost.

[0092] In some alternative implementation manners of this embodiment, the apparatus further includes a second generation unit, configured to: obtain keywords for describing various visual attributes of an image; expand the keywords of the various visual attributes to obtain keyword sub-lists of the various visual attributes; generate a keyword list based on the keyword sub-lists of the various visual attributes. Thereby, a rich keyword list can be obtained, which helps the model learn more expressions.

[0093] In some alternative implementation manners of this embodiment, the third training unit is further configured to: train the third multi-modal large model based on a third sample set and a fifth sample set to obtain a fourth multi-modal large model, where the fifth sample set includes a fifth sample image, a second question text about the fifth sample image, and a second answer text for the second question text. By simultaneously using the third sample set and the fifth sample set for model fine-tuning, the trained fourth multi-modal large model can correctly respond to various open questions, which may include but are not limited to questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0094] In some alternative implementation manners of this embodiment, the third training unit is further configured to: merge the third sample set and the fifth sample set to obtain a sixth sample set; iteratively execute the following training steps: extract a target sample image from the sixth sample set; use the question text corresponding to the target sample image as the target question text, use the answer text for the target question text as the target answer text, input the target sample image and the target question text into the third multi-modal large model to obtain a to-be-tested text output by the third multi-modal large model; determine a loss value of the third multi-modal large model based on the to-be-tested text and the target answer text; train based on the third multi-modal large model; end the iteration and obtain a fourth multi-modal large model when a target condition is met. Through the above method, the parameters of the third multi-modal large model can be fine-tuned so that it can correctly respond to various open questions, which may include but are not limited to questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0095] In the device provided by the above embodiment of the present application, the first sample set includes first sample images and short texts for describing the first sample images, the second sample set includes second sample images and long texts for describing the second sample images, and the third sample set includes third sample images, first question texts about the underlying visual information of the third sample images, and first response texts for the first question texts. First, the mapper module in the first multi-modal large model is trained through the first sample set, so that the trained second multi-modal large model can spatially align the image features with the text features processed by the large language model module, thereby mapping the image features into the vision-language space applicable to the large language model module. Then, the visual encoder module and the large language model module in the second multi-modal large model are trained through the second sample set, so that the trained third multi-modal large model has more image visual knowledge, and thus can perceive and describe richer visual information. Finally, the third multi-modal large model is trained through the third sample set, enabling the third multi-modal large model to learn the response strategy for questions related to image quality assessment, thereby obtaining a fourth multi-modal large model capable of evaluating the underlying visual information of images. Since the first question texts about the underlying visual information of the images and their first response texts are used for model fine-tuning during the training of the multi-modal large model, the trained multi-modal large model can understand the underlying visual information of the images and have the ability to evaluate the underlying visual information, improving the accuracy and richness of the multi-modal large model's description of image quality.

[0096] It should be noted that for the image quality assessment method provided by the embodiment of the present application, the execution subject may be an image quality assessment device. In the embodiment of the present application, taking the image quality assessment device executing the image quality assessment method as an example, the image quality assessment device provided by the embodiment of the present application is described.

[0097] As Figure 5 shown, the model training device 500 in this embodiment includes: an acquisition unit 501, configured to acquire a to-be-tested image and a question text, where the question text is used to ask the fourth multi-modal large model about the image quality of the to-be-tested image; an image quality assessment unit 502, configured to input the to-be-tested image and the question text into the fourth multi-modal large model to obtain an image quality assessment result of the to-be-tested image output by the fourth multi-modal large model, where the image quality assessment result includes a description text of the underlying visual information of the to-be-tested image.

[0098] It can be understood that the units described in the device 500 correspond to the respective steps in the above image quality assessment method. Therefore, the operations, features, and beneficial effects described above for the image quality assessment method also apply to the device 500 and the units included therein, and will not be elaborated herein.

[0099] The model training device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an Augmented Reality (AR) / Virtual Reality (VR) device, a robot, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a Personal Computer (PC), a Television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0100] The model training device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0101] The model training device provided in the embodiments of the present application can implement Figure 1 or Figure 3 each process implemented by the method embodiments described above. To avoid repetition, it will not be elaborated here.

[0102] Optionally, as Figure 6 shown, the embodiments of the present application further provide an electronic device 600, including a processor 601 and a memory 602. A program or instruction that can run on the processor 601 is stored on the memory 602. When the program or instruction is executed by the processor 601, it implements each step of the above model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0103] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0104] Figure 7 Schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0105] The electronic device 700 includes, but is not limited to, components such as a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710.

[0106] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 710 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 7 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0107] Among them, the processor 710 is used to train the mapper module in the first multi-modal large model based on the first sample set to obtain a second multi-modal large model. The first sample set includes a first sample image and a short text for describing the first sample image; based on the second sample set, train the visual encoder module and the large language model module in the second multi-modal large model to obtain a third multi-modal large model. The second sample set includes a second sample image and a first long text for describing the second sample image; based on the third sample set, train the third multi-modal large model to obtain a fourth multi-modal large model. The fourth multi-modal large model is used to evaluate the underlying visual information of the image. The third sample set includes a third sample image, a first question text about the underlying visual information of the third sample image, and a first reply text for the first question text.

[0108] Training the mapper module in the first multi-modal large model with the first sample set enables the trained second multi-modal large model to spatially align image features with the text features processed by the large language model module, thereby mapping the image features into the vision-language space applicable to the large language model module. Training the visual encoder module and the large language model module in the second multi-modal large model with the second sample set enables the trained third multi-modal large model to possess more image visual knowledge, thereby being able to perceive and describe richer visual information. Training the third multi-modal large model with the third sample set enables the third multi-modal large model to learn response strategies for questions related to image quality assessment, thereby obtaining a fourth multi-modal large model capable of evaluating the underlying visual information of images. Since the first question text and its first response text regarding the underlying visual information of images are used for model fine-tuning during the training of the multi-modal large model, the trained multi-modal large model can understand the underlying visual information of images and possess the ability to evaluate the underlying visual information, improving the accuracy and richness of the multi-modal large model's description of image quality.

[0109] Optionally, the processor 710 is further configured to obtain a keyword list, where the keyword list includes keywords for describing the visual attributes of an image; extract, from the fourth sample set, a target long text containing at least one keyword in the keyword list, where the fourth sample set includes fourth sample images and second long texts for describing the fourth sample images; use the fourth sample images corresponding to the target long text as the third sample images, and generate a prompt message based on the third sample images, the target long text, and the keywords in the target long text; input the prompt message into the single-modal large model to obtain a first question text regarding the underlying visual information of the third sample images and a first response text for the first question text; and store the third sample images, the first question text, and the first response text as a sample in the third sample set. By explicitly indicating the keywords related to the underlying visual information of the image in the prompt message, the relevance between the question-answer text generated by the single-modal large model and the underlying visual information of the image is greatly enhanced. In addition, by leveraging the existing sample set and the single-modal large model to generate the third sample set, high-quality samples can be efficiently obtained, reducing the difficulty of sample acquisition and the labor cost.

[0110] Optionally, the processor 710 is further configured to obtain keywords for describing various visual attributes of an image; expand the keywords for the various visual attributes to obtain keyword sub-lists for the various visual attributes; and generate a keyword list based on the keyword sub-lists for the various visual attributes. Thus, a rich keyword list can be obtained, which helps the model learn more expressions.

[0111] Optionally, the processor 710 is further configured to train the third multi-modal large model based on the third sample set and the fifth sample set to obtain a fourth multi-modal large model. The fifth sample set includes fifth sample images, second question texts about the fifth sample images, and second response texts for the second question texts. By simultaneously using the third sample set and the fifth sample set for model fine-tuning, the trained fourth multi-modal large model can correctly respond to various open questions, including but not limited to questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0112] Optionally, the processor 710 is further configured to merge the third sample set and the fifth sample set to obtain a sixth sample set; iteratively execute the following training steps: extract target sample images from the sixth sample set; use the question text corresponding to the target sample image as the target question text, use the response text for the target question text as the target response text, input the target sample image and the target question text into the third multi-modal large model to obtain the text to be tested output by the third multi-modal large model; determine the loss value of the third multi-modal large model based on the text to be tested and the target response text; train based on the third multi-modal large model; end the iteration when the target condition is met to obtain the fourth multi-modal large model. In the above manner, the parameters of the third multi-modal large model can be fine-tuned so that it can correctly respond to various open questions, including but not limited to questions related to image quality assessment, enriching the functions and usage scenarios of the fourth multi-modal large model.

[0113] In addition, the processor 710 can also be used to obtain a to-be-tested image and a question text, where the question text is used to ask the fourth multi-modal large model about the image quality of the to-be-tested image; input the to-be-tested image and the question text into the fourth multi-modal large model to obtain the image quality assessment result of the to-be-tested image output by the fourth multi-modal large model, and the image quality assessment result includes a description text of the underlying visual information of the to-be-tested image.

[0114] It should be understood that in the embodiments of the present application, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes the image data of static pictures or videos obtained by an image capturing device (such as a camera) in a video capturing mode or an image capturing mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also referred to as a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0115] The memory 709 can be used to store software programs and various data. The memory 709 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 can include a volatile memory or a non-volatile memory, or the memory 709 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0116] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 710 either.

[0117] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0118] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk, or optical disc, etc.

[0119] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0120] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0121] The embodiments of the present application provide a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned model training method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0122] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0123] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0124] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A model training method, characterized in that: The method comprises: Based on a first sample set, training a mapper module in a first multimodal large model to obtain a second multimodal large model, wherein the first sample set includes a first sample image and a short text for describing the first sample image; Based on the second sample set, training the visual encoder module and the large language model module in the second multimodal large model to obtain a third multimodal large model, wherein the second sample set includes a second sample image and a first long text for describing the second sample image; Based on the third sample set, the third multimodal large model is trained to obtain a fourth multimodal large model, wherein the fourth multimodal large model is used to evaluate the underlying visual information of the image, and the third sample set includes a third sample image, a first question text about the underlying visual information of the third sample image, and a first answer text to the first question text.

2. The method according to claim 1, characterized in that Before training the third multimodal large model based on the third sample set, the method further includes: Obtaining a keyword list, wherein the keyword list includes keywords for describing visual attributes of an image; extracting a target long text containing at least one keyword in the keyword list from a fourth sample set, wherein the fourth sample set includes a fourth sample image and a second long text for describing the fourth sample image; using a fourth sample image corresponding to the target long text as a third sample image, and generating prompt information based on the third sample image, the target long text, and keywords in the target long text; Inputting the prompt information into the unimodal large model to obtain a first question text about the underlying visual information of the third sample image and a first answer text for the first question text; The third sample image, the first question text and the first answer text are taken as a sample and stored in a third sample set.

3. The method according to claim 1, characterized in that The training of the third multimodal large model based on the third sample set to obtain a fourth multimodal large model includes: The third multimodal large model is trained based on the third sample set and the fifth sample set to obtain a fourth multimodal large model, wherein the fifth sample set includes a fifth sample image, a second question text regarding the fifth sample image, and a second answer text to the second question text.

4. The method according to claim 3, characterized in that The training of the third multimodal large model based on the third sample set and the fifth sample set to obtain a fourth multimodal large model includes: Combining the third sample set and the fifth sample set to obtain a sixth sample set; Iteratively perform the following training steps: extract a target sample image from the sixth sample set; use the question text corresponding to the target sample image as the target question text, use the answer text for the target question text as the target answer text, input the target sample image and the target question text into the third multimodal large model, and obtain the text to be tested output by the third multimodal large model; determine the loss value of the third multimodal large model based on the text to be tested and the target answer text; and perform training based on the third multimodal large model; When the target conditions are met, the iteration is terminated to obtain the fourth multimodal large model.

5. A method for evaluating image quality, characterized in that: The method comprises: Acquire an image to be tested and a question text, wherein the question text is used to inquire about the image quality of the image to be tested of a fourth multimodal large model, wherein the fourth multimodal large model is trained using the model training method according to any one of claims 1 to 4; The image to be tested and the question text are input into the fourth multimodal large model to obtain an image quality assessment result of the image to be tested output by the fourth multimodal large model, wherein the image quality assessment result includes a descriptive text of the underlying visual information of the image to be tested.

6. A model training device, characterized in that: The device comprises: A first training unit is used to train a mapper module in a first multimodal large model based on a first sample set to obtain a second multimodal large model, wherein the first sample set includes a first sample image and a short text for describing the first sample image; a second training unit, configured to train the visual encoder module and the large language model module in the second multimodal large model based on a second sample set to obtain a third multimodal large model, wherein the second sample set includes a second sample image and a first long text for describing the second sample image; A third training unit is used to train the third multimodal large model based on a third sample set to obtain a fourth multimodal large model, wherein the fourth multimodal large model is used to evaluate underlying visual information of an image, and the third sample set includes a third sample image, a first question text about the underlying visual information of the third sample image, and a first answer text to the first question text.

7. The device according to claim 6, characterized in that The device also includes a first generating unit, configured to: Obtaining a keyword list, wherein the keyword list includes keywords for describing visual attributes of an image; extracting a target long text containing at least one keyword in the keyword list from a fourth sample set, wherein the fourth sample set includes a fourth sample image and a second long text for describing the fourth sample image; using a fourth sample image corresponding to the target long text as a third sample image, and generating prompt information based on the third sample image, the target long text, and keywords in the target long text; Inputting the prompt information into the unimodal large model to obtain a first question text about the underlying visual information of the third sample image and a first answer text for the first question text; The third sample image, the first question text and the first answer text are taken as a sample and stored in a third sample set.

8. The device according to claim 6, characterized in that The third training unit is further used for: The third multimodal large model is trained based on the third sample set and the fifth sample set to obtain a fourth multimodal large model, wherein the fifth sample set includes a fifth sample image, a second question text regarding the fifth sample image, and a second answer text to the second question text.

9. The device according to claim 8, characterized in that The third training unit is further used for: Combining the third sample set and the fifth sample set to obtain a sixth sample set; Iteratively perform the following training steps: extract a target sample image from the sixth sample set; use the question text corresponding to the target sample image as the target question text, use the answer text for the target question text as the target answer text, input the target sample image and the target question text into the third multimodal large model, and obtain the text to be tested output by the third multimodal large model; Determining a loss value of the third multimodal large model based on the text to be tested and the target answer text; Training based on the third multimodal large model; When the target conditions are met, the iteration is terminated to obtain the fourth multimodal large model.

10. A picture quality assessment device, characterized in that: The device comprises: an acquisition unit, used for acquiring an image to be tested and a question text, wherein the question text is used for inquiring a fourth multimodal large model about the image quality of the image to be tested, wherein the fourth multimodal large model is trained by the model training method according to any one of claims 1 to 4; An image quality assessment unit is used to input the image to be tested and the question text into the fourth multimodal large model to obtain an image quality assessment result of the image to be tested output by the fourth multimodal large model, wherein the image quality assessment result includes a description text of the underlying visual information of the image to be tested.