Quality evaluation method and system for AI generated image
By fine-tuning the local multimodal large model and using the multimodal large language model to analyze the quality, alignment and authenticity of the images generated by AI, it solves the problem of difficulty in effectively evaluating the quality of AI-generated images in the prior art, and achieves high accuracy and interpretability evaluation results.
Patent Information
- Application Number
- CN202411950382.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively evaluate the quality of AI-generated images, especially in terms of alignment and logical authenticity of image content with generated prompt words.
A quality evaluation method for AI-generated images is proposed. By fine-tuning the local multimodal large model, the multimodal large language model is used to analyze the quality, alignment and authenticity of the image, and a detailed evaluation description is generated.
A comprehensive, fine-grained and interpretable evaluation of the quality of AI-generated image is achieved, and the accuracy and interpretability of evaluation results are improved.
Smart Images

Figure CN119941645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a quality assessment method and system for AI-generated images. Background Art
[0002] The quality assessment of natural images usually considers factors such as image clarity, color, lighting, composition, and artistry. In addition to the above factors, the quality assessment of AI-generated images also needs to consider factors such as the alignment of image content and the cue words used to generate it, and whether the image is logical. Traditional image quality assessment algorithms usually construct image and user score datasets and perform regression predictions on trained neural networks. These algorithms focus on the local or global details of the image itself and can better handle the quality assessment of the image itself, but it is difficult to effectively evaluate the additional cue word alignment and authenticity involved in AI-generated images. Therefore, some studies have begun to model the degree of cue word alignment based on the CLIP algorithm.
[0003] However, most existing solutions focus on using the CLIP algorithm to calculate the similarity between the prompt words used when AI generates images and the image itself, and then use this similarity score as the basis for evaluation; at the same time, some solutions use the partial prompting capabilities of the text encoder in the CLIP algorithm to perform a graded evaluation of the quality of AI-generated images, but the text prompting capabilities in CLIP are limited and cannot obtain more detailed evaluation results. Therefore, a comprehensive, effective, fine-grained and explainable AI-generated image quality evaluation method is needed.
[0004] In view of this, the present invention is proposed. Summary of the invention
[0005] The present invention proposes a quality assessment method and system for AI-generated images. The present invention can effectively analyze the AI-generated images from multiple perspectives based on the characteristics that distinguish them from ordinary images, and provide reliable assessment results.
[0006] In a first aspect, the present invention proposes a method for fine-tuning a local multimodal large model, comprising: S11. A data set E is preset, and the data set E consists of n pictures P i , for the picture P i The corresponding description word T i And the picture P i Score S i Composition, wherein the picture P i is formed by i The corresponding description word T i Generated by AI, i=1, 2, ..., n, for the image P i Score Si Including the picture P i The quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui; S12. The picture P i , for the picture P i The corresponding description word T i And the picture P i Score S i Write the constructed prompt word template T d The prompt word tdi is obtained from the prompt word template T d for: Please analyze this picture with the prompt word T i Image P generated by AI i ; Its quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui. The score range is 0-N, and the higher the score, the better; Please return your analysis results in JSON format, which contains the following three properties: quality_explanation: describes the overall quality of the image in terms of structure, clarity, visible defects, etc. alignment_explanation: describes the alignment between the content in the image and the cue words used, specifically which elements are present and which are missing; authenticity_explanation: Describe how similar the image is to an authentic image. Please indicate any parts of the image that are unofficial or have obvious artifacts. Please do not include specific ratings for this type of image in your response. Here is a sample response: { "quality_explanation": "The overall quality of the image is very good, including...", "alignment_explanation": "The alignment between the image and the prompt word is good, among which...", "authenticity_explanation": "The image doesn't look real, like something out of a dream..." } Among them, the prompt word template T d Where N is an integer greater than 0, qi, ai, aui are all real numbers belonging to 0-N; S13. The picture P i and the prompt word tdi are input into the online multimodal large model A to obtain the image Pi The text description Tqi of the quality qi of the image P i The text description Tai of the alignment degree ai, the image P i The authenticity of aui's textual description of Taui; S14. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the quality qi is written into the constructed prompt word template Tqi cq The prompt word tcqi is obtained from cq for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image quality. The prompt word used for its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 11>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 12>; S15. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the alignment degree ai is written as the prompt word template T constructed by Tai ca The prompt word tcai is obtained from ca for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image alignment. The prompt word used to create it is T i , according to this prompt word T i You can see from this picture P i What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 13>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 14>; S16. The picture P i , for the picture P i The corresponding description word T i And the picture P i Authenticity aui text description Taui input build prompt word template T cau The prompt word tcaui is obtained from cau for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 15>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 16>; Among them, in steps S14-S16, <Reply 11>, <Reply 13>, <Reply 15> are the analysis results given by the online multimodal large model A, and <Reply 12>, <Reply 14>, <Reply 16> are the analysis results given by the online multimodal large model A according to the image P. i The real scores in the data set E are generated by intervals; S17. Constructing a dialogue dataset C q , C a , C au , the dialogue dataset C q , C a , C au The data in the dialog dataset C are derived from the prompt words tcqi, tcai, and tcaui. q , C a , C au The data in the image P are respectively split into training sets Trq, Tra, Trau and validation sets Vq, Va, Vau, and the local multimodal large model B is fine-tuned until the local multimodal large model B is able to adapt to the image P i The analysis results of the text description of quality qi, alignment ai, and authenticity aui converge on the validation sets Vq, Va, and Vau, respectively, and the best results are achieved on the validation sets Vq, Va, and Vau, respectively.
[0007] Furthermore, the online multimodal large model A described in S13 of this embodiment can be any one of GPT-4V and Gemini.
[0008] Furthermore, the local multimodal large model B in S17 of this embodiment is MiniCPM v2.5.
[0009] Furthermore, in this embodiment S17, the local multimodal large model B can be fine-tuned using any one of LoRA, AdaLoRA, Adapter-Tuning, Prefix Tuning, and Prompt Tuning.
[0010] In a second aspect, the present invention proposes a quality assessment method for AI-generated images, which adopts the method of fine-tuning the local multimodal large model as described in the first aspect, including: S21. A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; S22. A fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1, and obtain the data that is consistent with the dialogue data set C q , C a , C au Similar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; S23. Construct texts t1cqj, t1caj, and t1cauj, wherein the texts t1cqj, t1caj, and t1cauj respectively intercept the first three dialogues in the prompt words tcqj, tcaj, and tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P jWhat do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; S24. The text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final encoding of S25. The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final encoding of S26. The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis results auj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final encoding of S27. Distribute the logits to Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P jPredicted scores for image quality qj, alignment aj, and authenticity auj.
[0011] Furthermore, in S24-S27, the mean square error loss function (MSE) and the mean absolute error (MAE) are used to update the learnable parameters in the xLSTM and MLP regression layers using gradient descent.
[0012] In a third aspect, the present invention provides a quality assessment system for AI-generated images, which adopts a quality assessment method for AI-generated images as described in the second aspect, including: D31. Dataset acquisition module: A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; D32. Prompt word acquisition module: a fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1 to obtain the data that is consistent with the dialogue data set C q , C a , C au Similar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is Tj , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; D33. Constructing text module: Constructing text t1cqj, t1caj, t1cauj, wherein the text t1cqj, t1caj, t1cauj respectively intercepts the first 3 items of the dialogue in the prompt words tcqj, tcaj, tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; D34.logits distribution final encoding acquisition module: the text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final encoding of The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final encoding of The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis results auj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final encoding of D35. Image prediction and scoring module: distribute the logits to Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P j Predicted scores for image quality qj, alignment aj, and authenticity auj.
[0013] In a fifth aspect, the present invention proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a method for fine-tuning a local multimodal large model or a method for evaluating the quality of an AI-generated image as described above.
[0014] In a sixth aspect, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for fine-tuning a local multimodal large model or a method for quality assessment of AI-generated images as described above.
[0015] Compared with the prior art, the present invention has the following beneficial effects: Compared with the traditional AI-generated image quality assessment algorithm, the AI-generated image assessment method based on a multimodal large language model proposed in the present invention can obtain more accurate assessment results, thereby improving the accuracy and interpretability of the AI-generated image assessment results. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification are used to explain the principles of the present invention. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0017] Figure 1 It is an example schematic diagram of the online large model proposed in the present invention to generate a text description reflecting the quality, alignment, and authenticity of an AI-generated image.
[0018] Figure 2 This is an example schematic diagram of fine-tuning the local multimodal large model proposed in the present invention.
[0019] Figure 3 It is a flow chart of the quality assessment method of AI generated images proposed by the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.
[0022] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0023] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0024] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the commodity or device including the elements.
[0025] The optional embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0026] In a first aspect, the present invention proposes a method for fine-tuning a local multimodal large model, comprising: S11. A data set E is preset, and the data set E consists of n pictures P i , for the picture P i The corresponding description word T i And the picture P i Score S i Composition, wherein the picture P i is formed by i The corresponding description word T i Generated by AI, i=1, 2, ..., n, for the image P i Score S i Including the picture P iThe quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui; S12. The picture P i , for the picture P i The corresponding description word T i And the picture P i Score S i Write the constructed prompt word template T d The prompt word tdi is obtained from the prompt word template T d for: Please analyze this picture with the prompt word T i Image P generated by AI i ; Its quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui. The score range is 0-N, and the higher the score, the better; Please return your analysis results in JSON format, which contains the following three properties: quality_explanation: describes the overall quality of the image in terms of structure, clarity, visible defects, etc. alignment_explanation: describes the alignment between the content in the image and the cue words used, specifically which elements are present and which are missing; authenticity_explanation: Describe how similar the image is to an authentic image. Please indicate any parts of the image that are unofficial or have obvious artifacts. Please do not include specific ratings for this type of image in your response. Here is a sample response: { "quality_explanation": "The overall quality of the image is very good, including...", "alignment_explanation": "The alignment between the image and the prompt word is good, among which...", "authenticity_explanation": "The image doesn't look real, like something out of a dream..." } Among them, the prompt word template T d Where N is an integer greater than 0, qi, ai, aui are all real numbers belonging to 0-N; S13. The picture P i and the prompt word tdi are input into the online multimodal large model A to obtain the image P i The text description Tqi of the quality qi of the image Pi The text description Tai of the alignment degree ai, the image P i The authenticity of aui's textual description of Taui; Among them, since the prompt word tdi gives the AI generated image P i The text description given by the online multimodal large model A will be close to human preferences based on the actual human subjective quality score, otherwise the online multimodal large model A will tend to give a positive evaluation result; S14. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the quality qi is written into the constructed prompt word template Tqi cq The prompt word tcqi is obtained from cq for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image quality. The prompt word used for its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 11>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 12>; S15. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the alignment degree ai is written as the prompt word template T constructed by Tai ca The prompt word tcai is obtained from ca for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image alignment. The prompt word used to create it is T i , according to this prompt word T i You can see from this picture P i What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 13>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 14>; S16. The picture P i , for the picture P i The corresponding description word T i And the picture P i Authenticity aui text description Taui input build prompt word template T cau The prompt word tcaui is obtained from cau for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 15>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 16>; Among them, in steps S14-S16, <Reply 11>, <Reply 13>, <Reply 15> are the analysis results given by the online multimodal large model A, and <Reply 12>, <Reply 14>, <Reply 16> are the analysis results given by the online multimodal large model A according to the image P. i The real scores in the data set E are generated by intervals; S17. Constructing a dialogue dataset C q , C a , C au , the dialogue dataset C q , C a , C au The data in the dialog dataset C are derived from the prompt words tcqi, tcai, and tcaui. q , C a , C au The data in the image P are respectively split into training sets Trq, Tra, Trau and validation sets Vq, Va, Vau, and the local multimodal large model B is fine-tuned until the local multimodal large model B is able to adapt to the image P iThe analysis results of the text description of quality qi, alignment ai, and authenticity aui converge on the validation sets Vq, Va, and Vau, respectively, and the best results are achieved on the validation sets Vq, Va, and Vau, respectively.
[0027] like Figure 1 As shown, in this example, the image quality assessment of a picture is taken as an example for explanation. The image alignment assessment and image authenticity assessment of a picture are similar to this and will not be described in detail here.
[0028] The prompt word T i , the picture P i , for the picture P i Score S i Enter the constructed prompt word template T d Get the prompt word tdi: Tip word T i It is: "a bicycle on a boat"; The picture P i is the i “A bicycle on a boat” generated image; For the picture P i Score S i Yes: "Its quality score qi is 3.0, the prompt word alignment score ai is 3.0, and the logical authenticity score aui is 2.6. The score range is 0-5, and the higher the score, the better; For the picture P i Scoring requirements and sample responses are as follows: Please return your analysis results in JSON format, which contains the following three properties: quality_explanation: describes the overall quality of the image in terms of structure, clarity, visible defects, etc. alignment_explanation: describes the alignment between the content in the image and the cue words used, specifically which elements are present and which are missing; authenticity_explanation: Describe how similar the image is to an authentic image. Please indicate any parts of the image that are unofficial or have obvious artifacts. Please do not include a specific rating for the image in your response. Here is a sample response: { "quality_explanation": "The overall quality of the image is very good, including...", "alignment_explanation": "The alignment between the image and the prompt word is good, among which...", "authenticity_explanation": "The image doesn't look real, like something out of a dream..." }.
[0029] In this embodiment, the picture P i and the prompt word tdi are input into the online multimodal large model A to obtain the image P i The text descriptions Tqi, Tai, and Taui of the quality qi, alignment ai, and authenticity aui are: "quality_explanation": "The overall quality of the image is very good. It is well focused and well lit. There is a lot of detail on the bike and boat. The composition is simple but effectively shows the unusual placement of the bike." "alignment_explanation": "The image perfectly reflects the content of the prompt. It accurately expresses a bicycle on a boat without any ambiguity or error in the expression of the prompt." "authenticity_explanation": "This image looks like a real photo rather than being artificially modified or AI-generated. Some of its details, such as the rust on the bicycle and the stripes on the wooden boards on the boat, are very real. There are no signs of artificial modification or AI generation in the image."
[0030] In this embodiment, the picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the quality qi is written into the constructed prompt word template Tqi cq Get the prompt word tcqi: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image quality. The prompt word used for its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 11>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 12>; Where <Response 11> is the analysis result given by the online multimodal large model A, and <Response 12> is the analysis result given by the online multimodal large model A based on the image P i The true scores in dataset E are generated in intervals. For example, if image P i The score is 3.5 points. In the 5-point system, 3.5 belongs to the range of 3 to 4 points, which can be judged as "good".
[0031] In this embodiment, the fine-tuned local multimodal large model B gives an analysis result on the image quality of the picture Pi as follows: Figure 2 shown.
[0032] Furthermore, the online multimodal large model A described in S13 of this embodiment can be any one of GPT-4V and Gemini.
[0033] Furthermore, the local multimodal large model B in S17 of this embodiment is MiniCPM v2.5.
[0034] Furthermore, in this embodiment S17, the local multimodal large model B can be fine-tuned using any one of LoRA, AdaLoRA, Adapter-Tuning, Prefix Tuning, and Prompt Tuning.
[0035] Second, as Figure 3 As shown, the present invention proposes a quality assessment method for AI-generated images, which adopts the method of fine-tuning the local multimodal large model as described in the first aspect, including: S21. A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; S22. A fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1, and obtain the data that is consistent with the dialogue data set C q , C a , C auSimilar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; S23. Construct texts t1cqj, t1caj, and t1cauj, wherein the texts t1cqj, t1caj, and t1cauj respectively intercept the first three dialogues in the prompt words tcqj, tcaj, and tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; S24. The text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final encoding of S25. The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final encoding of S26. The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis resultsauj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final encoding of S27. Distribute the logits to Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P j Predicted scores for image quality qj, alignment aj, and authenticity auj.
[0036] Among them, since in generating the dialogue dataset C1 q 、C1 a 、C1 au In the process, two rounds of dialogue were used, which included the image P j The text analysis results from a certain angle of image quality, alignment, and authenticity can be seen as an enhancement of the original data, that is, the image itself and the prompt words used to generate the image. At the same time, in the process of generating dialogues, meaningful logits distribution (probability of output words) is also obtained, which can be seen as the distribution of AI-generated images and their prompt words in the word table space, providing a basis for subsequent scoring predictions.
[0037] Furthermore, in S24-S27, the mean square error loss function (MSE) and the mean absolute error (MAE) are used to update the learnable parameters in the xLSTM and MLP regression layers using gradient descent.
[0038] In the present invention, firstly, a large multimodal language model is used to process the data, and the capability of the stronger online large multimodal model is distilled into a local large multimodal model.
[0039] Traditional AI-generated image quality assessment algorithms train image feature extraction networks, which have limited performance in feature extraction and can only handle the quality assessment of the image itself while ignoring the quality assessment standards unique to AI-generated images.
[0040] The present invention uses a multi-modal large language model (MLLM), which retains the powerful contextual understanding ability of the large language model while introducing the ability to analyze and understand images. Prompt words are constructed according to the images and their scores in the data set, and a mature online multi-modal large model is used to generate multi-angle evaluation descriptions of AI-generated images. Subsequently, the local multi-modal large language model is fine-tuned using the evaluation descriptions of the AI-generated images generated by the online multi-modal large model so that the local multi-modal large language model acquires its capabilities.
[0041] Compared with traditional AI-generated image quality assessment algorithms, the use of a multimodal large language model can analyze the features of an image and express them in the form of natural language. It can also provide a basis for subsequent reasoning of the local multimodal large model, thereby utilizing the reasoning capabilities of the multimodal large model to obtain more accurate evaluation results.
[0042] Secondly, the present invention uses a local multimodal large model for image understanding and reasoning.
[0043] The existing AI-generated image quality assessment method using a multimodal large model only uses the image analysis capability of the multimodal large model and directly obtains the assessment results, eliminating the intermediate process.
[0044] This solution uses a local multimodal large model to generate evaluation texts for various aspects of AI-generated images through prompt words. With the help of the thinking chain ability of the large language model, the local multimodal large model can fully "think" before making a decisive assessment of the quality of AI-generated images, thereby improving the accuracy and interpretability of the evaluation results.
[0045] Finally, the sequence output of the multimodal large model is used for training, which is aligned with human perception.
[0046] In the present invention, although the local multimodal large model has been fine-tuned, when generating a dialogue dataset, the generated responses still have a certain gap from the factual scores. Therefore, an enhanced long short-term memory network xLSTM is introduced in the present invention. The sequence representation ability of the enhanced long short-term memory network xLSTM is used to align the output of the local multimodal large model with human preferences.
[0047] The large language model has a strong generative ability, but studies have pointed out that the next token prediction probability of the large language model usually deviates from the correct answer. Therefore, the present invention regards the multimodal large model as a joint encoder of images and texts. The inference intermediate process generated by the locally fine-tuned multimodal large model is used as the context, and then the AI-generated image itself is sent to the multimodal large model to obtain the logits distribution, and then the logits are combined with the sequence model and regression module to finally obtain the prediction score.
[0048] In a third aspect, the present invention provides a quality assessment system for AI-generated images, which adopts a quality assessment method for AI-generated images as described in the second aspect, including: D31. Dataset acquisition module: A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; D32. Prompt word acquisition module: a fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1 to obtain the data that is consistent with the dialogue data set C q , C a , C au Similar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; D33. Constructing text module: Constructing text t1cqj, t1caj, t1cauj, wherein the text t1cqj, t1caj, t1cauj respectively intercepts the first 3 items of the dialogue in the prompt words tcqj, tcaj, tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; D34.logits distribution final encoding acquisition module: the text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final encoding of The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final encoding of The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis results auj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final encoding of D35. Image prediction and scoring module: distribute the logits to Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P j Predicted scores for image quality qj, alignment aj, and authenticity auj.
[0049] In a fifth aspect, the present invention proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a method for fine-tuning a local multimodal large model or a method for evaluating the quality of an AI-generated image as described above.
[0050] In a sixth aspect, the present invention proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for fine-tuning a local multimodal large model or a method for quality assessment of AI-generated images as described above.
[0051] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program code stored in the storage medium.
[0052] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0053] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0054] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0055] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fine-tuning a local multimodal large model, characterized in that: include: S11. A data set E is preset, and the data set E consists of n pictures P i , for the picture P i The corresponding description word T i And the picture P i Score S i Composition, wherein the picture P i is formed by i The corresponding description word T i Generated by AI, i=1, 2, ..., n, for the picture P i Score S i Including the picture P i The quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui; S12. The picture P i , for the picture P i The corresponding description word T i And the picture P i Score S i Write the constructed prompt word template T d The prompt word tdi is obtained from the prompt word template T d for: Please analyze this picture with the prompt word T i Image P generated by AI i ; Its quality score is qi, the prompt word alignment score is ai, and the logical authenticity score is aui. The score range is 0-N, and the higher the score, the better; Please return your analysis results in JSON format, which contains the following three properties: quality_explanation: describes the overall quality of the image in terms of structure, clarity, visible defects, etc. alignment_explanation: describes the alignment between the content in the image and the cue words used, specifically which elements are present and which are missing; authenticity_explanation: Describe how similar the image is to an authentic image. Please indicate any parts of the image that are unofficial or have obvious artifacts. Please do not include specific ratings for this type of image in your response. Here is a sample response: { "quality_explanation": "The overall quality of the image is very good, including...", "alignment_explanation": "The alignment between the image and the prompt word is good, among which...", "authenticity_explanation": "The image doesn't look real, like something out of a dream..." } Among them, the prompt word template T d Where N is an integer greater than 0, qi, ai, aui are all real numbers belonging to 0-N; S13. The picture P i and the prompt word tdi are input into the online multimodal large model A to obtain the image P i The text description Tqi of the quality qi of the image P i The text description Tai of the alignment degree ai, the image P i The authenticity of aui's textual description of Taui; S14. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the quality qi is written into the constructed prompt word template Tqi cq The prompt word tcqi is obtained from cq for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image quality. The prompt word used for its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 11>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your rating: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 12>; S15. The picture P i , for the picture P i The corresponding description word T i And the picture P i The text description of the alignment degree ai is written as the prompt word template T constructed by Tai ca The prompt word tcai is obtained from ca for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image alignment. The prompt word used to create it is T i , according to this prompt word T i You can see from this picture P i What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 13>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 14>; S16. The picture P i , for the picture P i The corresponding description word T i And the picture P i Authenticity aui text description Taui input build prompt word template T cau The prompt word tcaui is obtained from cau for: User: Please carefully analyze this AI-generated image P i , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T i , according to this prompt word T i You can see from this picture P i What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 15>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 16>; Among them, in steps S14-S16, <Reply 11>, <Reply 13>, <Reply 15> are the analysis results given by the online multimodal large model A, and <Reply 12>, <Reply 14>, <Reply 16> are the analysis results given by the online multimodal large model A according to the image P. i The real scores in the data set E are generated by intervals; S17. Constructing a dialogue dataset C q , C a , C au , the dialogue dataset C q , C a , C au The data in the dialog dataset C are derived from the prompt words tcqi, tcai, and tcaui. q , C a , C au The data in the image P are respectively split into training sets Trq, Tra, Trau and validation sets Vq, Va, Vau, and the local multimodal large model B is fine-tuned until the local multimodal large model B is able to adapt to the image P i The analysis results of the text description of quality qi, alignment ai, and authenticity aui converge on the validation sets Vq, Va, and Vau, respectively, and the best results are achieved on the validation sets Vq, Va, and Vau, respectively.
2. A method for fine-tuning a local multimodal large model according to claim 1, characterized in that: The online multimodal large model A described in S13 can be any one of GPT-4V and Gemini.
3. The method for fine-tuning a local multimodal large model according to claim 1, characterized in that: The local multimodal large model B described in S17 is MiniCPM v2.
5.
4. The method for fine-tuning a local multimodal large model according to claim 1, characterized in that: In S17, the local multi-modal large model B can be fine-tuned using any of LoRA, AdaLoRA, Adapter-Tuning, Prefix Tuning, and Prompt Tuning.
5. A method for evaluating the quality of AI-generated images, using the method for fine-tuning a local multimodal large model as claimed in claim 1, characterized in that: include: S21. A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; S22. A fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1, and obtain the data that is consistent with the dialogue data set C q , C a , C au Similar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your rating: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; S23. Construct texts t1cqj, t1caj, and t1cauj, wherein the texts t1cqj, t1caj, and t1cauj respectively intercept the first three dialogues in the prompt words tcqj, tcaj, and tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your rating: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; S24. The text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final code of S25. The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final code of S26. The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis results auj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final code of S27. Distribute the logits Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P j Predicted scores for image quality qj, alignment aj, and authenticity auj.
6. The quality assessment method of AI-generated images according to claim 5, characterized in that: In S24-S27, the mean square error (MSE) loss function and the mean absolute error (MAE) are used to update the learnable parameters in the xLSTM and MLP regression layers using gradient descent.
7. A quality assessment system for AI-generated images, using the quality assessment method for AI-generated images as claimed in claim 5, comprising: D31. Dataset acquisition module: A data set E1 is preset, and the data set E1 consists of n1 pictures P j , for the picture P j The corresponding description word T j And the picture P j Score S j Composition, wherein the picture P j is formed by j The corresponding description word T j Generated by AI, i=1, 2, ..., n1, for the picture P j Score S j Including the picture P j The quality score is qj, the prompt word alignment score is aj, and the logical authenticity score is auj; D32. Prompt word acquisition module: a fine-tuned local multimodal large model B1 is preset to infer the data in the data set E1 to obtain the data that is consistent with the dialogue data set C q , C a , C au Similar dialogue dataset C1 q 、C1 a 、C1 au , where the conversation dataset C1 q 、C1 a 、C1 au , the data in the prompt words are from the prompt words tcqj, tcaj, tcauj, and the prompt word template T1 corresponding to the prompt words tcqj, tcaj, tcauj cq 、T1 ca 、T1 cau The details are as follows: The prompt word template T1 cq for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your rating: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; AI Assistant: <Reply 22>; Wherein, <Response 21> and <Response 22> are generated by the local multimodal large model B1; The prompt word template T1 ca for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; AI Assistant: <Reply 24>; Wherein, <Response 23> and <Response 24> are generated by the local multimodal large model B1; The prompt word template T1 cau for: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; AI Assistant: <Reply 26>; Wherein, <Response 25> and <Response 26> are generated by the local multimodal large model B1; D33. Constructing text module: Constructing text t1cqj, t1caj, t1cauj, wherein the text t1cqj, t1caj, t1cauj respectively intercepts the first 3 items of the dialogue in the prompt words tcqj, tcaj, tcauj, as follows: The text t1cqj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image quality. The prompt word used for its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall quality? AI Assistant: <Reply 21>; User: Based on your analysis, how would you rate the quality of this image? Please select one from the following list as your rating: [very poor, poor, average, good, very good], where the words in the list represent the quality from low to high; Wherein, <Response 21> is generated by the local multimodal large model B1; The text t1caj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image alignment. The prompt word used to create it is T j , according to this prompt word T j You can see from this picture P j What do you infer from the above? How do you assess its overall alignment? AI Assistant: <Reply 23>; User: Based on your analysis, how would you rate the alignment of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the alignment from low to high; Wherein, <Response 23> is generated by the local multimodal large model B1; The text t1cauj is: User: Please carefully analyze this AI-generated image P j , and gives the analysis results based on its image authenticity. The prompt word used in its creation is T j , according to this prompt word T j You can see from this picture P j What do you infer from it? How do you assess its overall truth? AI Assistant: <Reply 25>; User: Based on your analysis, how do you rate the authenticity of this image? Please select one from the following list as your evaluation result: [very poor, poor, average, good, very good], where the words in the list represent the authenticity from low to high; Wherein, <Response 25> is generated by the local multimodal large model B1; D34.logits distribution final encoding acquisition module: the text t1cqj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image quality gives the logits distribution Q of the analysis results qj , the logits distribution Q qj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q qj The final code of The text file t1caj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image alignment gives the logits distribution Q of the analysis results alj , the logits distribution Q alj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q alj The final code of The text t1cauj and the picture P j Input the local multimodal large model B1 to obtain the image P j Image authenticity gives the logits distribution Q of the analysis results auj , the logits distribution Q auj The input is input into the enhanced long short-term memory network xLSTM for sequence encoding, and the feature extraction output average pooling result of the xLSTM is taken as the logits distribution Q auj The final code of D35. Image prediction and scoring module: distribute the logits to Q qj The final encoding, logits distribution Q alj The final encoding, logits distribution Q auj The final encodings of are respectively input into the regression layer composed of the multi-layer perceptron MLP to obtain the image P j Predicted scores for image quality qj, alignment aj, and authenticity auj.
8. An electronic device, characterized in that , including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for fine-tuning a local multimodal large model as described in claim 1 or a method for quality assessment of an AI-generated image as described in claim 5 is implemented.
9. A computer-readable storage medium, characterized in that , a computer program is stored thereon, which, when executed by a processor, implements a method for fine-tuning a local multimodal large model as described in claim 1 or a method for quality assessment of an AI-generated image as described in claim 5.
Citation Information
Cited By
Quality evaluation method and system for image content reasoning enhancement in mine limited environment
CN120526429A
A Method and System for AIGI Aesthetic Evaluation Based on Multimodal Large Model
CN122574482A