An art painting analysis method based on a multimodal large model

By integrating visual encoder and large language model in multimodal large models and performing supervised fine-tuning, the problems of insufficient analysis capabilities and language bias in art painting analysis are solved, and higher quality art painting analysis is achieved.

CN119580055BActive Publication Date: 2025-06-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410941788.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-06-10
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

The existing multimodal large models lack comprehensive and in-depth analysis in artistic painting analysis, and are prone to erroneous text generation due to language deviation problems caused by visual hallucinations.

Method used

By obtaining images from the art painting library and filtering and annotating data using closed-source large language models, a multimodal large model, including visual encoder, multimodal projector and large language model, supervised fine-tuning, and obtaining an art painting generative pre-trained model.

Benefits of technology

The ability of multimodal large models in artistic painting analysis has been significantly improved, the language deviation problems caused by visual hallucinations have been reduced, the quality of generated analytical texts has been improved, and the generalization ability of the model in artistic painting analysis tasks has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580055B_ABST
    Figure CN119580055B_ABST
Patent Text Reader

Abstract

The present invention discloses an art painting analysis method based on a multimodal large model. First, images are obtained from an art painting library, and a closed-source large language model is used to screen out images that have both a title and the name of the artist. Then, the corresponding overall art analysis paragraphs that only focus on visual features are annotated, and annotations for art analysis at the professional level such as composition, color, light and shadow, etc. are made, thereby constituting art painting analysis data. Using the collected art painting analysis data, a multimodal large model is trained and fine-tuned to obtain a generative pre-trained model for art painting analysis, which is used for the generation of art painting analysis. Through experimental research, the present invention significantly improves the art painting analysis ability of the multimodal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of art painting analysis, and more specifically, relates to an art painting analysis method based on a multimodal large model. Background Art

[0002] Art painting analysis is an important part of art appreciation and creation, and usually requires a deep understanding of visual composition, artistic techniques, style, and historical context. In the past decade, with the great success of deep learning in computer vision and natural language processing, artificial intelligence for art painting analysis has developed rapidly.

[0003] Early work explored the classification and recognition of artistic paintings, relying mainly on handcrafted features. With the great success of pre-trained language models, visual language models were fine-tuned using image-text pairs, achieving some improvements in style classification, object detection, and multimodal retrieval of artistic paintings. Later, researchers proposed multiple datasets for the task of visual question answering of art, and introduced external knowledge and bias elimination strategies to allow visual language models to learn patterns for question answering of artistic paintings, achieving good performance.

[0004] Although these simple tasks have been studied in the era of deep learning, the models still lack the ability to conduct a comprehensive and in-depth analysis of art paintings due to the limitations of visual perception and language model generation capabilities. It is also expected that AI systems will be able to provide comprehensive analysis of art paintings, which will benefit art education and help humans write reviews.

[0005] In recent years, large language models and multimodal large models have demonstrated excellent understanding and generation capabilities in many fields, including text summarization, open-ended question answering, and contextual reasoning. These models can accurately perceive and understand visual content and generate detailed descriptions, such as visual storytelling and subtitle generation, showing the potential for comprehensive analysis of artistic paintings.

[0006] Large language models and multimodal large models have promoted progress in many research fields. However, although existing multimodal large models have strong natural image understanding and text generation capabilities, they still cannot accurately perform comprehensive and in-depth analysis of art paintings, including color, composition, lines, shapes, light and shadow based on visual understanding. In addition, existing multimodal large models face the problem of language bias caused by visual illusions, and are prone to mistakenly identify the current visual content as other, and then the language model generates language for the misidentified visual content based on the corresponding memory knowledge to obtain incorrect text. Therefore, for the task of art painting analysis, existing models tend to first recognize the given painting and then perform the corresponding analysis, while not paying attention to the visual content of the painting in the generation stage of the analysis text. This process of first recognition and then analysis depends heavily on the accuracy of recognition, and it is easy to fail when the given painting is unknown and does not exist in the model memory. Although the existing ShareGPT4V provides a dataset with high-quality image and text description pairs to fine-tune the open source LLaVA multimodal model to alleviate the language bias problem caused by visual illusions, it still cannot perform professional art analysis on art paintings. Summary of the invention

[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an artistic painting analysis method based on a multimodal large model.

[0008] To achieve the above-mentioned object of the invention, the present invention provides an artistic painting analysis method based on a multimodal large model, characterized by comprising the following steps:

[0009] (1) Data collection for art painting analysis

[0010] 1.1) Get an image from the art painting library and use a closed-source large language model to determine whether the title and artist name of the image, i.e. the art painting, are known at the same time. If so, keep it, otherwise discard it;

[0011] 1.2) Use two closed-source large language models to retrieve the learned knowledge through the title of the art painting and the name of the artist, delete the title of the art painting and the name of the artist, and each generate a paragraph of analysis text that only focuses on visual features;

[0012] 1.3) Extract the overall art analysis paragraph from the two generated analysis texts that only focus on visual features, and analyze and annotate them to obtain the corresponding command text. Extract the five most important professional level analysis paragraphs from the two generated analysis texts that only focus on visual features, select the intersection of the five professional level analysis paragraphs as the selected professional level analysis paragraph, and analyze and annotate them to obtain the command text of the selected professional level. In this way, high-quality art painting analysis data consisting of the image of the art painting, the overall analysis paragraph and the corresponding command text, and the analysis paragraphs at each professional level and the corresponding command text are obtained, wherein the professional level includes composition, light and shadow, color, shape, texture, symbol and icon, perspective, movement and posture, line quality, and scale ratio;

[0013] (2) Train to obtain a generative pre-trained model for art paintings

[0014] 2.1) Construct a multimodal large model consisting of a visual encoder, a multimodal projector, and a large language model. The visual encoder encodes the image of the art painting to obtain visual features. The multimodal projector uses a multi-layer perceptron to project the visual features into the semantic space of the language. The large language model decodes the visual features projected into the semantic space of the language in combination with the command text to obtain an overall analysis paragraph or a professional level analysis paragraph.

[0015] 2.2) Use the collected art painting images and corresponding art painting analysis data to supervisely fine-tune the multimodal large model: freeze the parameters of the visual encoder, fine-tune the multimodal projector and the large language model, set the learning rate of training fine-tuning to 2e-5, set the batch size to 16, and perform 10k steps of fine-tuning to obtain the art painting generative pre-training model;

[0016] (3) Reasoning

[0017] Input an art painting image I, a command text P containing n words for the instruction model to analyze, and output an analysis paragraph Output containing m words through the art painting generative pre-trained model.

[0018] The object of the present invention is achieved in this way.

[0019] Since there is currently no dataset for art painting analysis, the art painting analysis method based on a multimodal large model in the present invention first obtains images from an art painting library and uses a closed-source large language model to screen out images that have both a title and the name of the artist. Then, it annotates the corresponding overall art analysis paragraphs that only focus on visual features, and annotates art analysis at professional levels such as composition, color, light and shadow, etc., thus constituting art painting analysis data. Using the collected art painting analysis data, a multimodal large model is trained and fine-tuned to obtain a generative pre-trained model for art painting, which is used for the generation of art painting analysis. Through experimental research, the present invention significantly improves the art painting analysis ability of the multimodal large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a specific implementation manner of the art painting analysis method based on a multimodal large model in the present invention;

[0021] Figure 2 is Figure 1 a flowchart of a specific implementation manner of data collection for art painting analysis in

[0022] Figure 3 is a specific example diagram of the instruction command example and dialogue ability of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following describes the specific implementation manner of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0024] The present invention focuses on art painting analysis and innovatively proposes an art painting analysis method based on a multimodal large model to improve the art painting perception ability of the model and understand the art techniques and professional levels demonstrated by visual elements. Since there is currently no dataset for art painting analysis, in this embodiment, the present invention collected approximately 19k painting images, annotated the corresponding overall art analysis paragraphs that only focus on visual features according to the title of the painting and the name of the artist, and annotated art analysis at specific levels such as composition, color, light and shadow, etc., thus constituting 50k pieces of art painting analysis data. Using the collected dataset, the present invention trained and fine-tuned a multimodal large model to obtain a generative pre-trained model for art painting, abbreviated as GalleryGPT, which can be used for the generation of art painting analysis. Through experimental research, the present invention significantly improves the art painting analysis ability of the multimodal large model.

[0025] Figure 1 is a flowchart of a specific implementation manner of the art painting analysis method based on a multimodal large model in the present invention.

[0026] In this embodiment, as Figure 1 shown, the method for analyzing artistic paintings based on a multimodal large model of the present invention includes the following steps:

[0027] Step S1: Data collection for artistic painting analysis

[0028] The present invention aims to alleviate the language bias problem caused by visual hallucinations by fine-tuning a multimodal large model to pay more attention to the visual elements of paintings. Since there is currently no dataset for artistic painting analysis, the present invention designs a data collection method to obtain high-quality data for artistic painting analysis, so as to fine-tune and train the multimodal large model to obtain an artistic painting generative pre-trained model, namely GalleryGPT.

[0029] However, manually annotating the analysis of artistic paintings requires annotators to have professional knowledge in art analysis, which is difficult and expensive, and it takes a lot of time to manually annotate a large number of artistic paintings. Therefore, considering that excellent closed-source large language models have sufficient knowledge, the present invention uses them to perform high-quality analysis and annotation on artistic painting images. The finally collected dataset includes artistic paintings and the artistic painting analysis data for each artistic painting. The specific data collection process is as Figure 2 shown, including the following steps:

[0030] Step S1.1: Obtain artistic painting images with titles and artist names

[0031] Obtain images from an artistic painting library, and use a closed-source large language model to determine whether it knows both the title and the artist name of the image, that is, the artistic painting. If it knows, keep it; otherwise, discard it.

[0032] With the development of the Internet, a large number of artistic paintings have been digitized and stored. In this embodiment, the present invention uses the Art Gallery as the source of paintings. To ensure that the closed-source large language models GPT-4 and Gemini can provide accurate analysis paragraphs of artistic paintings, first obtain 19,295 famous artistic painting images in the Art Gallery, and use the closed-source large language model Gemini to determine whether it knows the titles and artist names of these artistic paintings, so as to filter out paintings without specific titles and those judged as unknown by Gemini. Finally, 18,526 painting images are obtained. Among them, 13,526 are used as training images, and 5,000 relatively non-famous paintings are used as test images.

[0033] Step S1.2: Generate two paragraphs of analysis text that only focus on visual features

[0034] Use two closed-source large language models to retrieve learned knowledge respectively through the title and artist name of the art painting, delete the title and artist name of the art painting, and generate a paragraph of analysis text that only focuses on visual features for each.

[0035] In this embodiment, the present invention uses powerful closed-source large language models GPT-4 and Gemini to generate high-quality art analysis paragraphs. For each painting, only the title and artist name of the painting are provided, without inputting any visual information. The closed-source large language models GPT-4 and Gemini retrieve learned knowledge through the title and artist name of the painting, delete the title and artist name of the art painting, and generate a paragraph of analysis text that only focuses on visual features for each. To avoid the problem of language deviation caused by visual hallucinations, it is required that the closed-source large language models GPT-4 and Gemini do not mention the title and artist name of the art painting in the analysis text, so that the corresponding art painting works cannot be easily identified based on the specific analysis text, that is, to avoid the problem of language deviation caused by visual hallucinations.

[0036] Step S1.3: Generate analysis paragraphs for both overall art analysis and professional-level analysis

[0037] To make the art painting analysis data more diverse, generate analysis paragraphs in two aspects: extract the overall art analysis paragraph from the two generated paragraphs of analysis text that only focus on visual features, and perform analysis annotation to obtain the corresponding command text. Extract the most important 5 professional-level analysis paragraphs from the two generated paragraphs of analysis text that only focus on visual features respectively, select the intersection of the 5 professional-level analysis paragraphs as the selected professional-level analysis paragraphs, and perform analysis annotation to obtain the command text of the selected professional-level. In this way, high-quality art painting analysis data composed of the image of the art painting, the overall analysis paragraph and the corresponding command text, each professional-level analysis paragraph and the corresponding command text is obtained. Among them, the professional levels include composition, lighting, color, shape, texture, symbol and icon, perspective, movement and gesture, line quality, scale and proportion.

[0038] Step S2: Train to obtain an art painting generative pre-trained model

[0039] After collecting art painting analysis data, inspired by the multi-modal large model ShareGPT4V, in this embodiment, an art painting generative pre-trained model GalleryGPT is constructed and fine-tuned, which can analyze art painting works with a focus on visual elements.

[0040] Step S2.1: Construct a multi-modal large model

[0041] Construct a multimodal large model consisting of a visual encoder, a multimodal projector, and a large language model. The visual encoder encodes the image of the art painting to obtain visual features. The multimodal projector projects the visual features into the semantic space of language using a multi-layer perceptron. The large language model decodes the visual features projected into the language semantic space in combination with the command text to obtain the overall analysis paragraph or the professional-level analysis paragraph.

[0042] The present invention aims to make the open-source multimodal large model more focused on visual elements for analyzing art paintings. In this embodiment, an open-source multimodal large model ShareGPT4V-7B is used as the blueprint and the base model to construct a multimodal large model, which mainly consists of three parts: (1) The CLIP-Large visual encoder, the size of the input image is converted to 336×336, and then converted into 576 input tokens; (2) The multimodal projector, which projects the visual features into the semantic space of language using a multi-layer perceptron; (3) The Vicuna-v1.5 large language model, which adopts the Transformer decoder architecture.

[0043] Step S2.2: Use the collected art painting images and the corresponding art painting analysis data to perform supervised fine-tuning on the multimodal large model to obtain the art painting generative pre-training model.

[0044] In order to utilize the superior visual perception and description capabilities of the ShareGPT4V-7B base model, in this embodiment, the collected art painting images and the corresponding art painting analysis data are used to perform supervised fine-tuning on the multimodal large model: freeze the parameters of the CLIP-Large visual encoder, fine-tune the multimodal projector and the Vicuna-v1.5 large language model, set the learning rate of the fine-tuning training to 2e-5, set the batch size to 16, and perform 10k steps of fine-tuning to obtain the art painting generative pre-training model, which can perceive the art elements and techniques in the painting and output the corresponding description of the art analysis, that is, the analysis paragraph.

[0045] Step S3: Inference

[0046] Input an art painting image I and a command text P containing n words for the model to analyze, and output an analysis paragraph Output containing m words through the art painting generative pre-training model GalleryGPT.

[0047] In this embodiment, the command text can be expressed as P = {w 1 , w 2 , …, w n}, w iFor the i-th word among them, where i = 1, 2, …, n, the text analysis Output = {o 1 , o 2 , …, o m} is generated by the art painting generative pre-trained model GalleryGPT, where o j is the j-th word among them, and j = 1, 2, …, m.

[0048] Experimental results:

[0049] The present invention collects data for art painting analysis and fine-tunes to obtain the art painting generative pre-trained model GalleryGPT, which can generate an analysis focusing on visual elements for art paintings and alleviate the language deviation problem caused by visual hallucinations.

[0050] To evaluate the effectiveness of the art painting analysis method based on the multimodal large model of the present invention, in the experiment, first, experiments were conducted on the analysis test data of 5000 art paintings. Several popular open-source multimodal large models were also tested, including LLaVA-1.5, Qwen-VL-Chat, and ShareGPT4V. All multimodal models used the same instruction command text: "Please write a coherent art analysis for this painting." The present invention uses subtitle description evaluation metrics BLEU, GLEU, METEOR, and ROUGE to evaluate the quality of the generated analysis text.

[0051]

[0052] Table 1

[0053] Table 1 shows the experimental results on the art painting analysis test data. As shown in Table 1, the present invention is superior to all other open-source multimodal large models. Benefiting from the supervised fine-tuning of art painting analysis data, significant improvements have been achieved compared to the base model, which verifies the effectiveness of high-quality art analysis data collection and fine-tuning of multimodal large models.

[0054] To verify the classification and question-answering capabilities of the art painting generative pre-trained model GalleryGPT on art paintings, the present invention conducts test experiments on several existing art painting style classification and question-answering datasets.

[0055]

[0056] Table 2

[0057] Table 2 shows the experimental results on the art painting style classification and question-answering test datasets. As can be seen from Table 2, the art painting generative pre-trained model GalleryGPT of the present invention is also significantly superior to all baseline models, demonstrating its generalization ability for downstream art painting analysis tasks.

[0058] Figure 3 Examples of the instruction commands and dialogue capabilities of the present invention are shown. These examples show that the GalleryGPT model can follow the instructions for generating analysis of artistic paintings. Specifically, the GalleryGPT model can give an overall analysis of the input painting, the necessary level analysis, and the artistic style. In short, these qualitative examples show that the GalleryGPT model of the present invention has superior artistic painting analysis capabilities and illustrate that the collected dataset for artistic painting analysis has high quality.

[0059] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A method for analyzing artistic paintings based on a multimodal large model, characterized by comprising the following steps: (1) Data collection for art painting analysis 1.1) Get an image from the art painting library and use a closed-source large language model to determine whether the title and artist name of the image, i.e. the art painting, are known at the same time. If so, keep it, otherwise discard it; 1.2) Use two closed-source large language models to retrieve the learned knowledge through the title of the art painting and the name of the artist, delete the title of the art painting and the name of the artist, and each generate a paragraph of analysis text that only focuses on visual features; 1.3) Extract the overall art analysis paragraph from the two generated analysis texts that only focus on visual features, and analyze and annotate them to obtain the corresponding command text. Extract the five most important professional level analysis paragraphs from the two generated analysis texts that only focus on visual features, select the intersection of the five professional level analysis paragraphs as the selected professional level analysis paragraph, and analyze and annotate them to obtain the command text of the selected professional level. In this way, high-quality art painting analysis data consisting of the image of the art painting, the overall analysis paragraph and the corresponding command text, and the analysis paragraphs at each professional level and the corresponding command text are obtained, wherein, The professional aspects mentioned include composition, light and shadow, color, shape, texture, symbol and icon, perspective, movement and gesture, line quality, scale and proportion; (2) Train to obtain a generative pre-trained model for art paintings 2.1) Construct a multimodal large model consisting of a visual encoder, a multimodal projector, and a large language model. The visual encoder encodes the image of the art painting to obtain visual features. The multimodal projector uses a multi-layer perceptron to project the visual features into the semantic space of the language. The large language model decodes the visual features projected into the semantic space of the language in combination with the command text to obtain an overall analysis paragraph or a professional level analysis paragraph. 2.2) Use the collected art painting images and corresponding art painting analysis data to supervisely fine-tune the multimodal large model: freeze the parameters of the visual encoder, fine-tune the multimodal projector and the large language model, set the learning rate of training fine-tuning to 2e-5, set the batch size to 16, and perform 10k steps of fine-tuning to obtain the art painting generative pre-training model; (3) Reasoning Input an art painting image I, a command text P containing n words for the instruction model to analyze, and output an analysis paragraph Output containing m words through the art painting generative pre-trained model.

2. The method for analyzing artistic paintings based on a multimodal large model according to claim 1, characterized in that: Use the closed-source large language model Gemini to determine whether you know the titles and artist names of these art paintings.

3. The artistic painting analysis method based on a multimodal large model according to claim 1 is characterized in that: Use powerful closed-source large language models GPT-4 and Gemini to generate high-quality art analysis paragraphs.

4. The artistic painting analysis method based on a multimodal large model according to claim 1 is characterized in that: The multimodal large model is constructed using the open source multimodal large model ShareGPT4V-7B as a blueprint and base model.