Multimodal Large Model Evaluation Method and Device for Urban Governance Based on Multiple Layers of Text
By adopting a multi-level text evaluation method based on multi-modal large-scale model evaluation, including comprehensive evaluation of vocabulary unit similarity, low-level word frequency similarity, semantic similarity and high-level semantic similarity, the problem of inaccurate semantic differences evaluation of existing evaluation methods in urban governance scenarios is solved, and a more accurate multi-modal large-scale text output performance evaluation is achieved.
Patent Information
- Application Number
- CN202411443611.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-10-16
AI Technical Summary
Existing multimodal large model evaluation methods such as VQA are mainly evaluated through visual question and answer methods, and fail to effectively consider the semantic similarity between text outputs, resulting in inaccurate evaluation. Especially in urban governance scenarios, semantic differences in long text outputs are difficult to accurately evaluate through simple text similarity measurements.
The multi-level text-based evaluation method is adopted to divide the text into independent vocabulary through the urban management vocabulary, and the similarity of the vocabulary unit is calculated; the text is converted into a word frequency matrix based on the word bag model, and the similarity of the low-level word frequency is calculated; the semantic features are extracted using the trained language model to calculate the semantic similarity; and the high-level semantic similarity evaluation is carried out through the general large language model, and the score is finally comprehensively evaluated.
Through the multi-level evaluation method, the text output performance of multimodal large models can be more accurately measured, and the accuracy and reliability of evaluation can be improved. Especially in urban governance scenarios, semantic differences in long texts can be better captured.
Smart Images

Figure CN119399006B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data processing and smart city, and particularly relates to a method and device for evaluating a multi-modal large model for urban governance based on text multi-levels. Background Art
[0002] In recent years, large models have made breakthrough progress. A large model refers to a complex artificial neural network model with an ultra-large number of parameters, usually containing hundreds of millions to trillions of parameters. Due to its powerful learning and generalization capabilities, large models have shown extensive application potential in many fields. For example, the GPT (Generative Pre-trained Transformer, a pre-trained language model based on the Transformer architecture) series of large models developed by OpenAI can generate high-quality text content such as articles, stories, and news reports. A multi-modal large model is a further extension based on large models, which can not only process text data but also simultaneously process information in different modalities such as pictures, sounds, and videos. For example, a multi-modal large model can be used to understand the content of a picture and generate a corresponding text description accordingly. Evaluating the performance of multi-modal large models has become a subsequent task. An urban governance multi-modal large model is a multi-modal application customized for solving urban management scenarios based on a general large model as the basic foundation. An urban governance multi-modal large model can identify urban street view pictures and output a description related to urban management. For example, when a picture with illegal out-of-store business is input into the large model, the urban governance multi-modal large model outputs "The picture shows that the store merchant has expanded the business location outward, occupying public resources, which belongs to the problem of out-of-store business."
[0003] At present, the evaluation methods of existing multi-modal large models mostly adopt the evaluation method of Visual Question Answering (VQA). The VQA dataset is an open-ended question and answer dataset containing questions about pictures. These questions require an understanding of vision, language, and common sense to answer. VQA refers to generating a relevant answer given a picture and a natural language question related to that picture. The VQA dataset includes more than 200,000 pictures, with an average of 5 questions per picture, and each question has 10 ground-truth answers. However, the answer form of VQA is relatively single, and the answer types are divided into attributes, colors, yes / no, numerical values, etc., and the output text is also relatively short. For example: the question is "Which one in the picture wears glasses", and the answer is "man"; the question is "How many children are there on the bed", and the answer is "2". However, the output of large models is mostly long text. The evaluation method of VQA aims to measure the ability of the model to understand the content of the picture and give the correct answer according to the question. The VQA accuracy rate is the proportion of the answers given by the model that exactly match the manually annotated correct answers. The VQA accuracy rate is scored by comparing the matching degree between the answers predicted by the model and multiple answers provided by a group of human respondents. If the predicted answer is consistent with the content of a human answer (including consistency after simple format conversion, for example, the English word "two" is consistent with the Arabic numeral "2"), it is included in the calculation of the accuracy rate. Obviously, the VQA evaluation method does not consider the similarity between texts. Even if the text cosine similarity measure is added, it will lead to inaccurate similarity problems. On the other hand, the output of large models is mostly long text, and the similarity measure between long texts is also relatively complex. Simply using the text cosine similarity measure is difficult to evaluate the semantic similarity between two texts. Take a simple example: the text output by the large model is "I like you", while the true label text is "I don't like you". If the text similarity evaluation is used, the index is as high as 0.8. Obviously, the semantics of these two sentences are opposite. Therefore, simply using the text similarity measure is inaccurate.
[0004] In the context of smart city scenarios, the output text of the urban governance multimodal large model contains descriptions of targets and their attributes, and the text is relatively long. If only the low-level cosine similarity between the output text and the answer text is considered, and the output results of the model are not compared from a semantic perspective, it will lead to unreliable large model evaluation methods. For example, the output text of the large model is "There are unlicensed vendors in the picture. They are operating on the road, causing a crowd to gather, which may affect traffic and pose a certain safety hazard.", while the true label text is "The picture shows unlicensed vendors, but they are not selling goods, so there is no impact on the surrounding environment." If text similarity evaluation is used, the index is as high as 0.79. However, there are still significant differences between the output text of the large model and the true text. Therefore, it is inaccurate to use low-level text similarity to evaluate the performance of the multimodal large model text output in the urban governance scenario. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for evaluating an urban governance multimodal large model based on multiple levels of text, which solves the problems existing in the prior art.
[0006] The present invention is achieved through the following technical solutions:
[0007] On the one hand, the present invention provides a method for evaluating an urban governance multimodal large model based on multiple levels of text, including:
[0008] Obtain the true label text of the sample picture and the output text of the urban governance multimodal large model;
[0009] Use the urban management thesaurus to segment the true label text into first independent words, use the urban management thesaurus to segment the output text of the urban governance multimodal large model into second independent words, and obtain the lexical unit similarity score between the first independent words and the second independent words;
[0010] Create a bag-of-words model based on the urban management thesaurus, convert the true label text into a first term frequency matrix, convert the output text of the urban governance multimodal large model into a second term frequency matrix, and obtain the low-level term frequency similarity score between the first term frequency matrix and the second term frequency matrix;
[0011] Use the trained language model to extract the semantic features with context in the true label text and convert them into a first vector representation, use the trained language model to extract the semantic features with context in the output text of the urban governance multimodal large model and convert them into a second vector representation, and obtain the semantic similarity score between the first vector representation and the second vector representation;
[0012] Based on the real label text and the output text of the urban governance multimodal large model, use a general large language model with a formatted request with business guiding words and evaluation requirements to obtain a high-level semantic similarity score;
[0013] Obtain the comprehensive evaluation score of the multimodal large model according to the lexical unit similarity score, low-level word frequency similarity, semantic similarity score, and high-level semantic similarity score.
[0014] Further, obtaining the real label text and the output text of the urban governance multimodal large model of the sample picture includes:
[0015] Obtain the sample picture for the governance multimodal large model and the real label text corresponding to the sample picture;
[0016] Input the sample picture into the urban governance multimodal large model to obtain the output text of the urban governance multimodal large model corresponding to the sample picture.
[0017] Further, use the urban management thesaurus to segment the real label text into first independent words, use the urban management thesaurus to segment the output text of the urban governance multimodal large model into second independent words, and obtain the lexical unit similarity score between the first independent words and the second independent words, including:
[0018] Use urban management professional terms to construct an urban management thesaurus;
[0019] Based on the urban management thesaurus, segment the real label text into first independent words, and segment the output text of the urban governance multimodal large model into second independent words;
[0020] Use the Jaccard coefficient and / or the method of edit distance to obtain the lexical unit similarity score between the first independent words and the second independent words.
[0021] Further, create a bag-of-words model based on the urban management thesaurus, convert the real label text into a first word frequency matrix, convert the output text of the urban governance multimodal large model into a second word frequency matrix, and obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix, including:
[0022] Use the word frequency method or the term frequency-inverse document frequency method to construct a bag-of-words model based on the urban management thesaurus;
[0023] Convert the real label text into a first word frequency matrix according to the bag-of-words model;
[0024] Convert the output text of the urban governance multimodal large model into a second word frequency matrix according to the bag-of-words model;
[0025] The cosine similarity method is used to obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix.
[0026] Furthermore, a pre-trained language model is used to extract the semantic features with context in the real label text and convert them into a first vector representation, and a pre-trained language model is used to extract the semantic features with context in the output text of the urban governance multi-modal large model and convert them into a second vector representation. Obtaining the semantic similarity score between the first vector representation and the second vector representation includes:
[0027] Taking the pre-trained language model as a tokenizer and loading the pre-trained language model;
[0028] Through the pre-trained language model, the real label text and the output text of the urban governance multi-modal large model are converted into vector representations, obtaining the first vector representation corresponding to the real label text and the second vector representation corresponding to the output text of the urban governance multi-modal large model;
[0029] Using the pre-trained language model to extract the semantic features with context in the first vector representation corresponding to the real label text and the second vector representation corresponding to the output text of the urban governance multi-modal large model, obtaining the first semantic feature corresponding to the real label text and the second semantic feature corresponding to the output text of the urban governance multi-modal large model;
[0030] The cosine similarity method is used to obtain the semantic similarity score between the first vector representation and the second vector representation.
[0031] Furthermore, the pre-trained language model includes a BERT layer, a semantic feature layer, and a classification result output layer connected in sequence.
[0032] Furthermore, based on the real label text and the output text of the urban governance multi-modal large model, a general large language model with a formatted request with business guiding words and evaluation requirements is used to obtain the high-level semantic similarity score, including:
[0033] According to the formatted request with business guiding words and evaluation requirements, the real label text and the output text of the urban governance multi-modal large model are concatenated into a target question;
[0034] Based on the target question, a question request is sent to the general large language model to obtain the response text result of the general large language model;
[0035] Analyzing the response text result of the general large language model to extract the value of the high-level semantic similarity in the response text, obtaining the high-level semantic similarity score.
[0036] Further, the general large language model is a large language model LLMs with Chinese text understanding ability and text generation ability.
[0037] Further, according to the lexical unit similarity score, low-level word frequency similarity, semantic similarity score, and high-level semantic similarity score, the comprehensive evaluation score of the multi-modal large model is obtained as:
[0038]
[0039] where N represents the number of sample images, i represents the index of the sample image, α represents the influence coefficient of the lexical unit similarity, represents the lexical unit similarity score, β represents the influence coefficient of the low-level word frequency similarity, represents the low-level word frequency similarity, γ represents the influence coefficient of the semantic similarity, represents the semantic similarity score, δ represents the influence coefficient of the high-level semantic similarity, represents the high-level semantic similarity score.
[0040] On the other hand, the present invention provides a multi-modal large model evaluation device for urban governance based on text multi-levels, including: a text acquisition module, a lexical unit similarity score acquisition module, a low-level word frequency similarity score acquisition module, a semantic similarity score acquisition module, a high-level semantic similarity score acquisition module, and a comprehensive evaluation score acquisition module;
[0041] The text acquisition module is used to acquire the true label text of the sample image and the output text of the urban governance multi-modal large model;
[0042] The lexical unit similarity score acquisition module is used to extract the semantic features with context in the true label text by using a trained language model and convert them into a first vector representation, extract the semantic features with context in the output text of the urban governance multi-modal large model by using a trained language model and convert them into a second vector representation, and obtain the semantic similarity score between the first vector representation and the second vector representation;
[0043] The low-level word frequency similarity score acquisition module is used to create a bag-of-words model based on the urban management thesaurus, convert the true label text into a first word frequency matrix, convert the output text of the urban governance multi-modal large model into a second word frequency matrix, and obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix;
[0044] The semantic similarity score acquisition module is used to convert the true label text and the text output by the urban governance multi-modal large model into two vector representations respectively by using a trained language model, and extract semantic features with context from the two vector representations to obtain the first semantic feature corresponding to the true label text and the second semantic feature corresponding to the text output by the urban governance multi-modal large model, and acquire the semantic similarity score between the first semantic feature and the second semantic feature;
[0045] The high-level semantic similarity score acquisition module is used to obtain the high-level semantic similarity score by using a formatted request general large language model with business guiding words and evaluation requirements based on the true label text and the text output by the urban governance multi-modal large model;
[0046] The comprehensive evaluation score acquisition module is used to obtain the comprehensive evaluation score of the multi-modal large model according to the lexical unit similarity score, low-level word frequency similarity, semantic similarity score and high-level semantic similarity score.
[0047] An evaluation method and device for an urban governance multi-modal large model based on text multi-levels provided by the present invention uses an urban management thesaurus to segment text into independent words, calculates the lexical unit similarity score, and considers the lexical units at the fine-grained level of the text output of the urban governance large model;
[0048] Create a bag-of-words model through the urban management thesaurus, convert the text data into a word frequency matrix through the bag-of-words model, and calculate the low-level word frequency similarity score to measure the vector space distance of the text output sentences of the urban governance large model;
[0049] Calculate the semantic similarity score through a trained language model, and pay attention to the context of the text output of the urban governance large model;
[0050] Obtain the high-level semantic similarity score through a general large language model to evaluate the high-level semantic understanding ability of the urban governance large model;
[0051] Through the comprehensive evaluation score of the multi-modal large model, the lexical unit similarity, low-level word frequency similarity, semantic similarity and high-level semantic similarity of the output text are fully considered, effectively enhancing the accuracy of the evaluation of the text output of the urban governance large model. Brief Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0053] Figure 1 This is a flowchart of a method for evaluating a multimodal large model for urban governance based on multi-level text in an embodiment of the present invention;
[0054] Figure 2 This is a schematic diagram of multi-level space evaluation in an embodiment of the present invention;
[0055] Figure 3 This is a flowchart of obtaining semantic similarity scores in an embodiment of the present invention;
[0056] Figure 4 This is a flowchart of obtaining high-level semantic similarity scores in an embodiment of the present invention;
[0057] Figure 5 This is a schematic diagram of a formatted request with business guiding words and evaluation requirements in an embodiment of the present invention;
[0058] Figure 6 This is a schematic diagram of the structure of a device for evaluating a multimodal large model for urban governance based on multi-level text in an embodiment of the present invention;
[0059] Marks in the drawings and corresponding component names:
[0060] 201 - Text acquisition module, 202 - Vocabulary unit similarity score acquisition module, 203 - Low-level word frequency similarity score acquisition module, 204 - Semantic similarity score acquisition module, 205 - High-level semantic similarity score acquisition module, 206 - Comprehensive evaluation score acquisition module. Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments and drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0062] As Figure 1 shown, an embodiment of the present invention provides a method for evaluating a multimodal large model for urban governance based on multi-level text, including:
[0063] S101. Obtain the true label text of the sample image and the output text of the multimodal large model for urban governance;
[0064] S102. Use the urban management thesaurus to segment the true label text into first independent words, use the urban management thesaurus to segment the output text of the multimodal large model for urban governance into second independent words, and obtain the vocabulary unit similarity score between the first independent words and the second independent words;
[0065] S103. Create a bag-of-words model based on the urban management thesaurus, convert the real label text into a first term frequency matrix, convert the text output by the urban governance multi-modal large model into a second term frequency matrix, and obtain the low-level term frequency similarity score between the first term frequency matrix and the second term frequency matrix;
[0066] S104. Use the trained language model to extract the semantic features with context in the real label text and convert them into a first vector representation, use the trained language model to extract the semantic features with context in the text output by the urban governance multi-modal large model and convert them into a second vector representation, and obtain the semantic similarity score between the first vector representation and the second vector representation;
[0067] S105. Based on the real label text and the text output by the urban governance multi-modal large model, use the general large language model with business guiding words and evaluation requirements to obtain the high-level semantic similarity score;
[0068] S106. Obtain the comprehensive evaluation score of the multi-modal large model according to the lexical unit similarity score, low-level term frequency similarity, semantic similarity score, and high-level semantic similarity score.
[0069] As Figure 2 shown, for further illustration of Figure 1 The first-level original text is the text content described in S101, the second-level lexical structure is the independent lexical content of S102, the third-level term frequency matrix is the term frequency matrix content of S103, the fourth-level semantic vector is the vector representation content of S104, and the fifth-level high-level semantics is the high-level semantic content of S105.
[0070] In a possible implementation, obtaining the real label text and the text output by the urban governance multi-modal large model of the sample picture includes: obtaining the sample picture for the multi-modal large model of governance and the real label text corresponding to the sample picture; inputting the sample picture into the urban governance multi-modal large model to obtain the text output by the urban governance multi-modal large model corresponding to the sample picture.
[0071] In a possible implementation, the urban governance multi-modal large model is a multi-modal large model obtained by fine-tuning and training for a large number of urban governance scenario picture data;
[0072] The multi-modal large model receives visual pictures and question text inputs and generates an answer text output result corresponding to the picture.
[0073] The urban governance multi-modal large model adopted in this embodiment is based on the general large language model LLMs, and the model is obtained through parameter setting, training, and processing.
[0074] In a possible implementation, the real label text is segmented into first independent words using an urban management thesaurus, the text output by the urban governance multimodal large model is segmented into second independent words using the urban management thesaurus, and the lexical unit similarity score between the first independent words and the second independent words is obtained, including:
[0075] Construct an urban management thesaurus using urban management professional terms;
[0076] Based on the urban management thesaurus, segment the real label text into first independent words and segment the text output by the urban governance multimodal large model into second independent words;
[0077] In a possible implementation, use the Jaccard coefficient or the edit distance method to obtain the lexical unit similarity score between the first independent words and the second independent words.
[0078] Construct an urban management thesaurus using urban management professional terms, such as: "bicycle lying down", "items for drying", "ground-standing signs set in the square", "slogan-like promotional materials", "mobile vendors in the shed", "operating on the road", "bagged garbage", "garbage overflowing", "scattered garbage", "cracking and damage", "piled-up materials", "scribbling and painting", "banner-like promotional materials", "vendor's tricycle", "vendor's four-wheel truck", "occupying the road for business", "large area of water accumulation", "sporadic garbage", "abandoned furniture and equipment", "illegal street slope", "outdoor advertising", "air model or arch", "junction box door not closed", "junction box bottom cover missing", "facility box door not closed", "anti-collision barrel lying down", "pavement pile lying down", "guardrail falling off", "occupying the road for decoration", "medical waste", "barrel damaged", "street tree missing", "fire hydrant abnormal", "pole bottom cover missing", "kitchen waste", "tree protection pool", "isolation body lying down", "pavement damaged", "curbstone", "construction site enclosure", "street seat", "spilled mud", "construction waste materials", "traffic sign", "waste collection", "poultry and livestock", "green space guardrail", "manhole cover missing", "grate missing".
[0079] Input a picture of an urban street scene into the urban governance multimodal large model in this implementation, and the text output by the multimodal large model is: This photo shows a vendor's four-wheel truck parked by the roadside, with some items on the truck, and several people standing beside it seemingly talking.
[0080] The real label text is: The street scene at an intersection is shown in the picture. Mobile vendors in the shed are parked by the roadside, and the cart is full of fruits. There are several passers-by buying fruits.
[0081] For the input urban street scene picture, in the urban street scene, the first independent term extracted is mobile vendors in a shed, and the second independent term extracted is the vendor's four-wheel truck. According to the first and second independent terms extracted, the lexical unit similarity score is calculated using the Jaccard coefficient. It is 0.43.
[0082] The urban management bag-of-words model only focuses on the frequency of term occurrences and ignores the order and grammatical structure of words in the text. Therefore, the embodiments of the present invention further process it.
[0083] In a possible implementation, a bag-of-words model is created based on the urban management thesaurus, the real label text is converted into a first term frequency matrix, the text output by the urban governance multi-modal large model is converted into a second term frequency matrix, and the low-level term frequency similarity score between the first term frequency matrix and the second term frequency matrix is obtained, including:
[0084] Construct a bag-of-words model based on the urban management thesaurus using the term frequency (TF) method or the term frequency–inverse document frequency (TF-IDF) method;
[0085] Convert the real label text into a first term frequency matrix according to the bag-of-words model;
[0086] Convert the text output by the urban governance multi-modal large model into a second term frequency matrix according to the bag-of-words model;
[0087] Use the cosine similarity method to obtain the low-level term frequency similarity score between the first term frequency matrix and the second term frequency matrix.
[0088] If the term frequency matrices of the real label text and the text output by the urban governance multi-modal large model are obtained separately, then the first term frequency matrix and the second term frequency matrix obtained are both vectors. It is worth noting that the term frequency matrices of the real label text and the text output by the urban governance multi-modal large model can also be obtained together. For example: the label truth value is "The urban management problem shown in the picture is that there are fruit and vegetable vendor tricycles engaged in mobile (non-motor vehicle) business activities in public places", and the text predicted and output by the large model is "The urban management problem that exists is fruit and vegetable vendor tricycles engaged in mobile business activities". They are respectively converted into term frequency matrices X using the TF-IDF method. The first row vector of the matrix corresponds to the label truth value, and the second row vector corresponds to the text output by the large model. Then, the similarity value is calculated using the cosine similarity method.
[0089] For the above input urban street scene pictures, in the urban street scene, the TF-IDF method is used to obtain the first term frequency matrix, and the TF-IDF method is used to obtain the second term frequency matrix, and the low-level term frequency similarity is calculated is 0.51.
[0090] In a possible implementation manner, a pre-trained language model is used to extract the semantic features with context in the real label text and convert them into a first vector representation, and a pre-trained language model is used to extract the semantic features with context in the output text of the urban governance multimodal large model and convert them into a second vector representation, and the semantic similarity score between the first vector representation and the second vector representation is obtained, including: using the pre-trained language model as a tokenizer and loading the pre-trained language model; converting the real label text and the output text of the urban governance multimodal large model into vector representations through the language model tokenizer to obtain the first vector representation corresponding to the real label text and the second vector representation corresponding to the output text of the urban governance multimodal large model; using the pre-trained language model to extract the semantic features with context in the first vector representation corresponding to the real label text and the second vector representation corresponding to the output text of the urban governance multimodal large model to obtain the first semantic feature corresponding to the real label text and the second semantic feature corresponding to the output text of the urban governance multimodal large model; using the cosine similarity method to obtain the semantic similarity score between the first vector representation and the second vector representation.
[0091] The specific calculation process of the above semantic similarity score includes:
[0092] Using the pre-trained language model to extract the semantic features with context in the real label text and convert them into a first vector representation, and then calculating the semantic similarity score;
[0093] Among them, the pre-trained language model adds a semantic feature layer and a classification result output layer based on the pre-trained language model. The pre-trained language model is a natural language processing model based on self-attention to extract the long-distance dependence relationship of the text, and the classification label of the training corpus data is the content of the urban management thesaurus.
[0094] The pre-trained language model uses BERT (Bidirectional Encoder Representations from Transformers). The BERT model is built based on the Encoder part of Transformer and adopts a bidirectional training method, which can consider the context information before and after the vocabulary at the same time, enabling the model to obtain a richer understanding of the context. Append a semantic feature layer (with a feature length of 512) to the last output layer of BERT, and then append a classification result output layer (the output result is classified into 50 categories, which is the content of the urban management thesaurus). The training corpus data uses the text data in the training of the urban governance large model, and the classification labels use the content of the urban management thesaurus, including the following steps:
[0095] Use the pre-trained language model as a tokenizer and load the pre-trained language model;
[0096] The language model tokenizer encodes the text to be tested and converts the text into a vector representation;
[0097] Use the pre-trained language model to extract the output of the penultimate layer of the model from the vector representation as the semantic feature with context;
[0098] Calculate the cosine similarity score for the semantic feature.
[0099] In a possible implementation, the pre-trained language model includes a BERT layer, a semantic feature layer, and a classification result output layer connected in sequence.
[0100] As Figure 3 shown, where the upper part of the dotted line is the training mode and the lower part of the dotted line is the inference mode. The pre-trained language model uses BERT (Bidirectional Encoder Representations from Transformers). The BERT model is built based on the Encoder part of Transformer and adopts a bidirectional training method, which can consider the context information before and after the vocabulary at the same time, enabling the model to obtain a richer understanding of the context. Append a semantic feature layer (with a feature length of 512) to the last output layer of BERT, and then append a classification result output layer (the output result is classified into 50 categories, which is the content of the urban management thesaurus). The training corpus data uses the text data in the training of the urban governance large model, and the classification labels use the content of the urban management thesaurus.
[0101] Based on the above input urban street scene pictures, in this urban street scene, input the scene description into the pre-trained language model, extract the semantic features with context respectively, and calculate the semantic similarity score Sim isema It is 0.78.
[0102] In a possible implementation, as Figure 4 shown, based on the real label text and the output text of the urban governance multi-modal large model, a general large language model with a formatted request with business guiding words and evaluation requirements is used to obtain a high-level semantic similarity score, including:
[0103] According to the formatted request with business guiding words and evaluation requirements, the real label text and the output text of the urban governance multi-modal large model are concatenated into a target question;
[0104] Based on the target question, a question request is sent to the general large language model to obtain the response text result of the general large language model;
[0105] Analyze the response text result of the general large language model to extract the value of the high-level semantic similarity in the response text, and obtain the high-level semantic similarity score.
[0106] In this implementation, the general large language model LLMs refers to a general large language model based on the generative pre-trained Transformer, which has the ability to understand Chinese text and generate text. The general large language model LLMs can use GPT-4 of OpenAI, or LLaMA3 of Meta AI, or ChatGLM-6B of Tsinghua University, or Qwen of Alibaba Qianwen.
[0107] The range of the high-level semantic similarity score is 0 - 1.0, including the boundary values. 0 means the semantics are completely opposite, and 1.0 means the semantics are completely the same.
[0108] As Figure 5 shown, the formatted request with business guiding words and evaluation requirements is as follows: Sentence 1: A. Sentence 2: B. Please evaluate the semantic similarity of the above two sentences from the perspective of urban management business and give a score between 0 and 1.0, where 0 means completely different and 1 means completely the same. Among them, A and B represent the label true value text and the large model prediction output text respectively.
[0109] In some possible implementation manners, an urban street scene picture (this picture is the same urban street scene picture as above, and this sentence should be repeated four times) is input into the urban governance multi-modal large model in this implementation. The text output of the multi-modal large model is: There is a pickup truck parked by the roadside on the urban street scene picture, and there are some items on the truck, and several people are standing beside and seem to be talking. The real label text is: The picture shows a street scene at an intersection, a silver pickup truck is parked by the roadside, and the truck bed is full of fruits, and several passers-by are buying fruits.
[0110] Process the pictures of urban street scenes to obtain the following formatted requests with business guiding words and evaluation requirements:
[0111] Sentence 1: This photo shows a pickup truck parked by the roadside, carrying some items, and several people standing beside it seem to be talking.
[0112] Sentence 2: The picture shows a street scene at an intersection. A silver pickup truck is parked by the roadside, and its cargo bed is full of fruits. Several passers-by are buying fruits.
[0113] Please evaluate the semantic similarity of the above two sentences and give a score between 0 and 1.0, where 0 means completely different and 1 means completely the same.
[0114] Based on the input urban street scene pictures above, in this urban street scene, according to the formatted requests with business guiding words and evaluation requirements, splice the real label text and the output text of the urban governance multimodal large model into a target question. Based on the target question, send a question request to the general large language model to obtain the response text result of the general large language model, parse the response text result of the general large language model to extract the numerical value of the high-level semantic similarity in the response text, and obtain the high-level semantic similarity score It is 0.8.
[0115] In a possible implementation manner, according to the lexical unit similarity score, low-level word frequency similarity, semantic similarity score, and high-level semantic similarity score, obtain the comprehensive evaluation score of the multimodal large model as:
[0116]
[0117] Among them, ACC represents the comprehensive evaluation score, N represents the number of sample pictures, i represents the index of the sample pictures, α represents the influence coefficient of the lexical unit similarity represents the lexical unit similarity score, β represents the influence coefficient of the low-level word frequency similarity represents the low-level word frequency similarity, γ represents the influence coefficient of the semantic similarity represents the semantic similarity score, δ represents the influence coefficient of the high-level semantic similarity represents the high-level semantic similarity score.
[0118] In some possible embodiments, a specific numerical example is given. Among them, the weight of high-level semantics is the largest, and the influence coefficient δ of high-level semantic similarity can be set to 0.5. The weight of semantic similarity is the second, and the influence coefficient γ of semantic similarity can be set to 0.3. The weight of low-level word frequency similarity is the third, and the influence coefficient β of low-level word frequency similarity can be set to 0.15. The weight of lexical unit similarity is the smallest, and the influence coefficient α of lexical unit similarity can be set to 0.05. Based on this, the lexical unit similarity score Sim word is 0.43, and the low-level word frequency similarity score Sim freq is 0.51, the semantic similarity score Sim sema is 0.78, and the high-level semantic similarity score Sim llms is 0.8.
[0119] Based on the above specific values, the calculation result of the comprehensive evaluation score formula is 0.732. If the low-level word frequency similarity is used alone as the measurement index, the similarity is only 0.51. The higher the comprehensive score, the higher the accuracy of the text output of the multi-modal large model. Obviously, the calculation result of 0.732 in the embodiments of the present invention can better reflect the accuracy of the text output of the multi-modal large model.
[0120] As Figure 6 shown, the embodiments of the present invention provide a multi-modal large model evaluation device for urban governance based on text multi-levels, including: a text acquisition module 201, a lexical unit similarity score acquisition module 202, a low-level word frequency similarity score acquisition module 203, a semantic similarity score acquisition module 204, a high-level semantic similarity score acquisition module 205, and a comprehensive evaluation score acquisition module 206;
[0121] The text acquisition module 201 is used to acquire the real label text of the sample picture and the text output by the urban governance multi-modal large model;
[0122] The lexical unit similarity score acquisition module 202 is used to segment the real label text into first independent words by using the urban management thesaurus, segment the text output by the urban governance multi-modal large model into second independent words by using the urban management thesaurus, and obtain the lexical unit similarity score between the first independent words and the second independent words;
[0123] The low-level word frequency similarity score acquisition module 203 is used to create a bag-of-words model based on the urban management thesaurus, convert the real label text into a first word frequency matrix, convert the text output by the urban governance multi-modal large model into a second word frequency matrix, and obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix;
[0124] The semantic similarity score acquisition module 204 is configured to use a trained language model to convert the true label text and the text output by the urban governance multimodal large model into two vector representations respectively, extract semantic features with context from the two vector representations, obtain the first semantic feature corresponding to the true label text and the second semantic feature corresponding to the text output by the urban governance multimodal large model, and acquire the semantic similarity score between the first semantic feature and the second semantic feature;
[0125] The high-level semantic similarity score acquisition module 205 is configured to use a general large language model with a formatted request including business guiding words and evaluation requirements, based on the true label text and the text output by the urban governance multimodal large model, to obtain the high-level semantic similarity score;
[0126] The comprehensive evaluation score acquisition module 206 is configured to obtain the comprehensive evaluation score of the multimodal large model according to the lexical unit similarity score, the low-level word frequency similarity, the semantic similarity score, and the high-level semantic similarity score.
[0127] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0128] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0129] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the function specified in the flowchart Figure 1 a flowchart or multiple flowcharts and / or block Figure 1 a block or multiple blocks.
[0130] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in the flowchart Figure 1 a flowchart or multiple flowcharts and / or block Figure 1 a block or multiple blocks.
[0131] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0132] It is obvious that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-modal large model evaluation method for urban governance based on multi-level text, characterized in that: include: Obtain the real label text of the sample image and the output text of the urban governance multimodal large model; The urban management vocabulary is used to segment the real label text into first independent words, the urban management vocabulary is used to segment the output text of the urban governance multimodal large model into second independent words, and the vocabulary unit similarity score between the first independent words and the second independent words is obtained; Create a bag-of-words model based on the urban management vocabulary, convert the real label text into the first word frequency matrix, convert the output text of the urban governance multimodal large model into the second word frequency matrix, and obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix; The trained language model is used to extract semantic features with context in the real label text and convert them into a first vector representation. The trained language model is used to extract semantic features with context in the output text of the urban governance multimodal large model and convert them into a second vector representation, and the semantic similarity score between the first vector representation and the second vector representation is obtained. Based on the real label text and the output text of the urban governance multimodal large model, a formatted request general large language model with business guide words and evaluation requirements is used to obtain a high-level semantic similarity score; The comprehensive evaluation score of the multimodal large model is obtained based on the vocabulary unit similarity score, low-level word frequency similarity, semantic similarity score, and high-level semantic similarity score.
2. According to the text-based multi-level urban governance multimodal large model evaluation method of claim 1, it is characterized by: The method of obtaining the real label text of the sample image and the output text of the urban governance multimodal large model includes: Obtain sample images for managing multimodal large models and the real label text corresponding to the sample images; The sample image is input into the multimodal large model of urban governance to obtain the output text of the multimodal large model of urban governance corresponding to the sample image.
3. The method for evaluating a multimodal large model of urban governance based on multi-level text according to claim 1 is characterized in that: The urban management vocabulary is used to segment the real label text into the first independent vocabulary, and the urban management vocabulary is used to segment the output text of the urban governance multimodal large model into the second independent vocabulary, and the vocabulary unit similarity score between the first independent vocabulary and the second independent vocabulary is obtained, including: Use professional urban management terminology to build an urban management vocabulary; Based on the urban management vocabulary, the real label text is segmented into a first independent vocabulary, and the output text of the urban governance multimodal large model is segmented into a second independent vocabulary; The Jaccard coefficient and / or edit distance method is used to obtain a vocabulary unit similarity score between the first independent vocabulary and the second independent vocabulary.
4. According to claim 1, the method for evaluating the multimodal large model of urban governance based on multi-level text is characterized in that: A bag-of-words model is created based on the urban management vocabulary, the real label text is converted into the first word frequency matrix, the output text of the urban governance multimodal large model is converted into the second word frequency matrix, and the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix is obtained, including: Use the word frequency method or the word frequency-inverse document frequency method to build a bag-of-words model based on the urban management vocabulary; Converting the real label text into a first word frequency matrix according to the bag-of-words model; Converting the output text of the urban governance multimodal large model into a second word frequency matrix according to the bag-of-words model; The cosine similarity method is used to obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix.
5. According to the text-based multi-level urban governance multimodal large model evaluation method of claim 1, it is characterized by: The trained language model is used to extract the semantic features with context in the real label text and convert it into a first vector representation. The trained language model is used to extract the semantic features with context in the output text of the urban governance multimodal large model and convert it into a second vector representation. The semantic similarity score between the first vector representation and the second vector representation is obtained, including: Use the trained language model as a word segmenter and load the trained language model; The real label text and the output text of the urban governance multimodal large model are converted into vector representations through the trained language model to obtain a first vector representation corresponding to the real label text and a second vector representation corresponding to the output text of the urban governance multimodal large model; The cosine similarity method is used to obtain the semantic similarity score between the first vector representation and the second vector representation.
6. The method for evaluating a multimodal large model of urban governance based on multi-level text according to claim 5 is characterized in that: The trained language model includes a BERT layer, a semantic feature layer, and a classification result output layer connected in sequence.
7. The method for evaluating a multimodal large model of urban governance based on multi-level text according to claim 1 is characterized in that: Based on the real label text and the output text of the urban governance multimodal large model, a formatted request general large language model with business guide words and evaluation requirements is used to obtain high-level semantic similarity scores, including: According to the formatted request with business guide words and evaluation requirements, the real label text and the output text of the urban governance multimodal large model are spliced into the target question; Based on the target question, a question request is sent to the general large language model to obtain a response text result of the general large language model; The response text result of the general large language model is parsed to extract the numerical value of the high-level semantic similarity in the response text to obtain the high-level semantic similarity score.
8. The method for evaluating a multimodal large model of urban governance based on multi-level text according to claim 7 is characterized in that: The universal large language model is a large language model LLMs with Chinese character understanding and text generation capabilities.
9. The method for evaluating a multimodal large model of urban governance based on multi-level text according to claim 1 is characterized in that: According to the vocabulary unit similarity score, low-level word frequency similarity, semantic similarity score and high-level semantic similarity score, the comprehensive evaluation score of the multimodal large model is obtained as follows: Among them, ACC represents the comprehensive evaluation score, N represents the number of sample images, i represents the index of the sample image, and α represents the influence coefficient of the vocabulary unit similarity. represents the similarity score of the vocabulary unit, β represents the influence coefficient of the low-level word frequency similarity, represents the low-level word frequency similarity, γ represents the semantic similarity influence coefficient, represents the semantic similarity score, δ represents the high-level semantic similarity influence coefficient, Represents the high-level semantic similarity score.
10. A multi-modal large model evaluation device for urban governance based on multi-level text, characterized in that: It includes: text acquisition module, vocabulary unit similarity score acquisition module, low-level word frequency similarity score acquisition module, semantic similarity score acquisition module, high-level semantic similarity score acquisition module and comprehensive evaluation score acquisition module; The text acquisition module is used to obtain the real label text of the sample image and the output text of the urban governance multimodal large model; The vocabulary unit similarity score acquisition module is used to use the urban management vocabulary to segment the real label text into first independent vocabulary, use the urban management vocabulary to segment the urban governance multimodal large model output text into second independent vocabulary, and obtain the vocabulary unit similarity score between the first independent vocabulary and the second independent vocabulary; The low-level word frequency similarity score acquisition module is used to create a bag-of-words model based on the urban management vocabulary, convert the real label text into a first word frequency matrix, convert the output text of the urban governance multimodal large model into a second word frequency matrix, and obtain the low-level word frequency similarity score between the first word frequency matrix and the second word frequency matrix; The semantic similarity score acquisition module is used to use the trained language model to extract semantic features with context in the real label text and convert them into a first vector representation, use the trained language model to extract semantic features with context in the output text of the urban governance multimodal large model and convert them into a second vector representation, and obtain the semantic similarity score between the first vector representation and the second vector representation; The high-level semantic similarity score acquisition module is used to obtain a high-level semantic similarity score based on the real label text and the output text of the urban governance multimodal large model, using a formatted request general large language model with business guide words and evaluation requirements; The comprehensive evaluation score acquisition module is used to obtain the comprehensive evaluation score of the multimodal large model based on the vocabulary unit similarity score, the low-level word frequency similarity, the semantic similarity score and the high-level semantic similarity score.
Citation Information
Patent Citations
Medium-law inter-translation quality evaluation method based on BERT
CN117034961A
Global visual guidance image description generation method based on cross-modal large model
CN118378623A
Cited By
Multi-modal large model-oriented city spatio-temporal data embedding method
CN120179883A
A method for embedding urban spatiotemporal data in multimodal large models
CN120179883B