Method and apparatus for evaluating a multi-modal large model for urban governance based on text tagging

By extracting and structured output text information from the multimodal model of urban governance and calculating the accuracy index of a single picture, the problem of insufficient reliability of the existing evaluation methods is solved and the accuracy and reliability of the evaluation is improved.

CN119322933BActive Publication Date: 2025-05-30CHENGDU ZHIHUI HENENG CITY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411372640.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-05-30
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

The existing multimodal large-modal model evaluation method of urban governance mainly relies on text similarity, resulting in the inability to evaluate the results in urban governance scenarios.

Method used

By extracting the object existence, classification labels and target position information in the output text of the multimodal model of urban governance, a structured dictionary list is constructed, and the accuracy indicators of a single picture are calculated based on these data, and the comprehensive evaluation indicators are finally obtained.

Benefits of technology

The accuracy of the multimodal large model evaluation of urban governance has been improved, and the object judgment ability, label classification ability and target positioning ability of the model are fully considered, and the unreliability of the pure text similarity evaluation indicators is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119322933B_ABST
    Figure CN119322933B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for evaluating a multi-modal large model for urban governance based on text tagging, which relates to the technical fields of data processing and smart cities. Sample data is obtained to get sample pictures and corresponding text tags. The sample pictures are input into the model, and the output text corresponding to the model is obtained. The object existence, classification tags, and target location information in the output text are extracted to obtain a list of structured dictionaries. According to the data in the list of structured dictionaries and the text tags corresponding to the sample pictures, the accuracy index of the model for a single sample picture is obtained. According to the accuracy indexes of multiple single sample pictures, the comprehensive evaluation index of the model on the sample data is obtained. It fully considers the object judgment ability, label classification ability, and target positioning ability of the multi-modal large model for urban governance, overcomes the problem of unreliable pure text similarity evaluation indexes, and improves the accuracy of evaluating the multi-modal large model for urban governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and smart city, and particularly relates to a method and device for evaluating a multi-modal large model for urban governance based on text tagging. Background Art

[0002] In recent years, breakthroughs have been made in large models. A large model refers to a complex artificial neural network model with an ultra-large number of parameters, usually containing hundreds of millions to trillions of parameters. Due to its powerful learning and generalization capabilities, large models have shown broad application potential in many fields. For example, OpenAI's GPT series of large models can generate high-quality text content such as articles, stories, and news reports. Based on the application potential of large models, multi-modal large models have been further developed. Multi-modal large models can not only process text data but also simultaneously process information in different modalities such as pictures, sounds, and videos. For example, a multi-modal large model can be used to understand the content of a picture and generate a corresponding text description based on it. Based on the ability of multi-modal large models to process complex pictures, urban governance multi-modal large models have been developed, relying on general large models as the basic base and being multi-modal applications specifically customized for solving urban management scenarios. For example, an urban governance multi-modal large model can identify urban street view pictures and output a description related to urban management.

[0003] Currently, existing evaluation methods for urban governance multi-modal large models mostly calculate the difference between the predicted output results of urban governance multi-modal large models and the real descriptions using methods related to text similarity. For example, the predicted output result of an urban governance multi-modal large model is that there are unlicensed vendors selling in the picture [[188,411,488,871]], while the real description of the picture is that there are no illegal unlicensed vendors in the picture [[190,425,480,866]]. If the calculation method of pure text similarity is used, the accuracy rate of the predicted result and the real result is 0.75. From a semantic perspective, the urban governance multi-modal large model outputs that there are problems in the picture while the real description shows that there are no problems in the picture. Obviously, the difference between the predicted result and the real result is huge. Therefore, using pure text similarity to evaluate urban governance multi-modal large models is unreliable in urban governance scenarios. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and device for evaluating a multi-modal large model for urban governance based on text tagging, which solves the problems existing in the prior art.

[0005] The present invention is achieved through the following technical solutions:

[0006] On the one hand, the present invention provides a method for evaluating a multi-modal large model for urban governance based on text tagging, including:

[0007] Obtain sample data for evaluating the urban governance multi-modal large model; wherein, the sample data includes sample pictures and text labels corresponding to the sample pictures;

[0008] Input the sample pictures in the sample data into the urban governance multi-modal large model to obtain the output text corresponding to the urban governance multi-modal large model;

[0009] Extract the object existence, classification label, and target location information from the output text, and obtain a list of structured dictionaries based on the object existence, classification label, and target location information;

[0010] Obtain the accuracy index of the urban governance multi-modal large model for a single sample picture according to the data in the list of structured dictionaries and the text labels corresponding to the sample pictures;

[0011] Obtain the comprehensive evaluation index for evaluating the urban governance multi-modal large model on the sample data according to the accuracy indexes of multiple single sample pictures.

[0012] In the embodiment of the present invention, the urban governance multi-modal large model is a multi-modal large model obtained through fine-tuning training for a large number of urban governance scenario picture data;

[0013] The multi-modal large model receives visual pictures and question text inputs and generates the text corresponding to the pictures.

[0014] In the embodiment of the present invention, extracting the object existence, classification label, and target location information from the output text includes:

[0015] For the description corresponding to any target in the output text, use the keyword recognition method to recognize the object existence and classification label corresponding to the target;

[0016] For the description corresponding to any target in the output text, use the identifier recognition method to recognize the target location information corresponding to the target;

[0017] Traverse each target in the output text to obtain the object existence, classification label, and target location information corresponding to each target.

[0018] In the embodiment of the present invention, the list of structured dictionaries is:

[0019] Construct a list, each element in the list corresponds to a picture, and the elements of the list are structured dictionaries;

[0020] Among them, the list of structured dictionaries stores the picture name, picture size attribute, existence keyword, label keyword, and target box.

[0021] In an embodiment of the present invention, according to the data in the structured dictionary list and the text labels corresponding to the sample pictures, the accuracy index of the urban governance multi-modal large model for a single sample picture is:

[0022]

[0023] Among them, IOU i represents the intersection over union of the bounding box of the i-th target in a single sample picture, represents the true classification label of the i-th target in a single sample picture, represents the predicted classification label of the i-th target in the text label corresponding to the sample picture, represents the predicted classification label and the true classification label the text cosine similarity between them, represents the quantization value of whether the target exists in the true result, represents the quantization value of whether the target exists in the predicted result corresponding to the text label, and N represents the total number of targets in a single sample picture.

[0024] In an embodiment of the present invention, the comprehensive evaluation index for evaluating the urban governance multi-modal large model on the sample data according to the accuracy indexes of multiple single sample pictures is:

[0025]

[0026] Among them, M represents the total number of sample pictures in the sample data, λ represents the influence coefficient of the single picture accuracy, represents the text label corresponding to the j-th sample picture in the sample data, the predicted result text description of the j-th sample picture corresponding to the urban governance multi-modal large model in the sample data, represents the true result text description and the predicted result text description the text cosine similarity between them, η represents the influence coefficient of the text similarity, is the accuracy of a single sample picture.

[0027] On the other hand, the present invention provides an evaluation device for an urban governance multi-modal large model based on text labeling, including: a data acquisition module, an output acquisition module, a structured dictionary acquisition module, a first index acquisition module, a second index acquisition module, and an evaluation result acquisition module;

[0028] The data acquisition module is used to acquire sample data for evaluating the urban governance multi-modal large model; among them, the sample data includes sample pictures and text labels corresponding to the sample pictures;

[0029] The output acquisition module is used to input the sample pictures in the sample data into the urban governance multi-modal large model to obtain the output text corresponding to the urban governance multi-modal large model;

[0030] The structured dictionary acquisition module is used to construct a list, where each element in the list corresponds to a picture, and the elements of the list are structured dictionaries; the structured dictionary list stores picture names, picture size attributes, existence keywords, label keywords, and target boxes;

[0031] The first index acquisition module is used to obtain the accuracy index of the urban governance multi-modal large model for a single sample picture according to the data in the structured dictionary list and the text label corresponding to the sample picture;

[0032] The second index acquisition module is used to obtain a comprehensive evaluation index for evaluating the urban governance multi-modal large model on the sample data according to the accuracy index of the single sample picture.

[0033] In an embodiment of the present invention, the structured dictionary acquisition module includes a data extraction unit and a dictionary acquisition unit;

[0034] The data extraction unit is used to, for the description corresponding to any target in the output text, identify the object existence and classification label corresponding to the target by using the keyword recognition method; for the description corresponding to any target in the output text, identify the target position information corresponding to the target by using the identifier recognition method; traverse each target in the output text to obtain the object existence, classification label, and target position information corresponding to each target;

[0035] The dictionary acquisition unit is used to construct a list, where each list element corresponds to a picture, and the list element stores the object existence, classification label, and target position information to obtain the structured dictionary stored in the list.

[0036] In an embodiment of the present invention, the first index acquisition module includes: a prediction data acquisition unit, a real data acquisition unit, and an accuracy index acquisition unit;

[0037] The prediction data acquisition unit is used to determine whether there are quantization values for the predicted prediction box, predicted classification label, and prediction result in a single sample picture according to the data in the structured dictionary to obtain the prediction data corresponding to the single sample picture;

[0038] The real data acquisition unit is used to determine whether there are quantization values for the real prediction box, real classification label, and real result in the text label according to the text label corresponding to the sample picture to obtain the real data corresponding to the single sample picture;

[0039] The accuracy metric acquisition unit is used to obtain the accuracy metric of a single sample image according to the prediction data corresponding to the single sample image and the ground truth data corresponding to the single sample image, as follows:

[0040]

[0041] where IOU i represents the intersection over union of the bounding box of the i-th object in the single sample image, represents the ground truth classification label of the i-th object in the single sample image, represents the predicted classification label of the i-th object in the text label corresponding to the sample image, represents the predicted classification label and the ground truth classification label the text cosine similarity between them, represents the quantization value of whether the object exists in the ground truth result, represents the quantization value of whether the object exists in the predicted result corresponding to the text label, and N represents the total number of objects in the single sample image.

[0042] In an embodiment of the present invention, the second metric acquisition module includes a text similarity prediction unit and a comprehensive evaluation metric acquisition unit;

[0043] The text similarity prediction unit is used to determine the text cosine similarity between the ground truth result text description and the predicted result text description according to the sample image in the sample data and the text label corresponding to the sample image;

[0044] The comprehensive evaluation metric acquisition unit is used to determine the comprehensive evaluation metric of the urban governance multi-modal large model according to the text cosine similarities corresponding to multiple single sample images and the accuracy metric, as follows:

[0045]

[0046] where M represents the total number of sample images in the sample data, λ represents the influence coefficient of the single accuracy, represents the text label corresponding to the j-th sample image in the sample data, the predicted result text description of the j-th sample image in the sample data corresponding to the urban governance multi-modal large model, represents the ground truth result text description and the predicted result text description the text cosine similarity between them, η represents the influence coefficient of the text similarity, is the accuracy of a single sample image.

[0047] A method and device for evaluating a multi-modal large model for urban governance based on text tagging provided by the present invention extract object existence, classification labels, and target location information from the text output by the multi-modal large model for urban governance; store the extraction results in a structured dictionary list, and obtain the accuracy index of a single picture according to the data recorded in the structured dictionary list; then consider the pure text similarity and calculate the comprehensive evaluation index, which fully considers the object judgment ability, label classification ability, and target positioning ability of the multi-modal large model for urban governance, overcomes the problem of unreliable pure text similarity evaluation index, and greatly improves the accuracy of evaluating the multi-modal large model for urban governance. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. In the drawings:

[0049] Figure 1 is a flowchart of a method for evaluating a multi-modal large model for urban governance based on text tagging provided by an embodiment of the present invention;

[0050] Figure 2 is a schematic structural diagram of a device for evaluating a multi-modal large model for urban governance based on text tagging provided by an embodiment of the present invention;

[0051] Figure 3 is a schematic diagram of the structured dictionary list provided by an embodiment of the present invention;

[0052] Figure 4 is a schematic diagram for calculating the accuracy index of a single picture provided by an embodiment of the present invention.

[0053] Marks in the drawings and corresponding component names:

[0054] Among them, 201 - data acquisition module, 202 - output acquisition module, 203 - structured dictionary acquisition module, 204 - first index acquisition module, 205 - second index acquisition module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To make the purpose, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not used to limit the present invention.

[0056] As Figure 1As shown in the figure, an embodiment of the present invention provides a method for evaluating a multi-modal large model for urban governance based on text tagging, including:

[0057] S101. Obtain sample data for evaluating the multi-modal large model for urban governance; wherein, the sample data includes sample pictures and text tags corresponding to the sample pictures;

[0058] S102. Input the sample pictures in the sample data into the multi-modal large model for urban governance to obtain the output text corresponding to the multi-modal large model for urban governance;

[0059] S103. Extract the object existence, classification labels, and target location information from the output text, and obtain a structured dictionary list according to the object existence, classification labels, and target location information;

[0060] S104. According to the data in the structured dictionary list and the text tags corresponding to the sample pictures, that is, the ground truth label information, obtain the accuracy index of the multi-modal large model for urban governance in a single sample picture;

[0061] S105. According to the accuracy indexes of multiple single sample pictures, obtain the comprehensive evaluation index for evaluating the multi-modal large model for urban governance on the sample data.

[0062] In the embodiment of the present invention, the multiple single sample pictures in S105 include all the pictures of the sample.

[0063] By using the structured dictionary list, accuracy index, and comprehensive evaluation index as the evaluation results together, it enables the staff to verify, check, and trace the evaluation results.

[0064] The method for evaluating a multi-modal large model for urban governance based on text tagging provided in this embodiment extracts the object existence, classification labels, and target location information from the output text of the multi-modal large model for urban governance; stores the extraction results in the structured dictionary list, and obtains the accuracy index of a single picture according to the data recorded in the structured dictionary list; then considers the pure text similarity and calculates the comprehensive evaluation index, fully considering the object judgment ability, label classification ability, and target positioning ability of the multi-modal large model for urban governance, overcomes the problem of unreliable pure text similarity evaluation indexes, and greatly improves the accuracy of evaluating the multi-modal large model for urban governance.

[0065] In the embodiment of the present invention, according to the structured dictionary list, accuracy index, and comprehensive evaluation index, in order to avoid data isolation, the data in the structured dictionary list is used to obtain the indexes.

[0066] In an embodiment of the present invention, the urban governance multi-modal large model is a multi-modal large model obtained by fine-tuning and training for a large amount of urban governance scenario picture data;

[0067] The multi-modal large model receives visual picture and question text inputs and generates text corresponding to the picture.

[0068] In this embodiment, Qwen-VL and CogVLM-17B can be adopted. Qwen-VL and CogVLM-17B are a kind of multi-modal large model and belong to the base large model. The urban governance multi-modal large model in this embodiment can be based on this model, train and optimize this model to obtain the urban governance multi-modal large model.

[0069] In an embodiment of the present invention, extracting the object existence, classification label, and target location information in the output text includes:

[0070] For the description corresponding to any target in the output text, use the keyword recognition method to recognize the object existence and classification label corresponding to the target;

[0071] For the description corresponding to any target in the output text, use the identifier recognition method to recognize the target location information corresponding to the target;

[0072] Traverse each target in the output text to obtain the object existence, classification label, and target location information corresponding to each target.

[0073] Optionally, in order for those skilled in the art of the present invention to better understand the technical solution described in the embodiment of the present invention, the technical solution of the present invention will be described by way of examples below.

[0074] The output text of the urban governance multi-modal large model corresponds one-to-one with the sample picture input to the model. The output text of the urban governance multi-modal large model includes existence description, classification label description, and target location description; if there are multiple targets and multiple labels, the descriptive text corresponds one-to-one with the number of targets or the number of labels; if the text output by the urban governance multi-modal large model does not meet the requirements, an exception handling will be performed.

[0075] Each target and each label corresponds to a separate sentence. Taking the case of multiple targets as an example, the output text of the urban governance multimodal large model should be "The picture shows that the urban management problem is that there are shops (including various kiosks) on the sidewalk that place items outside the shops and cross the doors to occupy the road for business [[603,121,761,360]]", where [[603,121,761,360]] corresponds to [[Xmin,Ymin,Xmax,Ymax]], which represents a rectangular positioning box, where Xmin is the x-coordinate of the upper left corner of the positioning box, Ymin is the y-coordinate of the upper left corner of the positioning box, Xmax is the x-coordinate of the lower right corner of the positioning box, and Ymax is the y-coordinate of the lower right corner of the positioning box. There are shops (including various kiosks) on the sidewalk that place items outside the shops and cross the doors to occupy the road for business [[367,175,497,325]]. It can be found that the prediction output of the urban governance multimodal model contains two targets (occupancy of road operations [[603,121,761,360]] and occupation of road operations [[367,175,497,325]]). The urban governance multimodal model outputs text with a period as a string delimiter. Each string after segmentation can only contain one target. If multiple labels or multiple positioning boxes appear in the segmented string, the urban governance multimodal model is forced to regenerate the text output according to the format requirements.

[0076] The embodiment of the present invention defines "existence" and "non-existence" as existence description keywords. For example, the urban governance multimodal large model outputs the text "The picture shows that the urban management problem is that there are unauthorized installation of illegal circular-framed long-base slogan-type propaganda materials [[271,473,346,773]]", then the existence information of a single object can be extracted based on the keyword "existence". For another example, the urban governance multimodal large model outputs the text "There is no packaged garbage that has not been poured into garbage containers in public places", then the existence information of a single object can be extracted based on the keyword "non-existence".

[0077] The embodiment of the present invention defines the label text as a classification label description keyword, such as: "fallen bicycle", "drying items", "ground signboards set up in the square", "slogan propaganda materials", "carport-type peddlers", "occupying the road for business", "packaged garbage", "overflowing garbage", "scattered garbage", "cracked and damaged", "randomly piled materials", "shared bicycles", "exposed loess", "graffiti", "banner-type propaganda materials", "peddlers' tricycles", "peddlers' four-wheeled trucks", "peddlers occupying the road for business", "randomly parked non-motor vehicles", "large areas of water accumulation". For example, if the output text of the urban governance multimodal large model is "there are carport-type peddlers engaged in mobile business in public places [[752,522,999,815]]", the single object classification label information "carport-type peddlers" can be extracted.

[0078] In the embodiments of the present invention, "[[]]" is defined as the target position description identifier. For example, if the output text of the urban governance multimodal large model is "There are disorderly piled materials of box type in public places [[445,735,538,907]]", the target position information of a single object "445,735,538,907" can be extracted.

[0079] Continue to loop through the extraction process of a single target, and extract the existence, classification labels, and target position information of all targets from the output text of the urban governance multimodal large model.

[0080] In the embodiments of the present invention, the structured dictionary list is as follows:

[0081] Construct a list, where each element in the list corresponds to a picture, and the elements of the list are structured dictionaries;

[0082] Among them, the structured dictionary list stores the picture name, picture size attribute, existence keyword, label keyword, and target box;

[0083] Specifically:

[0084] Such as Figure 3As shown in the figure, an embodiment of the present invention provides a schematic diagram of a structured dictionary list. If the number of pictures is N, then the length of the list is N. The elements of the list are structured dictionaries dict_elei, where the subscript i represents the picture index. The keys of the structured dictionary dict_ele include: picture name Name, picture size Size, existence Obj, label Label, and target box Box. If the target label does not exist in the picture, the key-value pair of the existence keyword is Obj: 0. Taking two pictures as an example to illustrate the structured dictionary list structure, the output of the urban governance multimodal large model corresponding to Picture 1 is "The picture shows that the urban management problem is that there is exposed loess where the green vegetation does not completely cover [[000,826,170,988]]", then the structured dictionary dict_ele1 is {Name: "name1.jpg", Size: [1600, 900], Obj: 1, Label: "Exposed Loess", Box: [000, 826, 170, 988]}. The output of the urban governance multimodal large model corresponding to Picture 2 is "The picture shows that the urban management problem is that there is packaged garbage in public places that has not been put into the garbage container [[245,833,305,916]]", then the structured dictionary dict_ele2 is {Name: "name2.jpg", Size: [1600, 900], Obj: 1, Label: "Packaged Garbage", Box: [245, 833, 305, 916]}. Then the output results of all pictures are a list [{Name: "name1.jpg", Size: [1600, 900], Obj: 1, Label: "Exposed Loess", Box: [000, 826, 170, 988]}, {Name: "name2.jpg", Size: [1600, 900], Obj: 1, Label: "Packaged Garbage", Box: [245, 833, 305, 916]}].

[0085] The present invention constructs a structured dictionary, which is more convenient for the urban governance multimodal large model

[0086] In the embodiment of the present invention, according to the data in the structured dictionary list and the text labels corresponding to the sample pictures, the accuracy index of the urban governance multimodal large model for a single sample picture is obtained as follows:

[0087]

[0088] Among them, IOU i represents the intersection over union of the target box of the i-th target in a single sample picture, represents the true classification label of the i-th target in a single sample picture, represents the predicted classification label of the i-th target in the text label corresponding to the sample picture, Indicates the predicted classification label and the true classification label The text cosine similarity between them Indicates the quantization value of whether the target exists in the true result Indicates the quantization value of whether the target exists in the predicted result corresponding to the text label. If the target exists in the result, then obj i = 1. If the target does not exist in the result, then obj i = 0. N represents the total number of targets in a single sample image

[0089] As Figure 4 shown, an embodiment of the present invention provides a schematic diagram for calculating the accuracy index of a single image. The true value of the image is "There are no unlicensed peddlers without violations shown in the picture [[190,425,480,866]]." The predicted value output by the urban governance multimodal large model is "There are peddlers selling goods shown in the picture [[188,411,488,871]]." Since the existence description in the true value is "does not exist", the quantization value obj t of whether the target exists in the true result is 0. Since the existence description in the predicted value is "exists", the quantization value obj p of whether the target exists in the predicted result is 1. Similarly, the true classification label lab t is unlicensed peddler, and the predicted classification label lab p is peddler. The intersection over union IOU (coordinate coincidence degree) of the target boxes of the true value coordinates "[190,425,480,866]" and the predicted value coordinates "[188,411,488,871]" is 0.55. The number of targets N in the picture is 1. Substituting the above values into the accuracy index formula of a single image, A simple can be calculated to be 0

[0090] In an embodiment of the present invention, the comprehensive evaluation index for evaluating the urban governance multimodal large model on the sample data according to the accuracy indexes of multiple single sample images is as follows:

[0091]

[0092] where M represents the total number of sample images in the sample data, λ represents the influence coefficient of the single image accuracy, represents the text label corresponding to the j-th sample image in the sample data, the predicted result text description of the j-th sample image in the sample data corresponding to the urban governance multimodal large model, represents the true result text description and the predicted result text description The cosine similarity between texts, where η represents the influence coefficient of text similarity, is the accuracy rate of a single sample picture.

[0093] Optionally, the accuracy rate of a single sample picture is relatively reliable, and the influence coefficient λ of the single-picture accuracy rate is set to 0.8; the credibility of the text cosine similarity Sim is not high, and the influence coefficient η of the text similarity is set to 0.2.

[0094] An embodiment of the present invention provides a case. In this case, the true label is: There are unlicensed peddlers violating regulations in the picture [316, 319, 613, 456]. The prediction output result of the urban governance multi-modal large model is: There are no peddlers selling goods in the picture [325, 302, 640, 443]. If the pure text similarity measurement method is adopted, the accuracy rate of a single sample picture is 0.69. There is only one target object in the picture, the intersection over union IOU of the coordinates is 0.72, the existence quantification value of the true label is 1, the existence quantification value of the prediction result of the urban governance multi-modal large model is 0, and the classification accuracy (text similarity) between the classification label "" and "unlicensed peddler" is 0.68. According to the single-picture accuracy rate index formula, the exclusive OR value of the existence quantification values is 0. Therefore, the accuracy rate of this picture is 0. Obviously, it is more accurate to calculate the index by using the text labeling method.

[0095] Another embodiment of the present invention provides a case. In this case, the true label is: There are shops (including various kiosks) on the road surface, placing items outside the store and operating across the door and occupying the road [[066, 507, 178, 707]]. The predicted text result of the urban governance multi-modal large model is: There are vendors engaging in road occupation behavior in public places [[273, 587, 423, 926]]. According to the calculation of the intersection over union IOU of the positioning box, it is 0, which indicates that the positioning box calculation fails to match the predicted box and the true box. Therefore, the accuracy rate of a single picture is 0.

[0096] As Figure 2 shown, an embodiment of the present invention provides an evaluation device for an urban governance multi-modal large model based on text labeling, including: a data acquisition module 201, an output acquisition module 202, a structured dictionary acquisition module 203, a first index acquisition module 204, and a second index acquisition module 205;

[0097] The data acquisition module 201 is used to acquire sample data for evaluating the urban governance multi-modal large model; among them, the sample data includes sample pictures and text labels corresponding to the sample pictures;

[0098] The output acquisition module 202 is used to input the sample pictures in the sample data into the urban governance multi-modal large model to obtain the output text corresponding to the urban governance multi-modal large model;

[0099] The structured dictionary acquisition module 203 is used to construct a list, where each element in the list corresponds to an image, and the elements of the list are structured dictionaries; the structured dictionary list stores the image name, image size attributes, existence keywords, label keywords, and target boxes.

[0100] The first metric acquisition module 204 is used to obtain the accuracy metric of the urban governance multi-modal large model for a single sample image according to the data in the structured dictionary list and the text label corresponding to the sample image.

[0101] The second metric acquisition module 205 is used to obtain the comprehensive evaluation metric for evaluating the urban governance multi-modal large model on the sample data according to the accuracy metric of the single sample image.

[0102] The urban governance multi-modal large model evaluation device described in the embodiments of the present invention can execute the above method embodiments, and the beneficial effects produced are the same, so details will not be described here.

[0103] In the embodiments of the present invention, the structured dictionary acquisition module 203 includes a data extraction unit and a dictionary acquisition unit;

[0104] The data extraction unit is used to identify the object existence and classification label corresponding to the target for any description of the target in the output text by using the keyword recognition method; for any description of the target in the output text, the identifier recognition method is used to identify the target position information corresponding to the target; traverse each target in the output text to obtain the object existence, classification label, and target position information corresponding to each target.

[0105] The dictionary acquisition unit is used to construct a list, where each list element corresponds to an image, and the list element stores the object existence, classification label, and target position information, so as to obtain the structured dictionary stored in the list.

[0106] In the embodiments of the present invention, the first metric acquisition module 204 includes: a prediction data acquisition unit, a ground truth data acquisition unit, and an accuracy metric acquisition unit;

[0107] The prediction data acquisition unit is used to determine whether there are quantization values for the predicted bounding box, predicted classification label, and prediction result in a single sample image according to the data in the structured dictionary, so as to obtain the prediction data corresponding to the single sample image.

[0108] The real data acquisition unit is used to determine the real prediction box, real classification label, and whether there is a quantization value in the real result in the text label according to the text label corresponding to the sample picture, and obtain the real data corresponding to a single sample picture;

[0109] The accuracy rate index acquisition unit is used to obtain the accuracy rate index of a single sample picture according to the prediction data corresponding to the single sample picture and the real data corresponding to the single sample picture as:

[0110]

[0111] Among them, IOU i represents the intersection over union of the target box of the i-th target in a single sample picture, represents the real classification label of the i-th target in a single sample picture, represents the predicted classification label of the i-th target in the text label corresponding to the sample picture, represents the predicted classification label and the real classification label the text cosine similarity between them, represents the quantization value of whether the target exists in the real result, represents the quantization value of whether the target exists in the predicted result corresponding to the text label. If the target exists in the result, then obj i = 1. If the target does not exist in the result, then obj i = 0, and N represents the total number of targets in a single sample picture.

[0112] In the embodiment of the present invention, the second index acquisition module 205 includes a text similarity prediction unit and a comprehensive evaluation index acquisition unit;

[0113] The text similarity prediction unit is used to determine the text cosine similarity between the real result text description and the predicted result text description according to the sample picture in the sample data and the text label corresponding to the sample picture;

[0114] The comprehensive evaluation index acquisition unit is used to determine the comprehensive evaluation index of the urban governance multi-modal large model according to the text cosine similarity corresponding to multiple single sample pictures and the accuracy rate index as:

[0115]

[0116] Among them, M represents the total number of sample pictures in the sample data, λ represents the influence coefficient of the single accuracy rate, represents the text label corresponding to the j-th sample picture in the sample data, the predicted result text description of the urban governance multi-modal large model corresponding to the j-th sample picture in the sample data, Indicates the description of the true result text and the description of the predicted result text The cosine similarity between texts, η represents the influence coefficient of text similarity, is the accuracy rate of a single sample image.

[0117] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.

[0118] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0119] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0121] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0122] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for evaluating a large multimodal model of urban governance based on text labeling, characterized in that: include: Obtain sample data for evaluating a large multimodal model of urban governance; wherein the sample data includes sample images and text labels corresponding to the sample images; Input the sample images in the sample data into the urban governance multimodal large model to obtain the output text corresponding to the urban governance multimodal large model; Extracting object existence, classification labels, and target location information from the output text, and obtaining a structured dictionary list based on the object existence, classification labels, and target location information; According to the data in the structured dictionary list and the text labels corresponding to the sample images, the accuracy index of the urban governance multimodal large model in a single sample image is obtained, specifically including: Among them, IOU i represents the target box intersection-union ratio of the i-th target in a single sample image, represents the true classification label of the i-th target in a single sample image, Represents the predicted classification label of the i-th target in the text label corresponding to the sample image, Represents the predicted classification label and the true classification label The cosine similarity of the texts between A quantitative value indicating whether the target exists in the real result. A quantitative value indicating whether the target exists in the prediction result corresponding to the text label. N represents the total number of targets in a single sample image. Based on the accuracy indicators of multiple single sample images, comprehensive evaluation indicators for evaluating the multimodal large model of urban governance on sample data are obtained.

2. According to the text labeling-based urban governance multimodal large model evaluation method of claim 1, it is characterized in that: The urban governance multimodal large model is a multimodal large model obtained through fine-tuning training for urban governance scene image data; The multimodal large model receives visual images and question text inputs and generates text corresponding to the images.

3. According to the text labeling-based urban governance multimodal large model evaluation method of claim 1, it is characterized in that: Extracting object existence, classification labels, and target location information from the output text includes: For the description corresponding to any target in the output text, a keyword recognition method is used to identify the existence and classification label of the object corresponding to the target; For the description corresponding to any target in the output text, an identifier recognition method is used to identify the target position information corresponding to the target; Traverse each target in the output text to obtain the object existence, classification label and target location information corresponding to each target.

4. According to the text labeling-based urban governance multimodal large model evaluation method of claim 3, it is characterized in that: The structured dictionary list is: Build a list, each element in the list corresponds to a picture, and the elements of the list are structured dictionaries; The structured dictionary list stores image names, image size attributes, existence keywords, label keywords, and target boxes.

5. According to the text labeling-based urban governance multimodal large model evaluation method of claim 1, it is characterized in that: According to the accuracy index of multiple single sample images, the comprehensive evaluation index of evaluating the urban governance multimodal large model on the sample data is obtained as follows: Among them, M represents the total number of sample images in the sample data, λ represents the influence coefficient of single accuracy, Indicates the text label corresponding to the jth sample image in the sample data, The jth sample image in the sample data corresponds to the text description of the prediction result of the urban governance multimodal large model, Indicates the actual result text description And the prediction result text description The cosine similarity of the texts between them, η represents the influence coefficient of text similarity, is the accuracy of a single sample image.

6. A large multi-modal model evaluation device for urban governance based on text labeling, characterized in that: include: A data acquisition module, an output acquisition module, a structured dictionary acquisition module, a first indicator acquisition module, a second indicator acquisition module and an evaluation result acquisition module; The data acquisition module is used to acquire sample data for evaluating the multimodal large model of urban governance; wherein the sample data includes sample images and text labels corresponding to the sample images; The output acquisition module is used to input the sample images in the sample data into the urban governance multimodal large model to obtain the output text corresponding to the urban governance multimodal large model; The structured dictionary acquisition module is used to extract the object existence, classification label and target location information in the output text, and obtain a structured dictionary list according to the object existence, classification label and target location information; The first indicator acquisition module is used to obtain the accuracy indicator of the urban governance multimodal large model in a single sample image according to the data in the structured dictionary list and the text label corresponding to the sample image; The first indicator acquisition module includes: a prediction data acquisition unit, a real data acquisition unit and an accuracy indicator acquisition unit; The prediction data acquisition unit is used to determine whether there are quantized values ​​for the prediction box, the prediction classification label, and the prediction result in the single sample image according to the data in the structured dictionary, and obtain the prediction data corresponding to the single sample image; The real data acquisition unit is used to determine whether there are quantized values ​​for the real prediction box, the real classification label, and the real result in the text label according to the text label corresponding to the sample image, and obtain the real data corresponding to the single sample image; The accuracy index acquisition unit is used to obtain the accuracy index of a single sample image according to the predicted data corresponding to the single sample image and the real data corresponding to the single sample image: Among them, IOU i represents the target box intersection-union ratio of the i-th target in a single sample image, represents the true classification label of the i-th target in a single sample image, Represents the predicted classification label of the i-th target in the text label corresponding to the sample image, Represents the predicted classification label and the true classification label The cosine similarity of the texts between A quantitative value indicating whether the target exists in the real result. A quantitative value indicating whether the target exists in the prediction result corresponding to the text label. N represents the total number of targets in a single sample image. The second indicator acquisition module is used to obtain a comprehensive evaluation indicator for evaluating the urban governance multimodal large model on sample data based on the accuracy indicator of the single sample image.

7. The urban governance multimodal large model evaluation device based on text labeling according to claim 6 is characterized in that: The structured dictionary acquisition module includes a data extraction unit and a dictionary acquisition unit; The data extraction unit is used to use a keyword recognition method to identify the existence and classification label of the object corresponding to the target for the description corresponding to any target in the output text; and use an identifier recognition method to identify the target location information corresponding to the target for the description corresponding to any target in the output text; Traverse each target in the output text to obtain the object existence, classification label and target location information corresponding to each target; The dictionary acquisition unit is used to construct a list, each element in the list corresponds to a picture, and the elements of the list are structured dictionaries; wherein the structured dictionary list stores picture names, picture size attributes, existence keywords, label keywords and target boxes.

8. The urban governance multimodal large model evaluation device based on text labeling according to claim 7 is characterized in that: The second indicator acquisition module includes a text similarity prediction unit and a comprehensive evaluation indicator acquisition unit; The text similarity prediction unit is used to determine the text cosine similarity between the real result text description and the predicted result text description according to the sample pictures in the sample data and the text labels corresponding to the sample pictures; The comprehensive evaluation index acquisition unit is used to determine the comprehensive evaluation index of the urban governance multimodal large model according to the text cosine similarity and accuracy index corresponding to multiple single sample images: Among them, M represents the total number of sample images in the sample data, λ represents the influence coefficient of single accuracy, Indicates the text label corresponding to the jth sample image in the sample data, The jth sample image in the sample data corresponds to the text description of the prediction result of the urban governance multimodal large model, Indicates the actual result text description And the prediction result text description The cosine similarity of the texts between them, η represents the influence coefficient of text similarity, is the accuracy of a single sample image.

Citation Information

Patent Citations

  • Classification method of urban governance events, and training method and device of classification model

    CN116992323A

  • Instant messaging data processing method and device, storage medium and terminal equipment

    CN117278512A