Evaluation Method and Open Evaluation Platform for Vision-Language Models

A unified evaluation framework and platform for visual language models addresses the inconsistency in evaluating these models by using standardized data sets and metrics, ensuring fair and comprehensive assessments across tasks and domains.

CN119988915BActive Publication Date: 2025-07-15ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510481339.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-15
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In the prior art, the evaluation framework of visual language models is not unified, making it difficult to fairly evaluate the advantages and disadvantages of each visual language model in different visual language tasks, mainly due to the lack of unified standards for data sets and evaluation indicators.

Method used

Provide a public evaluation platform and evaluation method. By building a unified instruction follow data set and evaluation indicators, we ensure that the model is evaluated in a unified evaluation environment, including evaluation request reception, instruction follow data set construction, model inference and evaluation modules, and support the evaluation of multiple visual-language tasks.

Benefits of technology

A systematic and standardized evaluation system is realized, which can fairly evaluate the advantages and disadvantages of visual language models in different tasks, simplifies the evaluation process, reduces research costs, and provides a fair performance benchmark.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988915B_ABST
    Figure CN119988915B_ABST
Patent Text Reader

Abstract

This specification provides a method for evaluating a vision-language model and an open evaluation platform. The method includes: receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the open evaluation platform; each task in the task set corresponds to an instruction-following data set, and any sample in the instruction-following data set includes an image, an instruction, and an answer. Obtaining a target data set corresponding to the target task from the instruction-following data set corresponding to the task set, and the target model performs inference on the target data set to obtain an inference result. According to the inference result and the target evaluation metrics of each target task, determining a first evaluation score of the target model on each target task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and particularly to an evaluation method and an open evaluation platform for vision-language models. Background Art

[0002] A vision-language model (VLM) is a multi-modal artificial intelligence model that can process both visual and language information simultaneously. It is mainly applied to scenarios such as image classification, visual question answering, image captioning, and spatial change detection.

[0003] Nowadays, various frameworks of vision-language models emerge in an endless stream, and correspondingly, there is a need to evaluate the processing performance of various vision-language models on different vision-language tasks. When technicians evaluate the vision-language models they developed, they generally collect public datasets or build their own datasets alone and set evaluation metrics alone.

[0004] When different technicians evaluate the vision-language models they developed, even when evaluating the same vision-language task, the datasets and evaluation metrics used may be different. Due to the lack of a unified evaluation framework, such as the lack of a unified standard for the datasets used and inconsistent evaluation metrics, it is difficult to fairly evaluate the advantages and disadvantages of each vision-language model on different vision-language tasks. Summary of the Invention

[0005] To overcome the problems existing in the related art, this specification provides an evaluation method and an open evaluation platform for vision-language models.

[0006] According to the first aspect of the embodiments of this specification, an evaluation method for a vision-language model is provided. The method is applied to an open evaluation platform and includes:

[0007] Receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the open evaluation platform; each task in the task set corresponds to an instruction-following dataset, and any sample in the instruction-following dataset includes an image, an instruction, and an answer;

[0008] Obtaining a target dataset corresponding to the target task from the instruction-following dataset corresponding to the task set, and the target model performs inference on the target dataset to obtain an inference result;

[0009] Determining a first evaluation score of the target model on each target task according to the inference result and the target evaluation metrics of each target task.

[0010] According to the first aspect of the embodiments of this specification, an open evaluation platform is provided, including:

[0011] An evaluation request receiving module, configured to receive a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform;

[0012] An instruction following dataset construction module, configured to construct an instruction following dataset corresponding to each task in the task set, where any sample in the instruction following dataset includes an image, an instruction, and an answer;

[0013] A model inference module, configured to obtain a target dataset corresponding to the target task from the instruction following dataset corresponding to the task set, and the target model performs inference on the target dataset to obtain an inference result;

[0014] A model evaluation module, configured to determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation metrics of each target task.

[0015] The technical solution provided by the embodiments of this specification may include the following beneficial effects:

[0016] In the embodiments of this specification, for the user's need to evaluate the performance of a target model on at least one target task in a task set provided by a public evaluation platform, this solution provides the user with an instruction following dataset and evaluation metrics corresponding to the target task. The target model performs inference on the unified instruction following dataset to obtain an inference result, and based on the inference result and the unified target evaluation metrics, determines the evaluation score of the target model on the target task. It can be seen that this solution provides the user with a systematic and standardized evaluation system, which can ensure a fair evaluation of the advantages and disadvantages of each vision-language model on different tasks.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0019] Figure 1 is an application scenario diagram of a public evaluation platform shown according to an exemplary embodiment of this specification.

[0020] Figure 2 is a flowchart of a method for evaluating a vision-language model shown according to an exemplary embodiment of this specification.

[0021] Figure 3 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of this specification.

[0022] Figure 4 It is a block diagram of an evaluation device for a vision-language model shown in this specification according to an exemplary embodiment. Detailed implementation manners

[0023] A vision-language model (VLM) is a multi-modal artificial intelligence model that can process visual and language information simultaneously. Its main applications include scenarios such as image classification, visual question answering, image captioning, and spatial change detection.

[0024] Nowadays, various frameworks of vision-language models emerge in an endless stream, and correspondingly, there is a need to evaluate the processing performance of various vision-language models on different vision-language tasks. When technicians evaluate the vision-language models they have developed, they generally collect public datasets or build their own datasets alone and set evaluation metrics alone.

[0025] When different technicians evaluate the vision-language models they have developed, even when evaluating the same vision-language task, the datasets and evaluation metrics used may be different. Due to the lack of a unified evaluation framework, such as the lack of a unified standard for the datasets used and inconsistent evaluation metrics, it is difficult to fairly evaluate the advantages and disadvantages of each vision-language model on different vision-language tasks.

[0026] To address the above technical problems, this specification provides an evaluation method and an open evaluation platform for vision-language models, by providing a systematic and standardized evaluation system to ensure fair evaluation of the advantages and disadvantages of each vision-language model on different tasks.

[0027] Figure 1 It is an application scenario diagram of an open evaluation platform shown in this specification according to an exemplary embodiment. As Figure 1 shown, the server 10 communicates with each client 12 - 14 through the network 11, and the open evaluation platform can be deployed on the server 10. The server 10 can provide a front-end access page to the client to realize the interaction between the client and the open evaluation platform. For example, the front-end access page can at least provide a model upload module and a task selection module. Among them, the model upload module can be used to upload the model to be evaluated, and the task selection module can be used to display various vision-language tasks provided by the open evaluation platform. Users can select at least one target task from the task selection module. The front-end access page can submit the model input by the user and the selected target task to the open evaluation platform, and the open evaluation platform can return at least the evaluation results to the front-end access page after completing the evaluation.

[0028] The application scenario of the public evaluation platform can be the evaluation of visual language models. Corresponding instruction following datasets and evaluation indicators can be constructed on the public evaluation platform for different visual-language tasks. Among them, each visual-language task can correspond to multiple instruction following datasets and multiple evaluation indicators to ensure that the performance of the model in a wide range of scenarios is evaluated and the model performance is measured more comprehensively.

[0029] Of course, for the application scenarios of visual language models, multiple capability dimensions can also be designed for the performance of visual language models, and multiple visual-language tasks can be divided for each capability dimension to comprehensively evaluate the performance of the model from multiple aspects such as tasks, capability dimensions, and comprehensive dimensions, so as to identify the shortcomings or advantages of the model to be evaluated in specific tasks, and promote developers to optimize the model in a targeted manner. In addition, the public evaluation platform provides a unified capability dimension and task classification to provide a benchmark for performance comparison between different models, which promotes the transparency and fairness of evaluation in the community studying visual language models.

[0030] After receiving the target model to be evaluated and at least one target task selected by the user from the task set provided by the public evaluation platform, the public evaluation platform can evaluate the performance of the target model on each target task and return the evaluation result to the user.

[0031] Compared with technicians independently completing repetitive processes such as instruction following datasets and evaluation indicators, the public evaluation platform can integrate instruction following datasets and evaluation indicators of different visual-language tasks on the public evaluation platform for technicians to directly call, simplifying the evaluation process, and the instruction following datasets and evaluation indicators used by different technicians are consistent, thereby ensuring the consistency and comparability of the evaluation results. In addition, for technicians who are new to the field of visual language models, the public evaluation platform can provide ready-made datasets and evaluation indicators, thereby reducing the initial cost of research.

[0032] Next, we will introduce the functional modules in the public evaluation platform in detail, including:

[0033] The evaluation request receiving module is used to receive the target model to be evaluated and at least one target task selected by the user from the task set provided by the public evaluation platform.

[0034] The instruction following data set construction module is used to construct an instruction following data set corresponding to each task in the task set, and any sample in the instruction following data set includes an image, an instruction and an answer.

[0035] The model reasoning module is used to obtain a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and the target model performs reasoning on the target data set to obtain a reasoning result.

[0036] A model evaluation module, configured to determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation metrics of each target task.

[0037] In one embodiment, the evaluation request receiving module may further receive other parameters. The other parameters may include a target instruction-following dataset and target evaluation metrics selected by the user for the target task after selecting the target task. Then, the first evaluation score of the target model determined by the model evaluation module on each target task may be related to the dataset source and the target evaluation metrics. For example, the first evaluation score may be associated with the data source and the evaluation metrics. Of course, even if the user makes no selection, the data source and the evaluation metrics may also be associated when determining the first evaluation score.

[0038] Although vision-language models have achieved great success in the field of natural images, for images in some technical fields, they still lack corresponding instruction-following datasets. For example, for remote sensing images, most of the currently publicly available datasets in the remote sensing field are image annotation datasets, and any sample in the image annotation dataset includes an image and an annotation result. However, since the image annotation datasets in the remote sensing field cannot be directly applied to the evaluation of vision-language models, the application of vision-language models in the remote sensing field is restricted.

[0039] Therefore, this specification provides a method for converting an image annotation dataset into an instruction-following dataset:

[0040] In one embodiment, an image annotation dataset is obtained, where any sample in the image annotation dataset includes an image and an annotation result. The image annotation dataset is converted into an instruction-following dataset, where, for different task types, instructions and answers corresponding to the task types are determined according to the annotation results of the image annotation dataset. In this embodiment, by converting the image annotation datasets in different technical fields into corresponding instruction-following datasets for application to the research of vision-language models, it is not necessary to build an instruction-following dataset from scratch for this technical field, thereby saving the cost of building the instruction-following dataset.

[0041] For example, as shown in Table 1, the corresponding relationships between the image annotation dataset and the instruction-following dataset for different task types are shown:

[0042] Table 1

[0043]

[0044] The key to converting an image annotation dataset into an instruction-following dataset is to convert the annotation results in the image annotation dataset into instructions and answers. For example, for a dataset of an image classification task, which includes images and class annotations, for the class annotation of any image in the dataset, the class annotation can be converted into instructions and answers, and the instructions and answers can be replaced in the original image annotation dataset to obtain an instruction-following dataset. Another example is that for an image annotation dataset of a segmentation task, whose annotation results include the pixel point annotation of the target area and the corresponding class, then for the task type to be converted, the annotation results can be converted into instructions and answers corresponding to the task type. Of course, the conversion methods for datasets of other task types are similar and will not be elaborated here.

[0045] In one embodiment, an image annotation dataset applied to a certain original task is not limited to being converted into an instruction-following dataset of the same task, but can be extended to other task types and correspondingly generate instruction-following datasets of other task types. In other words, an instruction-following dataset of at least one task type can be generated according to an image annotation dataset of a certain task type, and the at least one task type includes the certain task type.

[0046] For example, as shown in Table 2, the corresponding relationship between the original task and the extended task of the image annotation dataset is shown:

[0047] Table 2

[0048]

[0049] For example, referring to Table 2, for an image annotation dataset of image description, the image annotation dataset can be converted into an instruction-following dataset of the image short description task and an instruction-following dataset of the image detailed description task. Similarly, for an image annotation dataset of a segmentation task, in addition to being able to convert the image annotation dataset into an instruction-following dataset of the segmentation task, it can also be converted into an instruction-following dataset of the target calculation task and an instruction-following dataset of the polygon area classification task.

[0050] Therefore, when converting an image annotation dataset into an instruction-following dataset, essentially the annotation results of the image dataset are converted into instructions and answers, and the specific content of the instructions and answers depends on the target task type to be converted. For example, for the same annotation result in an image annotation dataset of a segmentation task, when it is converted into different task types, the obtained instructions and answers are different. For example, corresponding to the segmentation task, instructions and answers can be obtained according to the annotation result, and also instructions and answers corresponding to the target counting task can be obtained, etc.

[0051] In one embodiment, when determining instructions and answers corresponding to different task types according to the annotation results of an image annotation dataset, for any task type, an instruction template and an answer template corresponding to the task type are obtained. If the instruction template includes a first placeholder, a first object is determined from the annotation results of the image annotation dataset, and the first placeholder in the instruction template is replaced with the first object to obtain the instruction; if the instruction template does not include a first placeholder, the instruction template is used as the instruction. A second object is determined from the annotation results of the image annotation dataset, and the second placeholder in the answer template is replaced with the second object to obtain the answer; wherein, the answer template includes a second placeholder for replacement with the second object.

[0052] For example, as shown in Table 3, instruction-answer templates for different task types are presented:

[0053] Table 3

[0054]

[0055] For example, for the horizontal box object detection task, the obtained instruction template corresponding to the horizontal box detection task is: "Detect all <class>in the image”; Answer template: " <region>”. Among them, the first placeholder in the instruction template is <class>, the second placeholder in the answer template is <region>. Determine a first object and a second object respectively from the annotation results in the image annotation dataset for the horizontal box object detection task. Wherein, the first object is " <class>", specifically the target category of the labeled horizontal box; the second object is " <box> <x1> <y1> <x2> <y2>< / y2> < / x2> < / y1> < / x1> < / box> ", (X1, Y1), (X2, Y2) are the coordinates of the upper left corner and the lower right corner of the horizontal box respectively. Replace the first placeholder in the instruction template with the first object as the instruction, and replace the second placeholder in the answer template with the second object as the answer. It should be noted that the above is an explanation for any horizontal box in the labeling result. For other horizontal boxes in the labeling result, their conversion methods can all be carried out according to the above method, which will not be elaborated here one by one.

[0056] For example, for the rotated box target detection task, the instruction template corresponding to the rotated box target detection task can be obtained as: "Detect all <class>in the image. Use oriented bounding boxes”; Answer template: " <region>”. Among them, the first placeholder in the instruction template is <class>, the second placeholder in the answer template is <region>。Determine a first object and a second object respectively from the annotation results in the image annotation dataset for the rotating box object detection task. Among them, the first object is " <class>", which is the target category of the marked rotation box; the second object is " <quad> <x1> <y1> <x2> <y2> <x3> <y3> <x4> <y4>< / y4> < / x4> < / y3> < / x3> < / y2> < / x2> < / y1> < / x1> < / quad> ", and (X1, Y1), (X2, Y2), (X3, Y3), and (X4, Y4) are the coordinates of the four vertices of the rotation box respectively. Replace the first placeholder in the instruction template with the first object to obtain the instruction, and replace the second placeholder in the answer template with the second object to obtain the answer.

[0057] For example, for an image classification task, the obtained instruction template corresponding to the image classification task is: "Classify the image. Use one or a few words"; the answer template is: " <class>”. If the first placeholder is not included in the instruction template, the instruction template is directly used as an instruction; determine the second object from the annotation results of the image annotation dataset as " <class>", as the category of the image, then replace the second placeholder in the response template with the second object and use it as the response.

[0058] Similarly, for the image annotation dataset of the visual localization task, the description corresponding to the bounding box and the coordinate information of the bounding box can be determined from the annotation results in the image annotation dataset. For the image annotation dataset of the segmentation task, the segmentation mask image in the annotation results can be converted into the text representation required by the template, and the target category corresponding to the mask can be determined from it.

[0059] For example, for the image annotation dataset of certain specific tasks, when constructing the corresponding instruction-following dataset, an instruction-following dataset with a different task type from it can be constructed. Specifically, refer to Table 2.

[0060] For the extended task type, the instruction template and response template corresponding to the extended task type can be obtained. For example, for the region description task extended from the visual localization task, the instruction template and response template corresponding to the region description task can be obtained. The instruction template is "Describe the <region>"in this image”, the answer template is" <description>”. The annotation results of the image annotation dataset for the visual localization task include the description of the target area <description>The positioning coordinates of the target <region>, such as the target description <description>is "a red vehicle”, and the positioning coordinates of the target <region>For " <box> <125><369><256><489>< / box> ", the first object determined from the annotation results of the image annotation dataset is " <box> <125><369><256><489>< / box> ". The first placeholder in the instruction template can be replaced with the first object to be used as an instruction. The second object determined from the annotation results of the image annotation dataset is "a red vehicle". The second placeholder in the answer template can be replaced with the second object to be used as an answer.

[0061] Similar to the above extended task, for the object counting task extended from the horizontal bounding box object detection task, obtain the instruction template and answer template corresponding to the object counting task. Among them, the instruction template is "Count the numberof <class>", the first placeholder in the instruction template is <class>; The response template is " <number>", the second placeholder in the answer template is " <number>”. Determine a first object from the annotation results of the image annotation dataset corresponding to the horizontal box object detection task, where the first object is of the type in the annotation results as <class>horizontal frame, such as this <class>It can be a car. In this case, the car is the first object. Replace the first placeholder in the instruction template with "car", and replace the second placeholder in the answer template with the number of cars as the answer. The number of cars can be obtained by counting in the annotation result.

[0062] Of course, for some extended tasks, the instruction template and the answer template may not be used. For example, there are tasks of existence judgment and quantity comparison. To judge whether a certain target exists in the image, such as "Is a forest present in the image?", the answer is "yes" or "no". Quantity comparison is to judge the quantitative relationship between two types of targets in the image, such as "Are there more roads than commercial buildings?", and the answer is "yes" or "no". Then the answer can be determined according to the annotation result in the image annotation dataset.

[0063] For the image annotation dataset of the image description task, it can be converted into an instruction-following dataset for the image detailed description task and / or an instruction-following dataset for the image short description task. Among them, obtain the image description text from the annotation result in the image annotation dataset, and perform deduplication on the image description text. If the number of sentences in the deduplicated image description text is not less than the first threshold or the number of words in the image description text is not less than the second threshold, then use the image description text as the answer of the instruction-following dataset for the detailed description task, and randomly select one sentence from the unused sentences as the answer of the instruction-following dataset for the image short description task. If the number of sentences in the deduplicated image description text is less than the first threshold or the number of words in the image description text is less than the second threshold, then use the deduplicated image description text as the answer of the instruction-following dataset for the image short description task.

[0064] Exemplarily, for the expansion of the image description task into the image short description task and the image detailed description task, due to the uneven quality of the description annotations in the open-source dataset and the existence of duplicate descriptions, a deduplication operation is first performed. For a piece of image description, calculate the similarity between each sentence through the word intersection ratio of the bag-of-words model. It is also possible to choose to use Jaccard similarity, TF-IDF cosine similarity or sentence vectors to measure the similarity of sentences. Merge the sentences judged to be dissimilar into one sentence, and judge whether it is a detailed description or a short description. Judgment rule: A description sentence with less than or equal to 4 sentences or a sentence length less than or equal to 30 words is defined as a short description, and a description sentence with more than 4 sentences or a sentence length greater than 30 words is defined as a detailed description. If the merged description is a detailed description and there are unused similar sentences, then select one sentence from the remaining similar sentences as the short description; if it is a short sentence, then discard all similar sentences.

[0065] For different task types, which can reflect the capabilities of a vision-language model in different ability dimensions, this specification provides an exemplary way to divide the ability dimensions for evaluating a vision-language model. Of course, other ways of dividing ability dimensions can also be adopted, and this specification does not impose any restrictions on this.

[0066] For example, as shown in Table 4, Table 4 shows tasks under different ability dimensions:

[0067] Table 4

[0068]

[0069] In one embodiment, when the model evaluation module determines the first evaluation scores of the target model on each target task according to the inference results and the target evaluation metrics of each target task, in addition to determining the first evaluation scores of the target model on each target task, it can also determine the target ability dimensions to which each target task belongs. For any target ability dimension, according to the first evaluation scores of each target task under this target ability dimension, determine the second evaluation score of the target model on this target ability dimension. And / or, according to the second evaluation scores on each target ability dimension, determine the third evaluation score of the target model.

[0070] In this embodiment, when evaluating the model's capabilities, not only the performance of a single task is concerned, but also the comprehensive performance of the model in multiple ability dimensions is examined, which helps to identify the overall advantages and potential weaknesses of the model.

[0071] In one embodiment, the evaluation platform of the present disclosure may further include an evaluation score storage module for storing the evaluation scores of the models whose evaluation is completed; and an evaluation result generation module for generating an evaluation result according to the evaluation scores of other models and the target model on the same task.

[0072] Among them, the evaluation result can be in the form of a chart to show the evaluation results of the target model and other models in different granularity dimensions such as different instruction-following data sets, evaluation metrics, tasks, and ability dimensions.

[0073] In one embodiment, the disclosed evaluation platform may further include a test trigger module for receiving a dataset / model update request submitted by a user to the disclosed evaluation platform. A model inference module is further configured to, when receiving a model update instruction, load the datasets of all tasks that the model can process in the instruction-following dataset and start an inference task; when receiving a dataset update, call all models capable of handling the tasks of the dataset and start an inference task. The instruction-following dataset can be centrally managed in a server through a distributed storage system, and the models are dynamically deployed on any node within the cluster. All model services can follow a preset communication interface template for communication between servers, supporting cross-server data interaction and result feedback (sending and receiving data and results). A model evaluation module is further configured to automatically collect inference results for metric calculation after the inference task is completed, complete multi-dimensional performance evaluation, and generate an evaluation result. An evaluation result generation module is further configured to generate a structured analysis report based on the evaluation result and return it to the client. The newly added models and datasets will be updated in the model list and dataset list, and the newly added data will be integrated and updated into the instruction-following dataset and centrally managed on the disclosed evaluation platform.

[0074] Figure 2 is a flowchart of a method for evaluating a vision-language model shown in accordance with an exemplary embodiment of the present specification. As Figure 2 shown, the method may be applied to a disclosed evaluation platform and specifically includes steps 201-203:

[0075] Step 201: Receive a target model to be evaluated and at least one target task selected by the user from a task set provided by the disclosed evaluation platform; each task in the task set corresponds to an instruction-following dataset, and any sample in the instruction-following dataset includes an image, an instruction, and an answer.

[0076] Step 202: Obtain a target dataset corresponding to the target task from the instruction-following dataset corresponding to the task set, and the target model performs inference on the target dataset to obtain an inference result.

[0077] Step 203: Determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation metrics of each target task.

[0078] In this embodiment, to meet the user's requirement of evaluating the performance of the target model on at least one target task in the task set provided by the public evaluation platform, this solution provides the user with an instruction-following dataset and evaluation metrics corresponding to the target task. The target model performs inference on the unified instruction-following dataset to obtain an inference result, and based on the inference result and the unified target evaluation metrics, determines the evaluation score of the target model on the target task. It can be seen that this solution provides the user with a systematic and standardized evaluation system, which can ensure a fair evaluation of the advantages and disadvantages of each vision-language model on different tasks.

[0079] In one embodiment, the evaluation metrics can include not only quantitative metrics but also human scoring. For example, in open-ended response tasks such as image captioning, it is difficult to reflect the true performance of the model based on quantitative metrics. Therefore, a human evaluation experiment can be conducted to obtain a more accurate evaluation. First, 50 different types of images are selected from a private high-score dataset, including different scenes such as cities, rural areas, and industrial zones, covering a large number of labels such as roads, grasslands, buildings, ponds, and farmlands. To ensure the accuracy and reliability of the evaluation, multiple volunteers are invited to score according to the designed questionnaire to evaluate the image captioning performance of the model. Each questionnaire includes 50 images, the image captioning results of the anonymized model, and the scoring criteria. Among them, the scoring criteria are designed from three dimensions, namely details, location, and hallucination description. Each dimension uses a four-level scoring system, namely A, B, C, and D. The specific criteria for each level are shown in Table 5. By quantifying the A-D ratings from 4 to 1, the performance of the model can be further quantitatively analyzed.

[0080] Table 5

[0081]

[0082] In one embodiment, the task set can be divided according to the ability dimension, and each ability dimension includes at least one task. The division results can be exemplarily referred to the aforementioned Table 4, which will not be elaborated herein.

[0083] In addition to determining the first evaluation score of the target model on each target task according to the inference result and the target evaluation metrics of each target task, it is also possible to determine the target ability dimension to which each of at least one target task belongs. For any target ability dimension, according to the first evaluation scores of the target tasks under this target ability dimension, determine the second evaluation score of the target model on this target ability dimension. And / or, according to the second evaluation scores on each target ability dimension, determine the third evaluation score of the target model.

[0084] In addition, if there are multiple instruction-following datasets corresponding to the target task, the first evaluation score of the target model on the target task can be calculated using the following formula:

[0085]

[0086] where is the first evaluation score of the target model on the target task, is the number of instruction-following datasets corresponding to the target task, is the evaluation score of the target model on the th instruction-following dataset.

[0087] To balance the difficulty of each target task and the problem of inconsistent scoring scales on different instruction-following datasets for different target tasks, the raw scores of the target model on each instruction-following dataset are uniformly scaled to obtain the evaluation scores of the target model on each instruction-following dataset.

[0088] For example, taking Table 4 as an example, assuming the target tasks selected by the user are the image classification task, the horizontal box object detection task, and the visual localization task, it can be determined that the target ability dimension to which the image classification task and the horizontal box object detection task belong is the global visual perception ability dimension, while the target ability dimension to which the visual localization task belongs is the fine-grained visual perception ability dimension.

[0089] For the global visual perception ability dimension, the second evaluation score thereon can be determined by the first evaluation score of the image classification task and the first evaluation score of the visual localization task; while for the fine-grained visual perception ability dimension, the second evaluation score thereon can be determined by the first evaluation score of the visual localization task.

[0090] For the third evaluation score of the target model, the third evaluation score thereon can be determined by the second evaluation score of the global visual perception ability dimension and the second evaluation score of the fine-grained visual perception ability dimension.

[0091] In this embodiment, by dividing the tasks into different ability dimensions, the performance of the target model can be evaluated at different granularities such as datasets, tasks, ability dimensions, and comprehensive dimensions.

[0092] In one embodiment, when determining the second evaluation score of the target model on the target ability dimension according to the average value of each first evaluation score, the first number of tasks not selected by the user under the target ability dimension can be determined, and the second evaluation score of the target model on the target ability dimension can be determined according to the first number and the average value of each first evaluation score. Wherein, the first number is used to determine the penalty weight.

[0093] For example, following the previous example, for the dimension of global visual perception ability, the tasks not selected by the user are the rotated box object detection task, the object segmentation task, and the visual question answering - existence task. Therefore, the first quantity can be obtained as 3.

[0094] Assume that the evaluation score of the target model on this target ability dimension is determined only based on the average value of each first evaluation score as , the first quantity is used to determine the penalty weight, whose value is not greater than 1 and greater than 0. The first quantity of the tasks not selected by the user and the penalty weight are negatively correlated. Finally, the second evaluation score of the target model on this target ability dimension determined according to the first quantity and the average value of each first evaluation score can be .

[0095] It can be seen that in this embodiment, when calculating the second evaluation score on the target ability dimension, the situation of other tasks not covered on the target ability dimension can be considered, and a penalty weight is added for balance, so that when evaluating the score of the target model on this target ability dimension, not only the performance of the model is considered, but also the breadth of the target model is considered.

[0096] To balance the performance and breadth of the target model on the target ability dimension, in one embodiment, the formula used to determine the second evaluation score of the target model on this target ability dimension according to the first quantity and the average value of each first evaluation score is:

[0097]

[0098] where, is the second evaluation score on this target ability dimension, is the number of target tasks selected under this target ability dimension, is the first evaluation score of the target model on the th target task, is the first weight, is the second weight, is the number of all tasks under this target ability dimension. Optionally, = 0.7, = 0.3.

[0099] In this embodiment, = , = .

[0100] In one embodiment, when determining the third evaluation score of the target model according to the second evaluation scores on each target ability dimension, the second quantity of the ability dimensions not covered in the task set may be determined; the third evaluation score of the target model is determined according to the second quantity and the second evaluation scores on each target ability dimension, where the second quantity is used to determine the penalty weight.

[0101] In this embodiment, assuming that the second quantity is not considered, the evaluation score of the target model can be determined as Y according to the second evaluation scores on each target ability dimension, and the second quantity can be used to determine the penalty weight , whose value is not greater than 1 but greater than 0, the second quantity of the uncovered ability dimensions is negatively correlated with the penalty weight, and finally the third evaluation score of the target model determined according to the second quantity and the second evaluation scores on each target ability dimension can be .

[0102] In this embodiment, when calculating the comprehensive ability score, the model is allowed to stand out in the advantageous dimensions, while adding penalties for the uncovered dimensions, avoiding high scores in a single dimension covering up the short board, and encouraging all-round development.

[0103] In one embodiment, the formula used to determine the third evaluation score of the target model according to the quantity and the second evaluation scores on each target ability dimension is:

[0104]

[0105] Wherein, is the third evaluation score, is the quantity of the target ability dimensions, is the second evaluation score of the target model on the th target ability dimension, is the second quantity of the ability dimensions not covered in the task set, is the third weight, which can be selected to be set to 20%.

[0106] In this embodiment, β = , .

[0107] Corresponding to the embodiment of the foregoing method, this specification also provides an embodiment of a device and a terminal to which the device is applied.

[0108] As Figure 3 shown, Figure 3 FIG. 0 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. At the hardware level, the electronic device 300 includes a processor 302, an internal bus 304, a network interface 306, a memory 308, and a non-volatile memory 310. Of course, it may also include other hardware required for other services. One or more embodiments of this specification can be implemented in software. For example, the processor 302 reads the corresponding computer program from the non-volatile memory 310 into the memory 308 and then runs it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic module, and can also be hardware or a logic device.

[0109] As Figure 4 shown, Figure 4 FIG. 7 is a block diagram of an evaluation device for a vision-language model shown according to an exemplary embodiment of this specification. The device can be applied to the electronic device 300 as Figure 3 shown to implement the technical solutions of this specification. The device includes:

[0110] A task receiving module 402, configured to receive a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; each task in the task set corresponds to an instruction-following data set, and any sample in the instruction-following data set includes an image, an instruction, and an answer.

[0111] An inference module 404, configured to obtain a target data set corresponding to the target task from the instruction-following data set corresponding to the task set, and the target model performs inference on the target data set to obtain an inference result.

[0112] An evaluation module 406, configured to determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task.

[0113] Optionally, the task set is divided according to ability dimensions, and each ability dimension includes at least one task. The evaluation module 406 is further configured to determine the target ability dimension to which each of the at least one target task belongs. For any target ability dimension, according to the first evaluation scores of the target tasks under the target ability dimension, determine a second evaluation score of the target model on the target ability dimension; and / or, according to the second evaluation scores on each target ability dimension, determine a third evaluation score of the target model.

[0114] Optionally, the evaluation module 406 is specifically configured to determine a first quantity of tasks that are not selected by the user under the target capability dimension; determine a second evaluation score of the target model on the target capability dimension according to the first quantity and an average value of each first evaluation score; wherein, the first quantity is used to determine a penalty weight.

[0115] Optionally, the formula used to determine the second evaluation score of the target model on the target capability dimension according to the first quantity and the average value of each first evaluation score is:

[0116]

[0117] Wherein, is the second evaluation score on the target capability dimension, is the quantity of target tasks selected under the target capability dimension, is the first evaluation score of the target model on the j-th target task, is the first weight, is the second weight, is the quantity of all tasks under the target capability dimension.

[0118] Optionally, the evaluation module 406 is specifically configured to determine a second quantity of capability dimensions not covered in the task set; determine a third evaluation score of the target model according to the second quantity and the second evaluation scores on each target capability dimension, wherein, the second quantity is used to determine a penalty weight.

[0119] Optionally, the formula used to determine the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target capability dimension is:

[0120]

[0121] Wherein, is the third evaluation score, is the quantity of target capability dimensions, is the second evaluation score of the target model on the -th target capability dimension, is the second quantity, is the third weight.

[0122] The implementation processes of the functions and roles of each module in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0123] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0124] This specification also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of any one of the foregoing evaluation methods of the visual language model provided by this application are implemented.

[0125] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.< / class> < / class> < / number> < / number> < / class> < / class> < / region> < / description> < / region> < / description> < / description> < / region> < / class> < / class> < / class> < / region> < / class> < / region> < / class> < / class> < / region> < / class> < / region> < / class>

Claims

1. A method for evaluating a visual language model, characterized in that, The method is applied to an open evaluation platform and includes: Receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the open evaluation platform; each task in the task set corresponds to an instruction-following data set, and any sample in the instruction-following data set includes an image, an instruction, and an answer; the task set is divided according to ability dimensions, and each ability dimension includes at least one task; Obtaining a target data set corresponding to the target task from the instruction-following data set corresponding to the task set, and the target model performs inference on the target data set to obtain an inference result; Determining a first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task; Determining a third evaluation score of the target model according to the second evaluation scores on each target ability dimension, including: determining a second quantity of the ability dimensions not covered in the task set; determining the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target ability dimension, where the second quantity is used to determine a penalty weight.

2. The method according to claim 1, wherein The method further includes: Determining the target ability dimension to which each of the at least one target task belongs, and for any target ability dimension, determining a second evaluation score of the target model on the target ability dimension according to the first evaluation scores of each target task under the target ability dimension.

3. The method according to claim 2, wherein The determining the second evaluation score of the target model on the target ability dimension according to the first evaluation scores of each target task under the target ability dimension includes: Determining a first quantity of the tasks not selected by the user under the target ability dimension; Determining the second evaluation score of the target model on the target ability dimension according to the first quantity and the average value of each first evaluation score; where the first quantity is used to determine a penalty weight.

4. The method according to claim 3, characterized in that, The formula used for determining the second evaluation score of the target model on the target ability dimension according to the first quantity and the average value of each first evaluation score is: Among them, is the second evaluation score on the target ability dimension, is the number of target tasks selected under the target ability dimension, is for the target model at the th target task's first evaluation score, is the first weight, is the second weight, is the number of all tasks under the target ability dimension.

5. The method according to claim 1, wherein The formula used for determining the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target ability dimension is: Among them, is the third evaluation score, is the number of target ability dimensions, is the second evaluation score of the target model on the k-th target ability dimension, is the second quantity, is the third weight.

6. An open evaluation platform, characterized in that, Including: An evaluation request receiving module, configured to receive a target model to be evaluated and at least one target task selected by a user from a task set provided by the open evaluation platform; The task set is divided according to ability dimensions, and each ability dimension includes at least one task; An instruction-following data set construction module, configured to construct an instruction-following data set corresponding to each task in the task set, and any sample in the instruction-following data set includes an image, an instruction, and an answer; A model inference module, configured to obtain a target data set corresponding to the target task from the instruction-following data set corresponding to the task set, and the target model performs inference on the target data set to obtain an inference result; A model evaluation module, configured to determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task; The model evaluation module is further configured to determine a third evaluation score of the target model according to the second evaluation scores on each target ability dimension, including: determining a second quantity of the ability dimensions not covered in the task set; and determining the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target ability dimension, where the second quantity is used to determine a penalty weight.

7. The platform according to claim 6, wherein Constructing the instruction-following data set corresponding to each task in the task set includes: Obtaining an image annotation data set, where any sample in the image annotation data set includes an image and an annotation result; Converting the image annotation data set into an instruction-following data set, where for different task types, instructions and answers corresponding to the task types are determined according to the annotation results of the image annotation data set.

8. The platform according to claim 7, characterized in that For different task types, determining instructions and answers corresponding to the task types according to the annotation results of the image annotation data set includes: For any task type, obtaining an instruction template and an answer template corresponding to the task type; If the instruction template includes a first placeholder, determining a first object from the annotation results of the image annotation data set, and using the instruction obtained by replacing the first placeholder in the instruction template with the first object; if the instruction template does not include a first placeholder, using the instruction template as the instruction; Determining a second object from the annotation results of the image annotation data set, and using the answer obtained by replacing the second placeholder in the answer template with the second object; where the answer template includes a second placeholder for replacement with the second object.

9. The platform according to claim 6, characterized in that, The platform further includes: An evaluation score storage module, configured to store the evaluation scores of the models for which the evaluation is completed; An evaluation result generation module, configured to generate an evaluation result according to the evaluation scores of other models and the target model on the same task.

Citation Information

Patent Citations

  • Multi-party model and multi-party financial model credibility evaluation method, device and equipment

    CN115292144A

  • Multi-dimensional evaluation method for visual basic model

    CN118887500A