Evaluation method of visual language model and public evaluation platform
Through the public evaluation platform and unified evaluation methods, the problem of inconsistent evaluation framework of visual language model is solved, and a fair and systematic evaluation effect is achieved.
Patent Information
- Application Number
- CN202510481339.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In the prior art, the evaluation framework of visual language models is not uniform, making it difficult to fairly evaluate the advantages and disadvantages of each visual language model in different visual-language tasks.
It provides an evaluation method and public evaluation platform for visual language models. By receiving target models and tasks selected by users, it constructs a unified instruction following data sets and evaluation indicators, and conducts systematic and standardized evaluation.
This achieves a fair assessment of the advantages and disadvantages of each visual language model in different tasks, simplifies the evaluation process, and ensures consistency and comparability of the evaluation results.
Smart Images

Figure CN119988915A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to an evaluation method and a public evaluation platform for a visual language model. Background Art
[0002] Vision-Language Models (VLM) are multimodal artificial intelligence models that can process visual and language information simultaneously. They are mainly used in scenarios such as image classification, visual question answering, image description, and spatial change detection.
[0003] Nowadays, various visual language model frameworks emerge one after another, and the corresponding demand arises to evaluate the processing performance of various visual language models on different visual-language tasks. When evaluating the visual language models they develop, technicians usually collect public data sets or build their own data sets and set evaluation indicators on their own.
[0004] When different technicians evaluate their own visual language models, even if they evaluate the same visual-language task, they may use different datasets and evaluation indicators. Due to the lack of unified evaluation frameworks, such as the lack of unified standards for the datasets used and inconsistent evaluation indicators, it is difficult to fairly evaluate the pros and cons of various visual-language models on different visual-language tasks. Summary of the invention
[0005] To overcome the problems existing in the related art, this specification provides an evaluation method and a public evaluation platform for a visual language model.
[0006] According to a first aspect of an embodiment of this specification, a method for evaluating a visual language model is provided. The method is applied to a public evaluation platform and includes: Receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; each task in the task set corresponds to an instruction following data set, and any sample in the instruction following data set includes an image, an instruction, and an answer; Acquire a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and perform reasoning on the target data set by the target model to obtain a reasoning result; According to the inference result and the target evaluation index of each target task, a first evaluation score of the target model on each target task is determined.
[0007] According to a first aspect of an embodiment of this specification, a public evaluation platform is provided, including: An evaluation request receiving module, used for receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; An instruction following data set construction module, used to construct an instruction following data set corresponding to each task in the task set, wherein any sample in the instruction following data set includes an image, an instruction and an answer; A model reasoning module, used for acquiring a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and the target model performs reasoning on the target data set to obtain a reasoning result; The model evaluation module is used to determine the first evaluation score of the target model on each target task based on the inference result and the target evaluation index of each target task.
[0008] The technical solutions provided by the embodiments of this specification may have the following beneficial effects: In the embodiments of this specification, in response to the user's need to evaluate the performance of the target model on at least one target task in the task set provided by the public evaluation platform, this solution provides the user with an instruction following data set and evaluation indicators corresponding to the target task, and the target model performs reasoning on the unified instruction following data set to obtain the reasoning result, and determines the evaluation score of the target model on the target task based on the reasoning result and the unified target evaluation indicator. It can be seen that this solution provides users with a systematic and standardized evaluation system that can ensure fair evaluation of the advantages and disadvantages of each visual language model on different tasks.
[0009] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0011] Figure 1 This is an application scenario diagram of a public evaluation platform shown in this specification according to an exemplary embodiment.
[0012] Figure 2 This is a flowchart of a visual language model evaluation method shown in this specification according to an exemplary embodiment.
[0013] Figure 3 It is a structural schematic diagram of an electronic device according to an exemplary embodiment of the present specification.
[0014] Figure 4It is a block diagram of a visual language model evaluation device according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION
[0015] Vision-Language Models (VLM) are multimodal artificial intelligence models that can process visual and language information simultaneously. They are mainly used in scenarios such as image classification, visual question answering, image description, and spatial change detection.
[0016] Nowadays, various visual language model frameworks emerge one after another, and the corresponding demand arises to evaluate the processing performance of various visual language models on different visual-language tasks. When evaluating the visual language models they develop, technicians usually collect public data sets or build their own data sets and set evaluation indicators on their own.
[0017] When different technicians evaluate their own visual language models, even if they evaluate the same visual-language task, they may use different datasets and evaluation indicators. Due to the lack of unified evaluation frameworks, such as the lack of unified standards for the datasets used and inconsistent evaluation indicators, it is difficult to fairly evaluate the pros and cons of various visual-language models on different visual-language tasks.
[0018] In response to the above technical problems, this specification provides a visual language model evaluation method and a public evaluation platform, which provides a systematic and standardized evaluation system to ensure fair evaluation of the advantages and disadvantages of various visual language models on different tasks.
[0019] Figure 1 This is an application scenario diagram of a public evaluation platform according to an exemplary embodiment of this specification. Figure 1 As shown, the server 10 communicates with each client 12-14 through the network 11, and the public evaluation platform can be deployed on the server 10. The server 10 can provide a front-end access page to the client to realize the interaction between the client and the public evaluation platform. For example, the front-end access page can provide at least a model upload module and a task selection module. Among them, the model upload module can be used to upload the model to be evaluated, and the task selection module can be used to display a variety of visual-language tasks provided by the public evaluation platform. The user can select at least one target task from the task selection module. The front-end access page can submit the model input by the user and the selected target task to the public evaluation platform, and the public evaluation platform can return the evaluation results to at least the front-end access page after completing the evaluation.
[0020] The application scenario of the public evaluation platform can be the evaluation of visual language models. Corresponding instruction following datasets and evaluation indicators can be constructed on the public evaluation platform for different visual-language tasks. Among them, each visual-language task can correspond to multiple instruction following datasets and multiple evaluation indicators to ensure that the performance of the model in a wide range of scenarios is evaluated and the model performance is measured more comprehensively.
[0021] Of course, for the application scenarios of visual language models, multiple capability dimensions can also be designed for the performance of visual language models, and multiple visual-language tasks can be divided for each capability dimension to comprehensively evaluate the performance of the model from multiple aspects such as tasks, capability dimensions, and comprehensive dimensions, so as to identify the shortcomings or advantages of the model to be evaluated in specific tasks, and promote developers to optimize the model in a targeted manner. In addition, the public evaluation platform provides a unified capability dimension and task classification to provide a benchmark for performance comparison between different models, which promotes the transparency and fairness of evaluation in the community studying visual language models.
[0022] After receiving the target model to be evaluated and at least one target task selected by the user from the task set provided by the public evaluation platform, the public evaluation platform can evaluate the performance of the target model on each target task and return the evaluation result to the user.
[0023] Compared with technicians independently completing repetitive processes such as instruction following datasets and evaluation indicators, the public evaluation platform can integrate instruction following datasets and evaluation indicators of different visual-language tasks on the public evaluation platform for technicians to directly call, simplifying the evaluation process, and the instruction following datasets and evaluation indicators used by different technicians are consistent, thereby ensuring the consistency and comparability of the evaluation results. In addition, for technicians who are new to the field of visual language models, the public evaluation platform can provide ready-made datasets and evaluation indicators, thereby reducing the initial cost of research.
[0024] Next, we will introduce the functional modules in the public evaluation platform in detail, including: The evaluation request receiving module is used to receive the target model to be evaluated and at least one target task selected by the user from the task set provided by the public evaluation platform.
[0025] The instruction following data set construction module is used to construct an instruction following data set corresponding to each task in the task set, and any sample in the instruction following data set includes an image, an instruction and an answer.
[0026] The model reasoning module is used to obtain a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and the target model performs reasoning on the target data set to obtain a reasoning result.
[0027] The model evaluation module is used to determine the first evaluation score of the target model on each target task based on the inference result and the target evaluation index of each target task.
[0028] In one embodiment, the evaluation request receiving module can also receive other parameters. The other parameters may include the target instruction following data set and target evaluation index selected by the user for the target task after selecting the target task. Then, the first evaluation score of the target model determined by the model evaluation module on each target task can be related to the data set source and the target evaluation index. For example, the first evaluation score can be associated with the data source and the evaluation index. Of course, even if the user does not make any selection, the data source and the evaluation index can also be associated when determining the first evaluation score.
[0029] Although the visual language model has achieved great success in the field of natural images, it still lacks corresponding instruction following datasets for images in certain technical fields. For example, for remote sensing images, most of the currently available datasets in the remote sensing field are image annotation datasets, in which any sample includes an image and an annotation result. However, since the image annotation dataset in the remote sensing field cannot be directly applied to the evaluation of the visual language model, the application of the visual language model in the remote sensing field is limited.
[0030] Therefore, this specification provides a method for converting an image annotation dataset into an instruction following dataset: In one embodiment, an image annotation dataset is obtained, wherein any sample in the image annotation dataset includes an image and an annotation result. The image annotation dataset is converted into an instruction following dataset, wherein for different task types, instructions and answers corresponding to the task type are determined according to the annotation results of the image annotation dataset. In this embodiment, by converting image annotation datasets in different technical fields into corresponding instruction following datasets for application in the study of visual language models, it is not necessary to construct an instruction following dataset for the technical field from scratch, thereby saving the cost of constructing an instruction following dataset.
[0031] For example, as shown in Table 1, the correspondence between image annotation datasets and instruction following datasets on different task types is shown: Table 1
[0032] The key to converting an image annotation dataset into an instruction following dataset is to convert the annotation results in the image annotation dataset into instructions and answers. For example, for a dataset for an image classification task, it includes image and category annotations. For the category annotation of any image in the dataset, the category annotation can be converted into instructions and answers, and the instructions and answers can be replaced with the original image annotation dataset to obtain an instruction following dataset. For another example, for an image annotation dataset for a segmentation task, the annotation results include pixel annotations in the target area and the corresponding categories. For the task type that needs to be converted, the annotation results can be converted into instructions and answers corresponding to the task type. Of course, the conversion method for datasets of other task types is similar and will not be repeated here.
[0033] In one embodiment, the image annotation dataset applied to a certain original task is not limited to being converted into an instruction following dataset of the same task, but can be extended to other task types, and the instruction following dataset of other task types can be generated accordingly. In other words, the instruction following dataset of at least one task type can be generated based on the image annotation dataset of a certain task type, and the at least one task type includes the certain task type.
[0034] For example, as shown in Table 2, the correspondence between the original tasks and the extended tasks of the image annotation dataset is shown: Table 2
[0035] For example, see Table 2. For an image annotation dataset for image description, the image annotation dataset can be converted into an instruction following dataset for the image short description task and an instruction following dataset for the image detailed description task. Similarly, for an image annotation dataset for a segmentation task, in addition to converting the image annotation dataset into an instruction following dataset for the segmentation task, it can also be converted into an instruction following dataset for the target computing task and an instruction following dataset for the polygon region classification task.
[0036] Therefore, when converting an image annotation dataset into an instruction following dataset, the annotation results of the image dataset are essentially converted into instructions and answers, and the specific contents of the instructions and answers depend on the target task type to be converted. For example, for the same annotation result in the image annotation dataset of the segmentation task, the instructions and answers obtained when converted to different task types are different. For example, the instructions and answers corresponding to the segmentation task can be obtained based on the annotation result, and the instructions and answers corresponding to the target counting task can also be obtained.
[0037] In one embodiment, for different task types, when determining the instructions and answers corresponding to the task type based on the annotation results of the image annotation data set, for any task type, obtain the instruction template and answer template corresponding to the task type. If the instruction template includes a first placeholder, a first object is determined from the annotation results of the image annotation data set, and the first placeholder in the instruction template is replaced with the first object as an instruction; if the instruction template does not include the first placeholder, the instruction template is used as an instruction. A second object is determined from the annotation results of the image annotation data set, and the second placeholder in the answer template is replaced with the second object as an answer; wherein the answer template includes a second placeholder for replacement with the second object.
[0038] For example, as shown in Table 3, instruction-answer templates for different task types are shown: Table 3
[0039] For example, for the horizontal frame object detection task, the instruction template corresponding to the horizontal frame detection task is obtained as: "Detect all <class>in the image"; the answer template is: " <region>". The first placeholder in the instruction template is <class>, the second placeholder in the answer template is <region>. Determine the first object and the second object respectively from the annotation results in the image annotation dataset for the horizontal frame target detection task. Wherein, the first object is " <class>", specifically, the target category of the marked horizontal box; the second object is " <box> <x1> <y1> <x2> <y2>< / y2> < / x2> < / y1> < / x1> < / box> ", (X1, Y1), (X2, Y2) are the coordinates of the upper left corner and the lower right corner of the horizontal box respectively. The first placeholder in the instruction template is replaced by the first object as the instruction, and the second placeholder in the answer template is replaced by the second object as the answer. It should be noted that the above is explained for any horizontal box of the annotation result. For other horizontal boxes in the annotation result, the conversion method can be carried out in the same way as above, which will not be repeated here.
[0040] For example, for the rotating frame target detection task, the instruction template corresponding to the rotating frame target detection task can be obtained as follows: "Detect all <class>in the image. Use oriented bounding boxes"; the answer template is: " <region>". The first placeholder in the instruction template is <class>, the second placeholder in the answer template is <region>. Determine the first object and the second object respectively from the annotation results in the image annotation dataset for the rotating box object detection task. Wherein, the first object is " <class>", is the target category of the marked rotation box; the second object is " <quad> <x1> <y1> <x2> <y2> <x3> <y3> <x4> <y4>< / y4> < / x4> < / y3> < / x3> < / y2> < / x2> < / y1> < / x1> < / quad> ", (X1, Y1), (X2, Y2), (X3, Y3) and (X4, Y4) are the coordinates of the four vertices of the rotation box respectively. The first placeholder in the instruction template is replaced by the first object as the instruction, and the second placeholder in the answer template is replaced by the second object as the answer.
[0041] For example, for an image classification task, the instruction template corresponding to the image classification task is: "Classify the image. Use one or a few words"; the answer template is: " <class>". Wherein, if the instruction template does not include the first placeholder, the instruction template is directly used as an instruction; the second object is determined to be " <class>" is the category of the image, then the second placeholder in the answer template is replaced with the second object as the answer.
[0042] Similarly, for image annotation datasets for visual localization tasks, the description corresponding to the bounding box and the coordinate information of the bounding box can be determined from the annotation results in the image annotation dataset. For image annotation datasets for segmentation tasks, the segmentation mask image in the annotation results can be converted into the text representation required by the template, and the target category corresponding to the mask can be determined from it.
[0043] For example, for some specific task image annotation datasets, when constructing the corresponding instruction following dataset, an instruction following dataset different from the task type can be constructed. For details, see Table 2.
[0044] For extended task types, you can obtain the instruction template and answer template corresponding to the extended task type. For example, for the region description task extended from the visual localization task, you can obtain the instruction template and answer template corresponding to the region description task. The instruction template is "Describe the <region>in this image", the answer template is " <description>The annotation results of the image annotation dataset for the visual localization task include a description of the target area. <description>and the target's location coordinates <region>, such as the target description <description>is "a redvehicle", the target's location coordinates <region>for" <box> <125><369><256><489>< / box> ". Then the first object determined from the annotation results of the image annotation dataset is " <box> <125><369><256><489>< / box> ", the first placeholder in the instruction template can be replaced with the first object as the instruction; the second object is determined to be "a red vehicle" from the annotation results of the image annotation dataset, and the second placeholder in the answer template can be replaced with the second object as the answer.
[0045] Similar to the above-mentioned extended task, for the target counting task extended from the horizontal frame target detection task, obtain the instruction template and answer template corresponding to the target counting task. Among them, the instruction template is "Count the number of <class>", the first placeholder in the instruction template is <class>; The answer template is " <number>", the second placeholder in the answer template is " <number>”. A first object is determined from the annotation results of the image annotation dataset corresponding to the horizontal frame target detection task. The first object is a type of <class>A horizontal box, such as this <class>It can be a car, then the car is the first object, the first placeholder in the instruction template is replaced with "car", and the second placeholder in the answer template is replaced with the number of cars as the answer, and the number of cars can be obtained from the statistics of the marking results.
[0046] Of course, for some extended tasks, instruction templates and answer templates are not needed, for example, existence judgment tasks and quantity comparison tasks. To judge whether a certain target exists in the image, such as "Is a forest present in the image?", the answer is "yes" or "no". Quantity comparison is to judge the quantitative relationship between two types of targets in the image, such as "Are there more roads than commercial buildings?", the answer is "yes" or "no". The answer can be determined based on the annotation results in the image annotation dataset.
[0047] For the image annotation dataset of the image description task, it can be converted into an instruction following dataset for the image detailed description task and / or an instruction following dataset for the image short description task. The image description text is obtained from the annotation results in the image annotation dataset, and the image description text is deduplicated. If the number of sentences in the deduplicated image description text is not less than a first threshold or the number of words in the image description text is not less than a second threshold, the image description text is used as the answer to the instruction following dataset for the detailed description task, and a sentence is randomly selected from the unused sentences as the answer to the instruction following dataset for the image short description task. If the number of sentences in the deduplicated image description text is less than the first threshold or the number of words in the image description text is less than the second threshold, the deduplicated image description text is used as the answer to the instruction following dataset for the image short description task.
[0048] Exemplarily, the image description task is expanded into the image short description task and the image detailed description task. Since the description annotation quality of the open source data set is uneven and there are repeated descriptions, deduplication is performed first. For a description of an image, the similarity between sentences is calculated by the word intersection ratio of the bag-of-words model. You can also choose to use Jaccard similarity, TF-IDF cosine similarity or sentence vector to measure the similarity of sentences. Sentences that are judged to be dissimilar are merged into one sentence to determine whether it is a detailed description or a short description. Judgment rule: A description sentence with less than or equal to 4 sentences or a sentence length of less than or equal to 30 words is defined as a short description, and a description sentence with more than 4 sentences or a sentence length of more than 30 words is defined as a detailed description. If the merged description is a detailed description and there are unused similar sentences, one sentence is selected from the remaining similar sentences as a short description; if it is a short sentence, all similar sentences are discarded.
[0049] For different task types, it can reflect the capabilities of the visual language model in different capability dimensions. This specification provides an exemplary method for dividing the capability dimensions for evaluating the visual language model. Of course, other capability dimension division methods can also be used, and this specification does not impose any restrictions on this.
[0050] For example, as shown in Table 4, Table 4 shows the tasks under different ability dimensions: Table 4
[0051] In one embodiment, when the model evaluation module determines the first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task, in addition to determining the first evaluation score of the target model on each target task, the target capability dimension to which each target task belongs can also be determined. For any target capability dimension, the second evaluation score of the target model on the target capability dimension is determined according to the first evaluation score of each target task under the target capability dimension. And / or, the third evaluation score of the target model is determined according to the second evaluation score on each target capability dimension.
[0052] In this embodiment, when evaluating the model capabilities, we not only focus on the performance of a single task, but also examine the comprehensive performance of the model in multiple capability dimensions, which helps to identify the overall strengths and potential weaknesses of the model.
[0053] In one embodiment, the evaluation platform disclosed herein may also include an evaluation score storage module for storing the evaluation scores of the evaluated models; and an evaluation result generation module for generating evaluation results based on the evaluation scores of other models and the target model on the same task.
[0054] The evaluation result may be in the form of a chart, showing the evaluation results of the target model and other models in different dimensions of different granularity, such as different instruction following data sets, evaluation indicators, tasks, and capability dimensions.
[0055] In one embodiment, the public evaluation platform may also include a test trigger module for receiving a data set / model update request submitted by a user to the public evaluation platform. The model reasoning module is also used to load the data sets of all tasks that the model can handle in the instruction-following data set when a model update instruction is received, and start the reasoning task; when a data set update is received, call all models that are capable of the data set task and start the reasoning task. The instruction-following data set can be centrally managed in the server through a distributed storage system, and the model is dynamically deployed on any node in the cluster. All model services can follow the preset communication interface template, which is used for communication between servers and supports cross-server data interaction and result return (sending and receiving data and results). The model evaluation module is also used to automatically collect the reasoning results for indicator calculation after the reasoning task is completed, complete multi-dimensional performance evaluation, and generate evaluation results. The evaluation result generation module is also used to generate a structured analysis report based on the evaluation results and return it to the client. The newly added model and data set will be updated in the model list and data set list, and the newly added data will be integrated and updated into the instruction-following data set, and centrally managed on the public evaluation platform.
[0056] Figure 2 FIG. 1 is a flowchart of a method for evaluating a visual language model according to an exemplary embodiment of the present specification. Figure 2 As shown, the method can be applied to a public evaluation platform, and specifically includes steps 201-203: Step 201: Receive a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; each task in the task set corresponds to an instruction following data set, and any sample in the instruction following data set includes an image, an instruction, and an answer.
[0057] Step 202: Obtain a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and perform reasoning on the target data set by the target model to obtain a reasoning result.
[0058] Step 203: Determine a first evaluation score of the target model on each target task based on the inference result and the target evaluation index of each target task.
[0059] In this embodiment, in response to the user's need to evaluate the performance of the target model on at least one target task in the task set provided by the public evaluation platform, this solution provides the user with an instruction following data set and evaluation indicators corresponding to the target task. The target model performs reasoning on the unified instruction following data set to obtain the reasoning result, and determines the evaluation score of the target model on the target task based on the reasoning result and the unified target evaluation indicator. It can be seen that this solution provides users with a systematic and standardized evaluation system that can ensure fair evaluation of the advantages and disadvantages of each visual language model on different tasks.
[0060] In one embodiment, the evaluation index may include not only quantitative indexes but also manual scoring. For example, in open-ended answer tasks such as image description, it is difficult for evaluation based on quantitative indexes to reflect the true performance of the model. Therefore, a more accurate evaluation can be obtained through human evaluation experiments. First, 50 images of different types are selected from a private high-scoring dataset, including different scenes such as cities, rural areas, and industrial areas, covering a large number of labels such as roads, grasslands, buildings, ponds, and farmlands. In order to ensure the accuracy and reliability of the evaluation, multiple volunteers are invited to score according to the designed questionnaire to evaluate the image description performance of the model. Each questionnaire includes 50 images, the image description results of the anonymous model, and the scoring criteria. Among them, the scoring criteria are designed from three dimensions, namely, details, location, and illusion description. Each dimension adopts a four-level scoring system, namely A, B, C, and D. The specific criteria for each level are shown in Table 5. By quantifying the AD rating into 4 to 1 points, the performance of the model can be further quantitatively analyzed.
[0061] Table 5
[0062] In one embodiment, the task set may be divided according to the capability dimension, and each capability dimension includes at least one task. The division result may be exemplified in Table 4 above, and will not be described in detail in this specification.
[0063] In addition to determining the first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task, the target capability dimension to which at least one target task belongs can also be determined. For any target capability dimension, the second evaluation score of the target model on the target capability dimension is determined according to the first evaluation score of each target task under the target capability dimension. And / or, the third evaluation score of the target model is determined according to the second evaluation score on each target capability dimension.
[0064] In addition, if there are multiple instruction following data sets corresponding to the target task, the first evaluation score of the target model on the target task can be calculated using the following formula:
[0065] in, is the first evaluation score of the target model on the target task, is the number of instruction-following datasets corresponding to the target task, The target model is The instructions are followed by evaluation scores on the dataset.
[0066] In order to balance the difficulty of each target task and the problem of inconsistent score scales on different instruction following datasets for different target tasks, the original scores of the target model on each instruction following dataset are uniformly scaled to obtain the evaluation scores of the target model on each instruction following dataset.
[0067] For example, taking Table 4 as an example, assuming that the target tasks selected by the user are image classification task, horizontal box target detection task and visual positioning task, it can be determined that the target capability dimension of the image classification task and the horizontal box target detection task is the global visual perception capability dimension, while the target capability dimension of the visual positioning task is the fine-grained visual perception capability dimension.
[0068] For the global visual perception ability dimension, the second evaluation score thereon can be determined by the first evaluation score of the image classification task and the first evaluation score of the visual positioning task; and for the fine-grained visual perception ability dimension, the second evaluation score thereon can be determined by the first evaluation score of the visual positioning task.
[0069] For the third evaluation score of the target model, the third evaluation score thereon may be determined by the second evaluation score of the global visual perception capability dimension and the second evaluation score of the fine-grained visual perception capability dimension.
[0070] In this embodiment, by dividing the tasks into different capability dimensions, the performance of the target model can be evaluated from different granularities such as data set, task, capability dimension, and comprehensive dimension.
[0071] In one embodiment, when determining the second evaluation score of the target model on the target capability dimension according to the average value of each first evaluation score, the first number of tasks not selected by the user under the target capability dimension can be determined, and the second evaluation score of the target model on the target capability dimension can be determined according to the first number and the average value of each first evaluation score. The first number is used to determine the penalty weight.
[0072] For example, following the previous example, for the global visual perception ability dimension, the tasks not selected by the user are the rotating box target detection task, the target segmentation task, and the visual question answering-does it exist task, so the first number is 3.
[0073] Assume that the evaluation score of the target model in the target capability dimension is determined based only on the average of the first evaluation scores: The first number is used to determine the penalty weight, whose value is not greater than 1 and greater than 0. The first number of tasks not selected by the user is related to the penalty weight is negatively correlated, and finally the second evaluation score of the target model on the target capability dimension is determined according to the average of the first quantity and each first evaluation score, which can be .
[0074] It can be seen that in this embodiment, when calculating the second evaluation score on the target capability dimension, the situation where other tasks on the target capability dimension are not covered can be considered, and penalty weights can be added for balance, so that when evaluating the score of the target model on the target capability dimension, not only the performance of the model is taken into account, but also the breadth of the target model.
[0075] In order to balance the performance and breadth of the target model in the target capability dimension, in one embodiment, the formula used to determine the second evaluation score of the target model in the target capability dimension according to the first quantity and the average value of each first evaluation score is:
[0076] in, is the second assessment score on the target capability dimension, is the number of target tasks selected under the target capability dimension, For the target model in The first evaluation score on the target task, is the first weight, is the second weight, is the number of all tasks under the target capability dimension. Optionally, =0.7, =0.3.
[0077] In this embodiment, , .
[0078] In one embodiment, when determining the third evaluation score of the target model based on the second evaluation scores on each target capability dimension, the second number of uncovered capability dimensions in the task set can be determined; the third evaluation score of the target model is determined based on the second number and the second evaluation scores on each target capability dimension, wherein the second number is used to determine the penalty weight.
[0079] In this embodiment, assuming that the second number is not considered, the evaluation score of the target model can be determined as Y according to the second evaluation scores on each target capability dimension, and the second number can be used to determine the penalty weight , whose value is not greater than 1 but greater than 0, the second number of uncovered capability dimensions is positively correlated with the penalty weight, and finally the third evaluation score of the target model is determined according to the second number and the second evaluation scores on each target capability dimension, which can be .
[0080] In this embodiment, when calculating the comprehensive ability score, the model is allowed to highlight the advantageous dimensions, while adding a penalty for uncovered dimensions to avoid a high score in a single dimension covering up weaknesses and to encourage all-round development.
[0081] In one embodiment, the formula used to determine the third evaluation score of the target model according to the quantity and the second evaluation scores in each target capability dimension is:
[0082] in, is the third assessment score, is the number of target capability dimensions, For the target model in The second assessment score on the target capability dimension, is the second number of capability dimensions not covered in the task set, For the third weight, you can choose to set it to 20%.
[0083] In this embodiment, , .
[0084] Corresponding to the embodiments of the aforementioned method, this specification also provides embodiments of a device and a terminal to which it is applied.
[0085] like Figure 3 As shown, Figure 3 It is a structural diagram of an electronic device shown in this specification according to an exemplary embodiment. At the hardware level, the electronic device 300 includes a processor 302, an internal bus 304, a network interface 306, a memory 308 and a non-volatile memory 310, and of course may also include hardware required for other services. One or more embodiments of this specification can be implemented based on software, such as the processor 302 reading the corresponding computer program from the non-volatile memory 310 into the memory 308 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic module, but can also be hardware or logic devices.
[0086] like Figure 4 As shown, Figure 4 This is a block diagram of a visual language model evaluation device according to an exemplary embodiment of the present specification. The device can be applied to Figure 3 In the electronic device 300 shown, the technical solution of this specification is implemented. The device includes: The task receiving module 402 is used to receive the target model to be evaluated and at least one target task selected by the user from the task set provided by the public evaluation platform; each task in the task set corresponds to an instruction following data set, and any sample in the instruction following data set includes an image, an instruction, and an answer.
[0087] The reasoning module 404 is used to obtain a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and the target model performs reasoning on the target data set to obtain a reasoning result.
[0088] The evaluation module 406 is used to determine a first evaluation score of the target model on each target task according to the inference result and the target evaluation index of each target task.
[0089] Optionally, the task set is divided according to capability dimensions, each capability dimension includes at least one task, and the evaluation module 406 is also used to determine the target capability dimension to which the at least one target task belongs. For any target capability dimension, the second evaluation score of the target model on the target capability dimension is determined based on the first evaluation scores of each target task under the target capability dimension; and / or the third evaluation score of the target model is determined based on the second evaluation scores on each target capability dimension.
[0090] Optionally, the evaluation module 406 is specifically used to determine a first number of tasks that have not been selected by the user under the target capability dimension; determine a second evaluation score of the target model on the target capability dimension based on the first number and an average of each first evaluation score; wherein the first number is used to determine a penalty weight.
[0091] Optionally, the formula used to determine the second evaluation score of the target model in the target capability dimension according to the first quantity and the average value of each first evaluation score is:
[0092] in, is the second assessment score on the target capability dimension, is the number of target tasks selected under the target capability dimension, is the first evaluation score of the target model on the jth target task, is the first weight, is the second weight, It is the number of all tasks under the target capability dimension.
[0093] Optionally, the evaluation module 406 is specifically used to determine a second number of uncovered capability dimensions in the task set; and determine a third evaluation score of the target model based on the second number and the second evaluation scores on each target capability dimension, wherein the second number is used to determine a penalty weight.
[0094] Optionally, the formula used to determine the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target capability dimension is:
[0095] in, is the third assessment score, is the number of target capability dimensions, For the target model in The second assessment score on the target capability dimension, is the second quantity, The third weight.
[0096] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0097] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. Ordinary technicians in this field can understand and implement it without paying creative work.
[0098] The present specification also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of any of the aforementioned visual language model evaluation methods provided in the present application are implemented.
[0099] Specifically, computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks.< / class> < / class> < / number> < / number> < / class> < / class> < / region> < / description> < / region> < / description> < / description> < / region> < / class> < / class> < / class> < / region> < / class> < / region> < / class> < / class> < / region> < / class> < / region> < / class>
Claims
1. A method for evaluating a visual language model, characterized in that: The method is applied to a public evaluation platform, including: Receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; each task in the task set corresponds to an instruction following data set, and any sample in the instruction following data set includes an image, an instruction, and an answer; Acquire a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and perform reasoning on the target data set by the target model to obtain a reasoning result; According to the inference result and the target evaluation index of each target task, a first evaluation score of the target model on each target task is determined.
2. The method according to claim 1, characterized in that The task set is divided according to capability dimensions, each capability dimension includes at least one task, and the method further includes: Determine the target capability dimension to which each of the at least one target task belongs, and for any target capability dimension, determine a second evaluation score of the target model on the target capability dimension according to the first evaluation scores of each target task under the target capability dimension; and / or, A third evaluation score of the target model is determined based on the second evaluation scores on each target capability dimension.
3. The method according to claim 2, characterized in that Determining the second evaluation score of the target model on the target capability dimension according to the first evaluation score of each target task under the target capability dimension includes: Determining a first number of tasks under the target capability dimension that are not selected by the user; Determine a second evaluation score of the target model on the target capability dimension according to the first number and an average value of each first evaluation score; wherein the first number is used to determine a penalty weight.
4. The method according to claim 3, characterized in that The formula used to determine the second evaluation score of the target model in the target capability dimension according to the first quantity and the average value of each first evaluation score is: in, is the second assessment score on the target capability dimension, is the number of target tasks selected under the target capability dimension, For the target model in The first evaluation score on the target task, is the first weight, is the second weight, It is the number of all tasks under the target capability dimension.
5. The method according to claim 2, characterized in that: Determining the third evaluation score of the target model according to the second evaluation scores on each target capability dimension includes: determining a second number of uncovered capability dimensions in the task set; A third evaluation score of the target model is determined based on the second number and the second evaluation scores on the respective target capability dimensions, wherein the second number is used to determine a penalty weight.
6. The method according to claim 5, characterized in that The formula used to determine the third evaluation score of the target model according to the second quantity and the second evaluation scores on each target capability dimension is: in, is the third assessment score, is the number of target capability dimensions, is the second evaluation score of the target model on the kth target capability dimension, is the second quantity, The third weight.
7. A public evaluation platform, characterized in that: include: An evaluation request receiving module, used for receiving a target model to be evaluated and at least one target task selected by a user from a task set provided by the public evaluation platform; An instruction following data set construction module, used to construct an instruction following data set corresponding to each task in the task set, wherein any sample in the instruction following data set includes an image, an instruction and an answer; A model reasoning module, used for acquiring a target data set corresponding to the target task from the instruction following data set corresponding to the task set, and the target model performs reasoning on the target data set to obtain a reasoning result; The model evaluation module is used to determine the first evaluation score of the target model on each target task based on the inference result and the target evaluation index of each target task.
8. The platform according to claim 7, characterized in that The step of constructing an instruction follow-up data set corresponding to each task in the task set includes: Acquire an image annotation dataset, wherein any sample in the image annotation dataset includes an image and an annotation result; The image annotation dataset is converted into an instruction following dataset, wherein, for different task types, instructions and answers corresponding to the task type are determined according to the annotation results of the image annotation dataset.
9. The platform according to claim 8, characterized in that The step of determining, for different task types, instructions and answers corresponding to the task types according to the annotation results of the image annotation dataset includes: For any task type, obtain the instruction template and answer template corresponding to the task type; If the instruction template includes a first placeholder, a first object is determined from the annotation result of the image annotation data set, and the first placeholder in the instruction template is replaced with the first object as an instruction; if the instruction template does not include the first placeholder, the instruction template is used as an instruction; A second object is determined from the annotation results of the image annotation data set, and a second placeholder in the answer template is replaced with the second object as an answer; wherein the answer template includes a second placeholder for being replaced with the second object.
10. The platform according to claim 7, characterized in that The platform also includes: An evaluation score storage module is used to store the evaluation scores of the models that have been evaluated; The evaluation result generation module is used to generate evaluation results based on the evaluation scores of other models and the target model on the same task.
Citation Information
Patent Citations
Multidisciplinary design optimization method for improving target cascading and Kriging model combination
CN113408099A
Multi-party model and multi-party financial model credibility evaluation method, device and equipment
CN115292144A
Performance evaluation method and device of visual language model in positioning task
CN118736355A
Multi-dimensional evaluation method for visual basic model
CN118887500A
Effect evaluation method for improving saline-alkali soil by desulfurized gypsum
CN118983025A