Evaluation methods, devices and electronic equipment

By acquiring the target image and question, inputting them into MLLM to obtain the reasoning answer, and based on the target question and reference information, this method solves the problem of difficulty in evaluating the conversational and visual capabilities of MLLM in the prior art, and achieves more accurate and efficient evaluation results.

CN119807040BActive Publication Date: 2025-10-31BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411845379.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-10-31
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing assessment methods for Multimodal Large Language Models (MLLM) mainly rely on multiple-choice and true/false questions, which cannot effectively assess conversational and visual abilities, resulting in inaccurate assessment results.

Method used

By acquiring the target image and the question, inputting them into MLLM to obtain the reasoning answer, and based on the target question, assessment reference information, and reasoning answer, the MLLM assessment results are obtained, including assessments of basic visual perception ability, fine-grained image understanding ability, and complex visual reasoning ability.

Benefits of technology

It enables a comprehensive evaluation of MLLM's recognition, perception, reasoning, and conversational abilities, improving the accuracy and efficiency of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807040B_ABST
    Figure CN119807040B_ABST
Patent Text Reader

Abstract

This disclosure provides an evaluation method, apparatus, and electronic device, relating to the field of computer technology, particularly to application areas such as artificial intelligence, large language models, multimodal large language models, generative search, search engines, information retrieval, and question-answering systems. The specific implementation scheme is as follows: acquiring target data and evaluation reference information related to the target data; wherein, the target data includes a target image and a target question; inputting the target image and target question into a multimodal large language model to obtain a reasoned answer to the target question based on the target image using the multimodal large language model; and obtaining an evaluation result for the multimodal large language model based on the target question, evaluation reference information, and reasoned answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to application areas such as artificial intelligence, large language models, multimodal large language models, generative search, search engines, information retrieval, and question answering systems. Specifically, it relates to an evaluation method, apparatus, and electronic device. Background Technology

[0002] In the field of artificial intelligence, multimodal large language models have attracted much attention because they can handle various types of data (e.g., text data, image data, etc.). However, it is precisely because of the diversity of their input types and the complexity of their model structures that evaluating multimodal large language models faces enormous challenges. Summary of the Invention

[0003] This disclosure provides a testing method, apparatus, and electronic device.

[0004] According to a first aspect of this disclosure, an evaluation method is provided, comprising:

[0005] Acquire target data and related evaluation reference information; the target data includes target images and target questions.

[0006] The target image and target question are input into a multimodal large language model, so that the multimodal large language model can be used to obtain the reasoning answer for the target question based on the target image.

[0007] Based on the target question, evaluation reference information, and inference answers, evaluation results are obtained for multimodal large language models.

[0008] According to a second aspect of this disclosure, a testing apparatus is provided, comprising:

[0009] The data acquisition unit is used to acquire target data and evaluation reference information related to the target data; the target data includes target images and target questions.

[0010] The reasoning unit is used to input the target image and the target question into the multimodal large language model, so as to use the multimodal large language model to obtain the reasoning answer for the target question based on the target image;

[0011] The evaluation unit is used to obtain evaluation results for a multimodal large language model based on the target question, evaluation reference information, and inference answer.

[0012] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0013] At least one processor;

[0014] The memory that is communicatively connected to the at least one processor;

[0015] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method provided in the first aspect of this disclosure.

[0016] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method provided according to a first aspect of this disclosure.

[0017] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to a first aspect of this disclosure.

[0018] Using this disclosure can improve the accuracy of evaluation results for MLLM.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0020] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0021] Figure 1 A flowchart illustrating an evaluation method provided in an embodiment of this disclosure;

[0022] Figure 2 A flowchart illustrating the completeness of an evaluation method provided in this disclosure embodiment;

[0023] Figure 3 The target image is used in a specific example provided in this embodiment of the disclosure;

[0024] Figure 4 This is a schematic diagram illustrating an application scenario of an evaluation method provided in an embodiment of this disclosure;

[0025] Figure 5 A schematic structural block diagram of an evaluation device provided in an embodiment of this disclosure;

[0026] Figure 6 This is a schematic structural block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0028] As described in the background section, in the field of artificial intelligence, multimodal large language models (MLLMs) have attracted much attention due to their ability to handle various types of data. However, it is precisely because of the diversity of their input types and the complexity of their model structures that evaluating MLLMs faces significant challenges.

[0029] Currently, assessments of MLLM primarily rely on multiple-choice and / or true / false question formats. In the assessment process, for multiple-choice questions, MLLM simply selects a target answer from multiple candidate options; for true / false questions, it selects a target result from a given set of correct and incorrect options. However, the inventors have found that this assessment method mainly evaluates MLLM's recognition, perception, and reasoning abilities, lacking direct assessment of its conversational and visual abilities. Therefore, it cannot obtain accurate assessment results for MLLM.

[0030] To address the above problems, this disclosure provides an evaluation method that can be applied to electronic devices. The electronic device can be a server, a workbench, a mainframe computer, a conventional computer (desktop computer, laptop computer, etc.), or other similar computing devices. The following will be combined with... Figure 1 The flowchart shown illustrates an evaluation method provided by an embodiment of this disclosure. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described in the flowchart may be performed in a different order.

[0031] Step S101: Obtain the target data and the evaluation reference information related to the target data.

[0032] The target data may include the target image and the target question.

[0033] Here, the target image can be any one of the following used to evaluate MLLM's basic visual perception ability, fine-grained image understanding ability, complex visual reasoning ability, or video understanding and reasoning ability. Correspondingly, when the target image is used to evaluate MLLM's basic visual perception ability, the target question can be a question related to the target image and posed specifically to MLLM's basic visual perception ability; when the target image is used to evaluate MLLM's fine-grained image understanding ability, the target question can be a question related to the target image and posed specifically to MLLM's fine-grained image understanding ability; when the target image is used to evaluate MLLM's complex visual reasoning ability, the target question can be a question related to the target image and posed specifically to MLLM's complex visual reasoning ability; and when the target image is used to evaluate MLLM's video understanding and reasoning ability, the target question can be a question related to the target image and posed specifically to MLLM's video understanding and reasoning ability.

[0034] Furthermore, in this embodiment of the disclosure, the evaluation reference information may include descriptive information about the target image and / or reference answers to the target question. That is, the evaluation reference information may be descriptive information about the target image, reference answers to the target question, or both descriptive information about the target image and reference answers to the target question.

[0035] Step S102: Input the target image and the target question into the MLLM, so that the MLLM can be used to obtain the reasoning answer for the target question based on the target image.

[0036] It is understood that, in the embodiments of this disclosure, after the target image and the target question are input into the MLLM, the MLLM can automatically obtain the reasoning answer for the target question based on the target image.

[0037] Step S103: Based on the target question, assessment reference information, and reasoning answer, obtain the assessment results for MLLM.

[0038] In one example, the reasoning answer can be evaluated based on the target question and assessment reference information to obtain an assessment result for MLLM. This assessment result can be a capability score, used to characterize the performance of MLLM.

[0039] The evaluation method provided in this disclosure can acquire target data including a target image and a target question, and feed the target image and target question as input data to an MLLM. The MLLM, based on the target image, obtains a reasoned answer to the target question. Then, based on the target question, evaluation reference information related to the target data, and the reasoned answer, an evaluation result for the MLLM is obtained. This evaluation method can not only evaluate the MLLM's recognition, perception, and reasoning abilities, but also, because the MLLM's input data includes a target image and a target question, and the output result is a reasoned answer to the target question, its conversational visual abilities are inevitably used in the process of obtaining the output result based on the input data. Therefore, this method can directly evaluate the MLLM's conversational and visual abilities, thereby improving the accuracy of the evaluation results for the MLLM.

[0040] Furthermore, it should be noted that in this embodiment of the disclosure, after obtaining multiple evaluation results for MLLM, a final evaluation result for MLLM can be obtained based on the multiple evaluation results.

[0041] It should also be noted that, in this embodiment of the present disclosure, when performing step S101, target data can be obtained first, and then evaluation reference information related to the target data can be obtained. Moreover, while obtaining the evaluation reference information related to the target data, the target image and target question can be input into MLLM, so that MLLM can be used to obtain the reasoning answer for the target question based on the target image, thereby improving the evaluation efficiency of MLLM.

[0042] In some optional implementations, step S101, i.e., "acquiring target data", may include:

[0043] Identify multiple application categories for MLLM;

[0044] Construct multiple candidate data sets that correspond one-to-one with multiple application categories; wherein each candidate data set includes multiple candidate data sets, and each candidate data set includes a candidate image and a candidate question;

[0045] Each candidate data point in multiple candidate data groups is used as the target data.

[0046] In this context, multiple application categories can be multiple first-level categories, and each of these first-level categories can have multiple second-level categories. Therefore, multiple application categories can also be multiple second-level categories. Furthermore, each of these second-level categories can have multiple third-level categories, so multiple application categories can also be multiple third-level categories.

[0047] In one example, the multiple first-level categories, the multiple second-level categories under each of the multiple first-level categories, and the multiple third-level categories under each of the multiple second-level categories can be represented as shown in Table 1 below:

[0048] Table 1

[0049]

[0050]

[0051] Furthermore, in one example, "constructing multiple candidate data sets corresponding one-to-one with multiple application categories" can include:

[0052] Get the number of instances that match the target category in the actual application of MLLM; where the target category is each of multiple application categories.

[0053] Based on the number of instances, determine the number of targets corresponding to the target category;

[0054] Construct candidate data groups corresponding to the target category, and ensure that the candidate data groups include the target number of candidate data.

[0055] In a specific example, after obtaining the number of instances matching the target category in the actual application of MLLM, a value positively correlated with the number of instances and located within a preset numerical range can be obtained as the target number corresponding to the target category. That is, in the actual application of MLLM, the more instances matching the target category, the larger the target number corresponding to the target category; conversely, the fewer instances matching the target category, the smaller the target number corresponding to the target category. The preset numerical range can be set according to application requirements, for example, it can be set to [30, 600], and this embodiment does not impose any limitations on this. In a more specific example, after obtaining the number of instances matching the target category in the actual application of MLLM, the product of the number of instances and a preset ratio can be calculated. If the product is less than the minimum value in the preset numerical range, the minimum value in the preset numerical range is used as the target number corresponding to the target category; if the product is within the preset numerical range, the product is used as the target number corresponding to the target category; if the product is greater than the maximum value in the preset numerical range, the maximum value in the preset numerical range is used as the target number corresponding to the target category. The preset ratio can be set according to application requirements. For example, it can be set to a ratio value less than or equal to 1, such as 0.01. This embodiment does not limit this.

[0056] After determining the number of targets corresponding to the target category based on the number of instances, a candidate data group corresponding to the target category can be constructed, ensuring that the candidate data group includes the target number of candidate data. The candidate data can include candidate images and candidate questions, both of which can be obtained from MLLM application instances, and will not be elaborated upon here.

[0057] For example, multiple application categories are multiple secondary categories, and the preset ratio is 0.01.

[0058] Suppose that in a practical application of MLLM, the number of instances matching the second-level category "natural image" is 8600. Then, when using "natural image" as the target category, the number of instances corresponding to the target category can be determined to be 8600 × 0.01 = 86. A candidate data set corresponding to the target category can be constructed, ensuring that the candidate data set includes 86 candidate data points. Further suppose that in a practical application of MLLM, the number of instances matching the second-level category "sense recognition" is 34100. Then, when using "sense recognition" as the target category... The number of targets corresponding to the target category can be determined to be 34100 × 0.01 = 341, and a candidate data group corresponding to the target category can be constructed, with the candidate data group including 341 candidate data. Furthermore, assuming that in the actual application of MLLM, the number of instances matching the secondary category "image creation" is 8400, then when the secondary category "common sense recognition" is taken as the target category, the number of targets corresponding to the target category can be determined to be 8400 × 0.01 = 84, and a candidate data group corresponding to the target category can be constructed, with the candidate data group including 84 candidate data.

[0059] Finally, we can obtain multiple candidate data sets as shown in Table 2 below.

[0060] Table 2

[0061]

[0062]

[0063] In other words, in the above example, 13 candidate data groups can be obtained, which contain a total of 1606 candidate data. Therefore, each candidate data in the 13 candidate data groups can be used as the target data, and steps S102 and S103 can be executed. That is, each candidate data in the 1606 candidate data can be used as the target data (including the target image and the target question), and the target image and the target question can be input into MLLM. Using MLLM, based on the target image, the reasoning answer for the target question can be obtained. Then, based on the target question, the evaluation reference information related to the target data, and the reasoning answer, the evaluation result for MLLM can be obtained.

[0064] Ultimately, 1606 test results can be obtained. Then, based on these 1606 results, a final assessment result for MLLM can be derived. For example, the 1606 results can be merged to obtain a fused result (specifically, if the assessment result is a competency score, the fused result can be the average of the 1606 results), and this fused result can then be used as the final assessment result for MLLM.

[0065] Through the above methods, in this embodiment of the disclosure, multiple application categories of MLLM can be determined, and multiple candidate data groups corresponding one-to-one with the multiple application categories can be constructed. Then, each candidate data in the multiple candidate data groups is used as target data (including target image and target question) to evaluate MLLM. In this way, it can be ensured that the evaluation of MLLM is broad and comprehensive, covering multiple application categories of MLLM, thus further improving the accuracy of the evaluation results for MLLM.

[0066] Furthermore, when constructing multiple candidate data sets corresponding one-to-one with multiple application categories, each application category can be used as the target category. The number of instances matching the target category in actual MLLM applications is then determined. Based on this number of instances, a target number corresponding to the target category is determined to construct the candidate data sets corresponding to the target categories, ensuring that each candidate data set includes the target number of candidate data points. In other words, when evaluating MLLM, the number of candidate data points in each of the multiple candidate data sets is determined based on the number of instances matching the application category corresponding to that candidate data set in actual MLLM applications. This ensures that the distribution of all candidate data is consistent with the instance distribution of MLLM in real-world application scenarios, thereby making the evaluation results for MLLM more representative and relevant to real-world applications.

[0067] In some optional implementations, step S103, namely, "obtaining the assessment results for MLLM based on the target question, assessment reference information, and reasoning answer," may include:

[0068] Obtain multiple assessment dimensions;

[0069] Generate assessment prompts that include multiple assessment dimensions;

[0070] Using the assessment model and following the assessment prompts, based on the target question, assessment reference information, and reasoned answers, the assessment results for MLLM are obtained.

[0071] The evaluation dimensions may include at least one of accuracy, relevance (specifically, the relevance between the inferred answer and the target question), and emotional expression; the evaluation prompts are used to instruct the evaluation model on how to obtain the evaluation results for MLLM based on the target question, evaluation reference information, and inferred answer; the evaluation model may be another MLLM with better performance, or it may be a large language model (LLM) with better performance.

[0072] In one example, "using the assessment model, following the assessment prompts, and based on the target question, assessment reference information, and inferred answers, to obtain the assessment results for MLLM" can include:

[0073] Using the assessment model and following the assessment prompts, based on the target question, assessment reference information, and reasoned answers, multiple assessment sub-results are obtained that correspond one-to-one with multiple assessment dimensions and are specific to MLLM.

[0074] Based on multiple evaluation sub-results, the evaluation results for MLLM are obtained.

[0075] For each of the multiple evaluation sub-results, the evaluation sub-result can be a single-dimensional score, used to characterize the performance of MLLM under the evaluation dimension corresponding to that evaluation sub-result. Based on this, in a specific example, after obtaining multiple evaluation sub-results, the sum of the multiple evaluation sub-results can be used as the evaluation result for MLLM.

[0076] In the above example, after obtaining multiple assessment dimensions and generating assessment prompts for these dimensions, the assessment model, based on the target question, assessment reference information, and inferred answers, generates multiple assessment sub-results that correspond one-to-one with the multiple assessment dimensions and are specific to MLLM. Based on these sub-results, the final assessment result for MLLM is obtained. Therefore, the assessment result for MLLM can comprehensively represent the performance strength of MLLM across multiple assessment dimensions and has strong reliability.

[0077] Furthermore, as described above, in this embodiment of the disclosure, the evaluation reference information may include descriptive information for the target image and / or reference answers for the target question.

[0078] The evaluation reference information can be obtained by processing the target image and / or target problem through automatic annotation and / or manual annotation.

[0079] In one example, the target image and / or target problem can be processed first using automatic annotation to obtain initial reference information, and then manually annotated to verify the initial reference information to obtain evaluation reference information. Automatic annotation can utilize other high-performance MLLMs to process the target image and / or target problem to obtain initial reference information.

[0080] Furthermore, it should be noted that in this embodiment of the disclosure, the evaluation reference information may include a reference answer to the target question only if there is an objective answer corresponding to the target question, and the reference answer may include an objective answer corresponding to the target question.

[0081] For example, when the assessment reference information includes descriptive information about the target image and / or a reference answer to the target question, the assessment prompts may include:

[0082]

[0083]

[0084] Furthermore, in one example, the evaluation reference information includes descriptive information about the target image. Therefore, preferably, "using the evaluation model, following the evaluation prompts, and based on the target question, evaluation reference information, and inferred answer, to obtain the evaluation result for MLLM" may include:

[0085] Obtain the LLM;

[0086] Using LLM as an assessment model, we can obtain assessment results for MLLM based on the target question, descriptive information, and inferred answers, according to the assessment prompts.

[0087] In a specific example, an assessment model can be used to obtain multiple assessment sub-results that correspond one-to-one with multiple assessment dimensions and are specific to MLLM, based on the assessment prompts, target questions, assessment reference information, and reasoned answers, according to the assessment prompts. Based on the multiple assessment sub-results, an assessment result for MLLM can be obtained.

[0088] For example, if the evaluation reference information is descriptive information about the target image, then the evaluation model can be used to obtain a credible answer to the target question based on the descriptive information and the evaluation prompts. The inferred answer can then be compared with the credible answer across multiple evaluation dimensions to obtain multiple evaluation sub-results that correspond one-to-one with each evaluation dimension and are specific to MLLM. Finally, based on these multiple evaluation sub-results, the evaluation result for MLLM can be obtained. As another example, if the evaluation reference information includes both descriptive information about the target image and a reference answer to the target question, then the evaluation model can be used to obtain a credible answer to the target question based on the descriptive information and the evaluation prompts. The inferred answer can then be compared with both the credible and reference answers across multiple evaluation dimensions to obtain multiple evaluation sub-results that correspond one-to-one with each evaluation dimension and are specific to MLLM. Finally, based on these multiple evaluation sub-results, the evaluation result for MLLM can be obtained.

[0089] In the above examples, the evaluation reference information may include descriptive information about the target image and / or a reference answer to the target question. Furthermore, when the evaluation reference information includes descriptive information about the target image, an LLM (Multiple Linear Model) can be obtained and used as the evaluation model. This model, based on the evaluation prompts, the target question, descriptive information, and inferred answer, yields the evaluation results for MLLM. Since LLM typically only has the capability to process text data and not image data, using an LLM as the evaluation model when the evaluation reference information includes descriptive information about the target image not only ensures a smooth evaluation process but also eliminates the need for image data processing, thereby improving the evaluation efficiency of MLLM.

[0090] In summary, through the above methods, this embodiment of the disclosure can obtain multiple evaluation dimensions and generate evaluation prompts including these dimensions. Then, using the evaluation model, based on the evaluation prompts, the target question, evaluation reference information, and inferred answers, an evaluation result for MLLM can be obtained. On the one hand, this allows the evaluation results for MLLM to comprehensively characterize the performance of MLLM across multiple evaluation dimensions, thus ensuring the reliability of the evaluation results. On the other hand, after obtaining multiple evaluation dimensions and generating evaluation prompts including these dimensions, the evaluation model can be used to obtain the evaluation results for MLLM based on the target question, evaluation reference information, and inferred answers, thereby improving the evaluation efficiency for MLLM.

[0091] The following will combine Figure 2 The complete process of an evaluation method provided by the embodiments of this disclosure will be described.

[0092] First, multiple application categories of MLLM are identified, and multiple candidate data sets are constructed, each corresponding to one of the application categories. Then, each candidate data set in the multiple candidate data sets is used as the target data. Each candidate data set includes multiple candidate data sets, and each candidate data set includes a candidate image and a candidate question; the target data may include a target image and a target question.

[0093] In one example, “constructing multiple candidate data sets corresponding one-to-one with multiple application categories” can include:

[0094] Get the number of instances that match the target category in the actual application of MLLM; where the target category is each of multiple application categories.

[0095] Based on the number of instances, determine the number of targets corresponding to the target category;

[0096] Construct candidate data groups corresponding to the target category, and ensure that the candidate data groups include the target number of candidate data.

[0097] Subsequently, evaluation reference information related to the target data can be obtained. This evaluation reference information may include descriptive information about the target image and / or reference answers to the target questions.

[0098] The evaluation reference information can be obtained by processing the target image and / or target problem through automatic annotation and / or manual annotation.

[0099] In one example, the target image and / or target problem can be processed first using automatic annotation to obtain initial reference information, and then manually annotated to verify the initial reference information to obtain evaluation reference information. Automatic annotation can utilize other high-performance MLLMs to process the target image and / or target problem to obtain initial reference information.

[0100] Furthermore, it should be noted that in this embodiment of the disclosure, the evaluation reference information may include a reference answer to the target question only if there is an objective answer corresponding to the target question, and the reference answer may include an objective answer corresponding to the target question.

[0101] Next, the target image and the target question are input into the MLLM so that the MLLM can be used to obtain the reasoning answer to the target question based on the target image.

[0102] Next, multiple assessment dimensions can be obtained, and assessment prompts including multiple assessment dimensions can be generated. Then, LLM can be used as the assessment model. By using the assessment model, according to the assessment prompts, based on the target question, assessment reference information and reasoning answer, multiple assessment sub-results corresponding one-to-one with multiple assessment dimensions and specific to MLLM can be obtained. Based on the multiple assessment sub-results, the assessment result for MLLM can be obtained.

[0103] For the specific functions and examples of each step in this embodiment, please refer to the relevant descriptions of the corresponding steps in the foregoing evaluation method embodiments, which will not be repeated here.

[0104] The evaluation method provided in the embodiments of this disclosure will be described below with reference to specific examples.

[0105] In this example, the target data includes the target image and the target question. The target image is as follows: Figure 3 As shown; the target question is: What's funny about this picture? Please describe it frame by frame.

[0106] In this example, the evaluation reference information related to the target data includes descriptive information about the target image: This image shows a humorous misuse. The image is divided into three parts: On the left is a mobile phone with a blue VGA computer monitor cable plugged into its charging port, which is obviously incompatible, as mobile phones typically use smaller charging cables. The upper right part is a packaging box labeled "CABLE MANIA" and "For Lightning Cable," meaning that this box was originally designed for Lightning charging cables, not VGA cables. The lower right part shows a hand holding a VGA connector, one end of which has been modified to fit into a Lightning charging port. The whole scene appears somewhat absurd and comical because this modification is neither practical nor scientific.

[0107] Inputting the target image and target question into MLLM, and leveraging MLLM, the reasoned answer to the target question based on the target image is: The absurdity of this image lies in its depiction of a mobile phone connected to a device that looks like a computer mouse, which in turn is connected to a device that looks like a computer keyboard. This unusual connection creates a comical scenario, as mobile phones are not typically connected to computer mice or keyboards. Furthermore, the image also shows a person holding the mobile phone, further emphasizing this unusual connection.

[0108] After obtaining multiple assessment dimensions, assessment prompts are generated that include these dimensions, specifically:

[0109]

[0110]

[0111] Finally, LLM is used as the assessment model. Using this model, and following the assessment prompts, based on the target question, assessment reference information, and inferred answers, multiple assessment sub-results corresponding one-to-one with multiple assessment dimensions and specific to MLLM are obtained. Based on these multiple assessment sub-results, the final assessment result for MLLM is obtained, specifically as follows:

[0112]

[0113] Please see Figure 4 This is a schematic diagram illustrating an application scenario of an evaluation method provided in an embodiment of this disclosure.

[0114] As described above, the evaluation method provided in this disclosure is applied to electronic devices. These electronic devices can be servers, workbenches, mainframe computers, conventional computers, or other similar computing devices.

[0115] Electronic devices are used for:

[0116] Acquire target data and related evaluation reference information; the target data includes target images and target questions.

[0117] Input the target image and the target question into the MLLM, and use the MLLM to obtain the reasoning answer for the target question based on the target image;

[0118] Based on the target question, assessment reference information, and reasoning answers, the assessment results for MLLM are obtained.

[0119] It should be noted that, in the embodiments disclosed herein, Figure 4 The schematic diagrams shown are for illustrative purposes only and are not restrictive. Those skilled in the art can use them as a basis for their own interpretation. Figure 4 The examples may be modified in various obvious ways and / or substitutions, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of this disclosure.

[0120] To better implement the evaluation method, this disclosure also provides an evaluation device that can be integrated into an electronic device. The electronic device can be a server, a workbench, a mainframe computer, a conventional computer, or other similar computing device. The following will be combined with... Figure 5 The schematic block diagram shown illustrates a testing device 500 provided in the disclosed embodiment.

[0121] The testing device 500 includes:

[0122] The data acquisition unit 501 is used to acquire target data and evaluation reference information related to the target data; wherein, the target data includes target images and target questions;

[0123] The reasoning unit 502 is used to input the target image and the target question into the MLLM, so as to use the MLLM to obtain the reasoning answer for the target question based on the target image;

[0124] The assessment unit 503 is used to obtain assessment results for MLLM based on the target question, assessment reference information, and reasoned answers.

[0125] In some alternative implementations, the evaluation unit 503 is used for:

[0126] Obtain multiple assessment dimensions;

[0127] Generate assessment prompts that include multiple assessment dimensions;

[0128] Using the assessment model and following the assessment prompts, based on the target question, assessment reference information, and reasoned answers, the assessment results for MLLM are obtained.

[0129] In some alternative implementations, the evaluation unit 503 is used for:

[0130] Using the assessment model and following the assessment prompts, based on the target question, assessment reference information, and reasoned answers, multiple assessment sub-results are obtained that correspond one-to-one with multiple assessment dimensions and are specific to MLLM.

[0131] Based on multiple evaluation sub-results, the evaluation results for MLLM are obtained.

[0132] In some alternative implementations, the evaluation reference information includes descriptive information for the target image and / or reference answers for the target question.

[0133] In some alternative implementations, the evaluation reference information includes descriptive information for the target image;

[0134] Assessment Unit 503 is used for:

[0135] Obtain a large language model;

[0136] Using a large language model as an evaluation model, the evaluation results for MLLM can be obtained based on the target question, descriptive information, and inference answer, according to the evaluation prompts.

[0137] In some optional implementations, the data acquisition unit 501 is used for:

[0138] Identify multiple application categories for MLLM;

[0139] Construct multiple candidate data sets that correspond one-to-one with multiple application categories; wherein each candidate data set includes multiple candidate data sets, and each candidate data set includes a candidate image and a candidate question;

[0140] Each candidate data point in multiple candidate data groups is used as the target data.

[0141] In some optional implementations, the data acquisition unit 501 is used for:

[0142] Get the number of instances that match the target category in the actual application of MLLM; where the target category is each of multiple application categories.

[0143] Based on the number of instances, determine the number of targets corresponding to the target category;

[0144] Construct candidate data groups corresponding to the target category, and ensure that the candidate data groups include the target number of candidate data.

[0145] In this embodiment of the disclosure, the specific functions and examples of each unit in the testing device 500 can be found in the relevant descriptions of the corresponding steps in the aforementioned testing method embodiments, and will not be repeated here.

[0146] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0148] Figure 6 A schematic structural block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as in-vehicle computing devices, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0149] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0150] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of renderers, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0151] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a GPU, various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as evaluation methods. For example, in some embodiments, the evaluation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the evaluation methods described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured as an evaluation method by any other suitable means (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] Program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM) or flash memory, optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a rendering device (e.g., a cathode ray tube (CRT) renderer or a liquid crystal display (LCD)) for rendering information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices are also used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0156] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0157] A computer system can include client and server components. Clients and servers are generally located far apart and typically interact via a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server in a distributed system, or a server incorporating blockchain technology.

[0158] This disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute an evaluation method.

[0159] This disclosure also provides a computer program product, including a computer program that implements an evaluation method when executed by a processor.

[0160] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure is achieved, and this is not limited herein. Furthermore, in this disclosure, relational terms such as "first," "second," and "third" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Additionally, "multiple" in this disclosure can be understood as at least two.

[0161] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An assessment method, comprising: Acquire target data and evaluation reference information related to the target data; wherein, the target data includes target images and target questions; The target image and the target question are input into a multimodal large language model, so that the multimodal large language model can be used to obtain a reasoning answer for the target question based on the target image. Based on the target question, the evaluation reference information, and the reasoning answer, the evaluation results for the multimodal large language model are obtained; The acquisition of target data includes: Determine multiple application categories for the multimodal large language model; Each of the multiple application categories is taken as the target category. The number of instances matching the target category in the actual application of the multimodal large language model is obtained. Based on the number of instances, the target number corresponding to the target category is determined. A candidate data group corresponding to the target category is constructed, and the candidate data group includes the target number of candidate data. The candidate data group includes multiple candidate data, and each candidate data includes a candidate image and a candidate question. Each candidate data in a plurality of candidate data groups is used as the target data; wherein, the plurality of candidate data groups correspond one-to-one with the plurality of application categories.

2. The method according to claim 1, wherein, The evaluation results for the multimodal large language model, obtained based on the target question, the evaluation reference information, and the reasoning answer, include: Obtain multiple assessment dimensions; Generate assessment prompts that include the multiple assessment dimensions; Using the assessment model, and following the assessment prompts, based on the target question, the assessment reference information, and the reasoning answer, an assessment result is obtained for the multimodal large language model.

3. The method according to claim 2, wherein, The step of using the assessment model, according to the assessment prompts, and based on the target question, the assessment reference information, and the reasoning answer, to obtain the assessment result for the multimodal large language model includes: Using the assessment model, and following the assessment prompts, based on the target question, the assessment reference information, and the reasoning answer, multiple assessment sub-results are obtained that correspond one-to-one with the multiple assessment dimensions and are specific to the multimodal large language model. Based on the multiple evaluation sub-results, the evaluation results for the multimodal large language model are obtained.

4. The method according to any one of claims 2, wherein, The evaluation reference information includes descriptive information about the target image and / or reference answers to the target question.

5. The method according to claim 4, wherein, The evaluation reference information includes descriptive information for the target image; The step of using the assessment model, according to the assessment prompts, and based on the target question, the assessment reference information, and the reasoning answer, to obtain the assessment result for the multimodal large language model includes: Obtain a large language model; The large language model is used as an evaluation model to obtain evaluation results for the multimodal large language model based on the target question, the descriptive information, and the inference answer, according to the evaluation prompts.

6. A testing device, comprising: A data acquisition unit is used to acquire target data and evaluation reference information related to the target data; wherein, the target data includes target images and target questions; The reasoning unit is used to input the target image and the target question into a multimodal large language model, so as to use the multimodal large language model to obtain a reasoned answer to the target question based on the target image; The evaluation unit is used to obtain evaluation results for the multimodal large language model based on the target question, the evaluation reference information, and the reasoning answer. The data acquisition unit is used for: Determine multiple application categories for the multimodal large language model; Each of the multiple application categories is taken as the target category. The number of instances matching the target category in the actual application of the multimodal large language model is obtained. Based on the number of instances, the target number corresponding to the target category is determined. A candidate data group corresponding to the target category is constructed, and the candidate data group includes the target number of candidate data. The candidate data group includes multiple candidate data, and each candidate data includes a candidate image and a candidate question. Each candidate data in a plurality of candidate data groups is used as the target data; wherein, the plurality of candidate data groups correspond one-to-one with the plurality of application categories.

7. The apparatus according to claim 6, wherein, The evaluation unit is used for: Obtain multiple assessment dimensions; Generate assessment prompts that include the multiple assessment dimensions; Using the assessment model, and following the assessment prompts, based on the target question, the assessment reference information, and the reasoning answer, an assessment result is obtained for the multimodal large language model.

8. The apparatus according to claim 7, wherein, The evaluation unit is used for: Using the assessment model, and following the assessment prompts, based on the target question, the assessment reference information, and the reasoning answer, multiple assessment sub-results are obtained that correspond one-to-one with the multiple assessment dimensions and are specific to the multimodal large language model. Based on the multiple evaluation sub-results, the evaluation results for the multimodal large language model are obtained.

9. The apparatus according to any one of claims 7, wherein, The evaluation reference information includes descriptive information about the target image and / or reference answers to the target question.

10. The apparatus according to claim 9, wherein, The evaluation reference information includes descriptive information for the target image; The evaluation unit is used for: Obtain a large language model; The large language model is used as an evaluation model to obtain evaluation results for the multimodal large language model based on the target question, the descriptive information, and the inference answer, according to the evaluation prompts.

11. An electronic device, comprising: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 5.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Question and answer method, device and equipment based on large model and storage medium

    CN118246537A

  • Evaluation method and device for visual question and answer task, medium and computer program product

    CN118467709A