Multimodal large model evaluation system and method based on cognitive psychology
By constructing a multi-modal large model evaluation system based on cognitive psychology, multi-dimensional evaluation indicators and question banks are built, which solves the problem of single-dimensional evaluation in traditional evaluation schemes, realizes comprehensive performance evaluation of multi-modal large models, and improves the model's processing performance and user experience in the target task.
Patent Information
- Application Number
- CN202511476246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-16
AI Technical Summary
Existing multimodal large model evaluation schemes mainly focus on single-dimensional performance evaluation, lacking consideration of the model's comprehensive capabilities across multiple dimensions. This results in an inability to fully measure the model's actual capabilities and application potential. Furthermore, task-oriented benchmarking cannot accurately reflect the model's true performance, affecting subsequent fine-tuning and application.
A multimodal large model evaluation system based on cognitive psychology is adopted. By constructing evaluation indicators and evaluation question banks, and combining information processing models, the system evaluates the multi-dimensional performance of the multimodal large model in terms of perception, attention, memory and reasoning, and generates objective performance evaluation results.
It enables objective and comprehensive performance evaluation of large multimodal models, accurately reflects the model's true capabilities, and facilitates targeted fine-tuning and optimization by users, thereby improving the model's processing performance and user experience in the target task.
Smart Images

Figure CN120929795B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a multi-modal large model evaluation system and method based on cognitive psychology. BACKGROUND
[0002] In recent years, with the continuous innovation of artificial intelligence technology, multi-modal large models break through the limitations of traditional single-modal (such as only text, only image) models and can simultaneously understand, correlate and generate information of multiple modalities (such as images, videos, audio, text, etc.), providing more rich and accurate information expression and attracting widespread attention. In the face of complex and highly integrated multi-modal large models, how to effectively evaluate the comprehensive performance of multi-modal large models so that users can debug the models for specific tasks to improve the processing performance of the models has become a problem to be solved.
[0003] However, current multi-modal large model performance evaluation schemes mostly only focus on single-dimensional performance evaluation. For example, only the performance of the model in language understanding or image recognition is concerned, and the multi-dimensional comprehensive ability of the model is not considered. This leads to the inability to comprehensively measure the actual ability and application potential of the model. Moreover, current evaluation schemes often rely on task-oriented benchmark tests and can only evaluate the performance of the model for specific tasks, i.e., the evaluation results reflect the knowledge reserve of the model in a specific field, rather than the objective and accurate true performance, thereby misleading users, affecting the subsequent fine-tuning and application of the model, slowing down the task progress of users, and affecting the user experience. SUMMARY
[0004] Therefore, the present application aims to provide a multi-modal large model evaluation system and method based on cognitive psychology to achieve objective and comprehensive performance evaluation of multi-modal large models and accurately reflect the true performance of multi-modal large models.
[0005] To achieve the above-mentioned purpose, the technical solutions of the present application are as follows:
[0006] The first aspect of the embodiment of the present application provides a multi-modal large model evaluation system based on cognitive psychology, which comprises:
[0007] The evaluation execution module is configured to extract corresponding evaluation questions from an evaluation question bank according to at least one evaluation index specified by a user; the evaluation index is determined based on an information processing model and cognitive psychology theory and is used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the evaluation questions are constructed based on the evaluation index; and the evaluation questions are input into the multi-modal large model to be evaluated to obtain output results.
[0008] An analysis module is configured to calculate an evaluation score of the multi-modal large model based on the output result and the correct answer of the evaluation question, and generate a performance evaluation result of the multi-modal large model based on the evaluation score.
[0009] Optionally, the system further comprises a theory mapping module and a question bank construction module.
[0010] The theory mapping module is configured to set a plurality of evaluation indexes based on an information processing model and a cognitive psychology theory, specifically including a plurality of first-level indexes and a plurality of second-level indexes belonging to each first-level index; the plurality of first-level indexes include perceptual ability, attention, memory, and reasoning ability; the plurality of second-level indexes include visual perception and spatial perception belonging to perceptual ability, attention concentration and naming recognition belonging to attention, short-term visual memory and episodic memory belonging to memory, and analogical reasoning ability and programming ability belonging to reasoning ability.
[0011] The question bank construction module is configured to perform the following steps:
[0012] Obtain a plurality of types of original data, including original multiple-choice questions, original question-answer questions, and original fill-in-the-blank questions;
[0013] Based on the second-level indexes, the complete structure of the original multiple-choice questions is retained, and the options of the original multiple-choice questions are rearranged in disorder to obtain first-type evaluation questions;
[0014] Based on the second-level indexes, the stems of the original question-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions as second-type evaluation questions;
[0015] Based on each second-level index, the first-type evaluation questions and the second-type evaluation questions are classified and stored to generate an evaluation question bank.
[0016] Optionally, the question bank construction module is configured to process the stems of the original question-answer questions or the original fill-in-the-blank questions based on the second-level indexes to generate multiple-choice questions as second-type evaluation questions, specifically including:
[0017] Based on the stem of the original question-answer question or the stem of the original fill-in-the-blank question, generate one correct option and a plurality of incorrect options, at least one of which has a similarity to the correct option greater than or equal to a first threshold;
[0018] According to the order of similarity from high to low, a plurality of high-similarity incorrect options are screened out;
[0019] Based on the correct option and the plurality of high-similarity incorrect options screened out, generate the second-type evaluation questions.
[0020] Optionally, the question bank construction module is further configured to, after retaining the complete structure of the original selection question and rearranging the options of the original selection question in a disordered sequence, perform the following steps:
[0021] obtain the correct option and the incorrect options of the original selection question;
[0022] calculate the similarity of each incorrect option to the correct option and compare the similarity to a second threshold value;
[0023] in a case where the similarity corresponding to all incorrect options is less than the second threshold value, generate at least one new incorrect option based on the correct option, the new incorrect option having a similarity greater than or equal to the second threshold value;
[0024] replace the incorrect options in the original selection question with the new incorrect options to obtain the first type of test question.
[0025] Optionally, the test execution module is configured to extract corresponding test questions from the test question bank according to at least one test index specified by a user, specifically including:
[0026] obtain the index weight corresponding to each test index specified by the user;
[0027] determine the number of questions to be extracted based on the index weight of each test index; the number of questions is greater than or equal to a third threshold value;
[0028] obtain a corresponding number of test questions from the test question bank based on the number of questions to be extracted corresponding to each test index;
[0029] mix and randomly rearrange the order of the obtained test questions corresponding to each test index to determine the order in which the multi-modal large model processes the test questions.
[0030] Optionally, the analysis module is configured to calculate the test score of the multi-modal large model based on the output result and the correct answer of the test question, specifically including:
[0031] determine the score weight corresponding to each test index based on the index weight corresponding to each test index;
[0032] calculate the accuracy of the output result of the multi-modal large model based on the test question corresponding to each secondary index, and determine the secondary score corresponding to each secondary index based on the accuracy;
[0033] calculate the primary score corresponding to the primary index to which each secondary index belongs based on each secondary score and the corresponding score weight.
[0034] Optionally, the analysis module is configured to generate a performance evaluation result of the multi-modal large model based on the evaluation score, specifically comprising:
[0035] determining an evaluation score interval to which the first-level score of each first-level indicator belongs, to obtain evaluation information corresponding to the evaluation score interval, and determining the evaluation information as the performance evaluation result of the first-level score;
[0036] generating a performance evaluation result of the multi-modal large model based on the performance evaluation result of each first-level score.
[0037] Optionally, the system further comprises an interaction module configured to perform the following steps:
[0038] displaying the selectable first-level indicators and second-level indicators on a user interface, and sending the at least one evaluation indicator specified by the user to the evaluation execution module;
[0039] displaying the evaluation questions and the output result of the multi-modal large model on a user interface;
[0040] after the analysis module generates the performance evaluation result, displaying the performance evaluation result and / or the score corresponding to each evaluation indicator on a user interface.
[0041] According to a second aspect of an embodiment of the present application, a multi-modal large model evaluation method based on cognitive psychology is provided, which is applied to the system provided by the first aspect of the present application, and the method comprises:
[0042] extracting corresponding evaluation questions from an evaluation question bank according to at least one evaluation indicator specified by a user; the evaluation indicator is determined based on an information processing model and cognitive psychology theory, and is used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the evaluation question is constructed based on the evaluation indicator;
[0043] inputting the evaluation question into the multi-modal large model to be evaluated to obtain an output result;
[0044] calculating an evaluation score of the multi-modal large model based on the output result and the correct answer of the evaluation question;
[0045] generating a performance evaluation result of the multi-modal large model based on the evaluation score.
[0046] Optionally, the method further comprises pre-setting a plurality of evaluation indicators, specifically comprising:
[0047] setting at least one first-level indicator based on an information processing model and cognitive psychology theory;
[0048] At least two secondary indexes belonging to the primary index are generated based on each primary index and the Wechsler Intelligence Scale for Children.
[0049] The multi-modal large model evaluation system based on cognitive psychology provided in the present application determines a plurality of evaluation indexes in advance based on an information processing model and a cognitive psychology theory, for a user to select for performance evaluation of a multi-modal large model, and constructs corresponding evaluation questions based on the evaluation indexes to obtain an evaluation question bank. After the user specifies the evaluation indexes of the multi-modal large model, the evaluation execution module extracts corresponding evaluation questions from the evaluation question bank based on the evaluation indexes selected by the user, and inputs the evaluation questions to the multi-modal large model to be evaluated for processing to obtain output results of the multi-modal large model for the evaluation questions. Then, the analysis module calculates corresponding evaluation scores based on the correct answers of the evaluation questions and the output results of the multi-modal large model. The performance evaluation results corresponding to the multi-modal large model are generated based on the evaluation scores.
[0050] The processing mode of human cognitive processing proposed by the present application based on the cognitive psychology theory and the analogy of human cognition and computer information processing flow by the information processing model are used to evaluate the data processing performance of the multi-modal large model. Specifically, the evaluation indexes and the evaluation question bank are constructed based on the cognitive psychology theory and the information processing model for model performance evaluation, thereby establishing the relationship between the deep capabilities of the model and the dimensions of human cognition. Compared with the traditional model performance evaluation scheme only for the knowledge level or single dimension of a specific field, the present application realizes objective, comprehensive and accurate evaluation of the cognitive mechanism and capability of the multi-modal large model, so that the user can fully understand the real performance of the evaluated model in all aspects, and the user can fine-tune and optimize the model according to the target task (for example, a graphic-text question answering task) to be processed, thereby improving the performance of the model for the target processing task, improving the task completion quality, and improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0052] Figure 1 is a schematic diagram of a multi-modal large model evaluation system based on cognitive psychology according to an embodiment of the present application;
[0053] Figure 2 is a workflow diagram of a multi-modal large model evaluation system according to an embodiment of the present application;
[0054] Figure 3is a flowchart of a multi-modal large model evaluation method based on cognitive psychology, which is proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present application.
[0056] It should be understood that the term “one embodiment” or “an embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, “in one embodiment” or “in an embodiment” appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0057] In various embodiments of the present application, it should be understood that the size of the serial number of the following processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0058] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent the same or similar elements unless otherwise represented. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatuses and methods consistent with some aspects as detailed herein.
[0059] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0060] A multi-modal large model is a deep learning model with a huge number of parameters (usually up to tens of billions or even hundreds of billions), which automatically learns the general rules and knowledge of language, image and other patterns from unannotated data through “self-supervised learning”. Its ability not only lies in the parameter scale, but also exhibits complex abilities not explicitly programmed in the training data, such as context learning, complex reasoning, code generation ability, etc., when the model size breaks through the critical point. However, the traditional multi-modal large model evaluation scheme has the following problems:
[0061] (1) Only focus on single dimension evaluation, for example, only pay attention to the performance of the model in language understanding or image recognition, lack of consideration of the comprehensive ability of the model in multiple dimensions. The widely used VQA (Visual Question Answering) and GQA (Graph-based Question Answering) evaluation benchmarks measure the accuracy of the model in basic perception and simple reasoning by predefining a single image and artificially labeled questions. The preset task data set (such as picture description generation, multiple choice knowledge questions) and the accuracy, recall rate and other indicators of the model on the task data set are calculated to measure the performance of the model. This way can only evaluate the single task performance in a static scene, and cannot determine the multi-dimensional processing performance of the model itself. For example, the COCO data set focuses on object positioning accuracy, and the ImageNet data set focuses on classification ability, neither of which involves the model's ability to handle cross-modal correlation, long context dependence, or dynamic sequence.
[0062] (2) The traditional evaluation scheme relies on task-oriented knowledge-based testing, ignoring the focus on model cognitive mechanisms and capabilities, so that the evaluation results reflect the domain knowledge reserve rather than the cognitive mechanism efficiency. For example, a model that performs well in medical image diagnosis may have obtained high scores due to knowledge bias rather than cognitive advantage, making it difficult for developers to determine whether the model's problems are perception errors (such as misreading image features), attention distraction (such as ignoring key areas), or reasoning errors (such as misjudging pathological correlations), thereby preventing developers from effectively fine-tuning the model based on the evaluation results, affecting task processing efficiency and user experience. For example, long text answers actually reflect the attention ability of large models, but traditional evaluation schemes can only evaluate the model's knowledge answering ability from the knowledge level, and the evaluation results cannot accurately reflect the model's attention ability.
[0063] In view of the problems of traditional evaluation schemes, such as single evaluation dimension and ignoring model cognitive process disassembly, the present application customizes multi-dimensional evaluation targets based on information processing models and cognitive psychology theories to establish the relationship between model internal capabilities and human cognitive dimensions, and constructs an evaluation question bank for performance index evaluation of multi-modal large models. By analyzing the processing results of the model on the evaluation questions, the performance of each dimension of the model is objectively and accurately evaluated, the synergy of large models in the human cognitive level is revealed, and users can accurately know the human-like intelligence level of the measured model.
[0064] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0065] Figure 1 is a schematic diagram of a multi-modal large model evaluation system 100 based on cognitive psychology according to an embodiment of the present application. As shown inFigure 1 The system comprises:
[0066] The evaluation execution module 101 is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation index specified by a user; the evaluation index is determined based on an information processing model and cognitive psychology theory, and is used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the evaluation question is constructed based on the evaluation index; and the evaluation question is input into the multi-modal large model to be evaluated to obtain an output result.
[0067] The analysis module 102 is configured to calculate an evaluation score of the multi-modal large model based on the output result and a correct answer of the evaluation question, and generate a performance evaluation result of the multi-modal large model based on the evaluation score.
[0068] In the embodiments of the present application, a plurality of evaluation indexes are determined in advance based on an information processing model and cognitive psychology theory, and these evaluation indexes are used to measure the performance of the multi-modal large model in different dimensions, specifically including perceptual ability, attention, memory, and reasoning ability. Corresponding evaluation questions are constructed based on each evaluation index, and an evaluation question bank is generated.
[0069] When evaluating the multi-modal large model, the user selects the model to be evaluated and the corresponding evaluation index. In actual application, the multi-modal large model to be evaluated can be a mature architecture model such as GPT-4o-mini, Qwen-2.5-vl, Gemini-2-flash, GLM-4v-plus, Doubao-1.5, or a model obtained by customizing or fine-tuning an existing multi-modal large model.
[0070] The evaluation execution module extracts corresponding evaluation questions from the evaluation question bank according to each evaluation index specified by the user to form a question set. Each evaluation question in the question set is preprocessed, the question is converted into a format that can be recognized by the multi-modal large model to be evaluated, and is input into the multi-modal large model for processing, while a timer is started to record the response time, the model output is captured, and the output result of the multi-modal large model is obtained. Optionally, the evaluation questions constructed in the embodiments are all multiple-choice questions. Before the multi-modal large model to be evaluated processes the evaluation questions, the output format of the model is constrained based on the number of options of the evaluation questions, and the model is forced to return one of the options. Optionally, the system uses the API (Application Programming Interface) of the Flask framework to provide model evaluation services to users, and the output result of the multi-modal large model is stored through the Mysql database, which facilitates subsequent analysis of the multi-dimensional performance of the model.
[0071] The analysis module calculates the evaluation scores of the multi-modal large model for each evaluation indicator based on the output results of the multi-modal large model and the correct answers of each evaluation question in the question set, and further generates the corresponding performance evaluation results based on the evaluation scores. Optionally, in the process of the multi-modal large model processing the evaluation questions, the analysis module uses the SQLAlchemy tool to write the output results of the model for each evaluation question into the database through the ORM (Object Relational Mapping) mode, so as to facilitate reading the output results of the model and the correct answers of the evaluation questions from the database for comparison and analysis. Based on the performance evaluation results corresponding to the evaluation indicators, the user can accurately know the real processing capability of the model in different dimensions, and then according to the target task to be executed, the model is fine-tuned to improve the processing performance of the model, and the execution efficiency and completion quality of the target task are improved.
[0072] In this embodiment, a specific human cognitive psychology theory and information processing model are applied to multi-dimensional evaluation of the capability of the multi-modal large model, the learning process of the multi-modal large model is analogized based on the processing mode of human cognitive processing, so as to objectively, comprehensively and accurately evaluate the cognitive mechanism and real capability of the multi-modal large model, and facilitate the user to clearly determine the subsequent optimization direction of the model, and then improve the task completion quality and user experience.
[0073] The system can be used for the whole life cycle of the multi-modal large model, and the user can fine-tune the model before, during and after the model executes the target task based on the performance evaluation results of the multi-modal large model generated by the system. For example, before the unmanned aerial vehicle manufacturer completes the unmanned aerial vehicle flight task by using the multi-modal large model, the necessary conditions for the unmanned aerial vehicle flight are specified to specify the corresponding evaluation indicators (such as spatial perception capability and visual perception capability), and then the performance of the model based on the evaluation indicators is obtained through the system, the direction of subsequent fine-tuning of the model is determined, and the completion quality of the unmanned aerial vehicle flight task is improved.
[0074] As an embodiment of the present application, the system further comprises a theory mapping module and a question bank construction module.
[0075] The theory mapping module is configured to set a plurality of evaluation indicators based on the information processing model and the cognitive psychology theory, specifically including a plurality of first-level indicators and a plurality of second-level indicators belonging to each first-level indicator; the plurality of first-level indicators include perception, attention, memory, and reasoning; the plurality of second-level indicators include visual perception and spatial perception belonging to perception, attention concentration and naming recognition belonging to attention, short-term visual memory and episodic memory belonging to memory, and analogical reasoning ability and programming ability belonging to reasoning.
[0076] The question bank construction module is configured to perform the following steps:
[0077] Obtain a plurality of types of raw data, including: raw multiple-choice questions, raw question-answer questions, and raw fill-in-the-blank questions;
[0078] Based on the secondary indicators, the complete structure of the raw multiple-choice questions is retained, and the options of the raw multiple-choice questions are rearranged in a disordered sequence to obtain first-type evaluation questions;
[0079] Based on the secondary indicators, the stems of the raw question-answer questions or the raw fill-in-the-blank questions are processed to generate multiple-choice questions as second-type evaluation questions;
[0080] Based on each secondary indicator, the first-type evaluation questions and the second-type evaluation questions are classified and stored to generate an evaluation question bank.
[0081] In the embodiments of the present application, the system pre-constructs the mapping relationship between human cognitive dimensions and the core capabilities of the multi-modal large model based on the information processing model and the cognitive psychology theory through the theoretical mapping module, and converts and applies the ideas and methods of specific psychological test paradigms to the evaluation of the cognitive dimensions of the multi-modal large model.
[0082] Specifically, the mapping relationship is established based on the core cognitive dimensions of the PASS cognitive model and the Neisser information processing model, the human cognition is analogized to the information processing system of a computer, the core dimensions of human cognitive processing, i.e., “planning”, “attention”, “simultaneous processing”, and “successive processing”, and the multiple discrete stages of processing data (including: environmental perception of stimuli, preliminary interpretation, to selective focusing of resources filtering irrelevant information by attention mechanism, then entering the memory system, and finally using input information and long-term memory knowledge for complex operations in the thinking, reasoning, and decision-making stages) are combined to determine a plurality of first-level indicators for measuring the multi-dimensional capabilities of the multi-modal large model, including: perceptual ability, attention, memory, and reasoning. Among them, the perceptual ability indicator corresponds to the visual / auditory feature extraction capability of the model; the attention indicator corresponds to the information filtering and focusing capability of the model; the memory indicator corresponds to the context understanding and sequence maintenance capability of the model; and the reasoning indicator corresponds to the cross-modal logical association capability of the model. By setting a plurality of first-level indicators, it is ensured that the evaluation dimensions of the system comprehensively cover the key capabilities of the model.
[0083] In this embodiment, on the basis of each first-level index, each first-level index is further subdivided based on the Wechsler Intelligence Scale for Children to determine a plurality of second-level indexes, and a hierarchical evaluation framework is formed. Through each second-level index, the key capabilities of each dimension of the model are deeply disassembled, the accuracy of the model performance evaluation is improved, the evaluation result can reveal the more detailed performance differences of the model in each key capability dimension, so as to facilitate the user to clearly understand the detailed direction of model optimization, and improve the efficiency of subsequent model training.
[0084] Specifically, the perceptual index is related to the analysis and understanding ability of the model to the input information, including the processing of multi-modal information such as voice, image and text. In this embodiment, based on the perceptual index, the measurement paradigm of spatial intelligence is combined with the building block design and visual puzzle of human cognition to further divide visual perception and spatial perception as subordinate second-level indexes. Among them, the visual perception index quantifies the extraction ability of the model to low-level visual features through fine-grained image recognition, simulating the encoding process of the human retina-cortical pathway to shape and texture; the spatial perception index evaluates the reasoning ability of the model to relative position and spatial topology through three-dimensional scene question and answer.
[0085] The attention index aims to simulate the selective attention process in the process of human brain processing information. In this embodiment, based on the attention index, the test of “visual-symbol” conversion efficiency is combined with the human cognitive coding sub-test to further divide attention concentration and naming recognition as subordinate second-level indexes to simulate the human selective attention mechanism.
[0086] In this embodiment, based on the memory index, the reproduction test of the “number-space” sequence in the memory span test of human cognition is combined to further divide short-term visual memory and situational memory as subordinate second-level indexes. The short-term visual memory index quantifies the retention ability of the model to short-term visual information through multiple rounds of image recognition and analysis tasks; the situational memory index adopts a situational long context and a multi-round dialogue task to evaluate the modeling ability of the model to context-dependent relationships.
[0087] In this embodiment, based on the reasoning index, the standardized evaluation paradigm of non-verbal logical reasoning is combined with the matrix reasoning test of human cognition to further divide analogical reasoning ability and programming ability as subordinate second-level indexes. The analogical reasoning index measures the abstract mapping ability of the model to implicit relationships by visual matrix completion; the programming index evaluates the decomposition and execution ability of the model to structured rules through algorithm and programming tasks.
[0088] In the traditional multi-modal large model evaluation scheme, the evaluation question is too simplified to effectively judge the performance of the model. For example, the multi-modal model comprehensive evaluation benchmark (MME) constructs a binary judgment question bank based on the "perception-cognition" dimension to evaluate object recognition ability and basic logical reasoning. However, the binary choice (yes / no) makes the random guess correct rate as high as 50%, so it is difficult to obtain the real understanding level of the model in practical application. Based on this, the multi-choice selection question based on each secondary index is constructed as the evaluation question in the embodiment, so as to avoid the influence of random guessing of the model on the accuracy of the evaluation result.
[0089] Specifically, based on each secondary index, the applicable data set is obtained from different sources as the original data, and the evaluation question corresponding to each evaluation index is constructed based on the original data. In actual application, public data sets can be collected from multiple third-party sources as original data according to actual needs, such as Google, Huggingface, Kaggle and the like. The obtained original data includes different types of original questions, such as selection questions, question and answer questions and fill-in-the-blank questions. In order to construct multi-choice selection questions matched with the evaluation index, different types of original questions need to be adjusted, as follows:
[0090] (1) For the original selection question, the complete structure is retained to ensure the integrity of the original context and task setting of the data, and the original options of the question are rearranged in disorder to change the position of the correct answer of the question, to generate the first type of evaluation question. By rearranging the original options in disorder, the situation of data leakage (i.e. the model has learned the data in advance, so it directly answers based on memory without thinking and reasoning) is avoided, and the accuracy of the evaluation result is improved;
[0091] (2) For the original question and answer question or the original fill-in-the-blank question, only the original data is used, without directly relying on the original question or the annotation, and by converting the stem of the original question into a multi-choice selection question matched with the evaluation index, the second type of evaluation question is generated, so as to keep consistent with the format of the first type of evaluation question.
[0092] The first type of evaluation question and the second type of evaluation question are merged and stored according to different secondary indexes, to form a standardized and normalized evaluation question bank, which is convenient for quickly querying and pulling the evaluation question corresponding to the evaluation index when evaluating the model.
[0093] In this embodiment, based on a plurality of first-level indicators measuring key capabilities, a plurality of second-level indicators subordinate thereto are further subdivided, thereby realizing deep-level disassembly of key capabilities of each dimension of the model, forming a hierarchical evaluation framework, and improving the accuracy of model performance evaluation. On this basis, based on each second-level indicator, the original data collected is processed to construct a standardized and anti-leakage evaluation question bank, and objective and comprehensive quantitative evaluation of the multi-modal large model is realized through multiple-choice questions matched with the evaluation indicators, and the deep-level human-like cognitive capabilities of the model are accurately reflected.
[0094] Optionally, when a user evaluates the multi-modal large model through the system, the user can specify a first-level indicator or a second-level indicator as needed. In the case where the user specifies a first-level indicator, the evaluation execution module determines all second-level indicators subordinate thereto based on the first-level indicator, and extracts corresponding evaluation questions from the evaluation question bank based on each second-level indicator; in the case where the user specifies a second-level indicator, the evaluation execution module directly extracts evaluation questions from the evaluation question bank according to the second-level indicator.
[0095] As an embodiment of the present application, the question bank construction module is configured to process the stems of the original question-answer questions or the original fill-in-the-blank questions based on the second-level indicators to generate multiple-choice questions as the second type of evaluation questions, specifically including:
[0096] Based on the stem of the original question-answer question or the stem of the original fill-in-the-blank question, one correct option and a plurality of incorrect options are generated, wherein at least one incorrect option has a similarity greater than or equal to a first threshold value with the correct option;
[0097] According to the order of similarity from high to low with the correct option, a plurality of incorrect options with higher similarity are screened out;
[0098] Based on the correct option and the plurality of incorrect options with higher similarity screened out, the second type of evaluation questions are generated.
[0099] In an embodiment, the question bank construction module constructs multiple-choice questions matched with the evaluation indicators as the second type of evaluation questions based on the stems of the original question-answer questions and the original fill-in-the-blank questions. Specifically, based on the stem of the original question-answer question or the stem of the original fill-in-the-blank question, a plurality of options of the multiple-choice question are generated, including one correct option and a plurality of incorrect options. When generating the options, one or more interference items with higher similarity to the correct option are generated, thereby improving the difficulty of the evaluation questions, reducing guessing during the problem-solving process of the multi-modal large model, excavating the deep-level understanding ability, analysis ability, comparison ability and logical judgment ability of the model, and improving the accuracy of the test results.
[0100] Specifically, based on the stem of the original fill-in-the-blank question / short answer question, a correct option and multiple incorrect options are generated. Among them, the similarity between at least one incorrect option and the correct option is not less than a first threshold, that is, at least one interference item is generated, and each interference item has a high similarity with the correct option. According to the similarity from high to low, all generated incorrect options are sorted, and based on the number of options required to generate the test questions, multiple incorrect options with high similarity are screened out. For example, the test question is a four-option selection question, and in addition to the correct answer, three incorrect options with high similarity need to be screened out. Based on the correct option and the multiple incorrect options screened out, the second type of test question is constructed.
[0101] Compared with the traditional multi-modal large model evaluation scheme, in the construction of the test question to be evaluated, the interference item with high similarity with the correct option is generated, thereby improving the complexity of the test question and enhancing the anti-data leakage ability of the test question bank. Moreover, based on the original non-selection question (including the original fill-in-the-blank question and the original short answer question), the second type of test question in the form of a selection question is generated, which can effectively expand the sample number of the test questions in the test question bank, and then a large number of test questions are used to fully detect the model performance and improve the accuracy of the detection result.
[0102] As an embodiment of the present application, the question bank construction module is further configured to, after retaining the complete structure of the original selection question and rearranging the options of the original selection question in a disordered manner, perform the following steps:
[0103] Obtain the correct option and the incorrect option of the original selection question;
[0104] Calculate the similarity of each incorrect option with the correct option and compare it with a second threshold;
[0105] In the case where the similarity of all incorrect options is less than the second threshold, at least one new incorrect option with a similarity greater than or equal to the second threshold is generated based on the correct option;
[0106] Replace the incorrect options in the original selection question with the new incorrect options to obtain the first type of test question.
[0107] In an embodiment, after the options of the original multiple-choice question are shuffled and rearranged, the interference items with high similarity are generated based on the correct option of the original multiple-choice question, and the interference items are used to replace the options other than the correct option, so as to avoid the model from obtaining the answer without analysis due to the leakage of the original multiple-choice question, and to enhance the anti-data leakage capability of the evaluation question bank. Moreover, by introducing the interference items of the correct option into the options of the original multiple-choice question, the complexity of the original multiple-choice question can be improved, and thus the deep cognitive process disassembly capability of the model can be mined when the model is evaluated, instead of being limited to the knowledge level.
[0108] Specifically, the correct option and the incorrect options in the original multiple-choice question are obtained, and the similarity between each incorrect option and the correct option is calculated. The calculated similarity is compared with the second threshold. If there is no incorrect option with a similarity greater than or equal to the second threshold, it is determined that there is no interference item at present, and one or more new incorrect options with a similarity not less than the second threshold are generated based on the correct option as interference items. The newly generated interference items are used to replace the same number of original incorrect options to generate the first type of evaluation question. If there is at least one original incorrect option with a similarity not less than the second threshold, it is indicated that there is an interference item at present, and all the current incorrect options are retained without generating new interference items.
[0109] Optionally, in the embodiments of the present application, other large models except the multi-modal large model to be evaluated can be selected to generate interference items according to actual needs, and the parameters of the large model used to generate the interference items are strictly isolated from the evaluation environment. For example, the Claude 3.5 sonnet model is selected to generate the interference items of the correct option.
[0110] As an embodiment of the present application, the evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation index specified by a user, specifically including:
[0111] obtaining the index weight corresponding to each evaluation index specified by the user;
[0112] determining the number of questions to be extracted corresponding to each evaluation index based on the index weight of the evaluation index; the number of questions is greater than or equal to a third threshold;
[0113] obtaining a corresponding number of evaluation questions from the evaluation question bank based on the number of questions to be extracted corresponding to each evaluation index;
[0114] mixing and randomly shuffling the obtained evaluation questions corresponding to each evaluation index to determine the order of the multi-modal large model processing the evaluation questions.
[0115] In an embodiment, when the user specifies the evaluation indicators, the user also sets corresponding indicator weights for each evaluation indicator. For example, the unmanned aerial vehicle manufacturer sets multiple evaluation indicators, including visual perception, spatial perception, short-term visual memory, and episodic memory. Considering that visual perception and spatial perception are more concerned during the flight of the unmanned aerial vehicle, the user sets relatively higher indicator weights for the two evaluation indicators of visual perception and spatial perception. Based on the evaluation indicators and corresponding indicator weights specified by the user, the evaluation execution module first determines the number of corresponding questions to be extracted according to the indicator weights. In actual applications, a mapping relationship between the indicator weights and the number of questions can be set according to actual needs. It is worth noting that, in order to ensure the reliability of the evaluation results, in this embodiment, the number of corresponding questions to be extracted determined based on each indicator weight is not less than a third threshold, that is, at least the third threshold of evaluation questions needs to be extracted for each evaluation indicator, so as to ensure that the multi-modal large model processes a sufficient number of questions and ensures the reliability of the evaluation results.
[0116] After obtaining the evaluation indicators and corresponding indicator weights specified by the user, the number of corresponding questions to be extracted is first determined based on each indicator weight. Then, according to each evaluation indicator and the number of corresponding questions, a corresponding number of evaluation questions are obtained from the evaluation question bank in the form of random extraction. The evaluation questions corresponding to each evaluation indicator are mixed and randomly shuffled to obtain a question set of the multi-modal large model to be output. The arrangement order of each evaluation question in the question set is the processing order of the multi-modal large model.
[0117] As an embodiment of the present application, the analysis module is configured to calculate the evaluation score of the multi-modal large model based on the output result and the correct answer of the evaluation question, specifically including:
[0118] determining the score weight corresponding to each evaluation indicator based on the indicator weight corresponding to each evaluation indicator;
[0119] calculating the accuracy of the output result of the multi-modal large model based on the evaluation question corresponding to each secondary indicator, and determining the secondary score corresponding to each secondary indicator based on the accuracy;
[0120] calculating the primary score corresponding to the primary indicator to which each secondary indicator belongs based on each secondary score and the corresponding score weight.
[0121] In an embodiment, the evaluation score is calculated based on the output result of the multi-modal large model and the correct answer of the evaluation question. Specifically, the score weight corresponding to each evaluation index is determined based on the index weight corresponding to the evaluation index. In this embodiment, the score weight corresponding to the index weight is set, so as to expand the score difference between different models in the evaluation index focused by the user, so that the user can more intuitively know the performance of different models in the dimension focused by the user, and facilitate the user to select and optimize the model. Specifically, for the evaluation index with higher index weight, a higher score weight is given. Based on the evaluation question corresponding to each secondary index, the accuracy of the output result of the multi-modal large model is calculated, and the accuracy is taken as the secondary score corresponding to the secondary index. Further, based on the secondary score corresponding to each secondary index and the corresponding score weight, the primary score corresponding to the primary index of each secondary index is calculated. The sum of the score weights corresponding to all secondary indexes belonging to the same primary index is 1.
[0122] Suppose the user specifies 2 secondary indexes under the primary index A: index B and index C, and performs performance evaluation on the multi-modal large model. Based on the output result of the multi-modal large model, the accuracy of index B is m, and the accuracy of index C is n. The primary score of index A is calculated as follows:
[0123] The primary score of index A = m x score weight of index B + n x score weight of index C.
[0124] It is worth noting that if the user only specifies part of the secondary indexes under the primary index for model performance evaluation, that is, there are secondary indexes without secondary scores, for the secondary indexes not specified by the user for evaluation, the secondary score is set to 0.
[0125] By calculating the primary score of each primary index based on different score weights, the performance of different models can be more obviously distinguished in the dimension focused by the user, which facilitates the user to select and optimize the model and improves the user experience.
[0126] In an embodiment, equal score weights (i.e. 0.5) are set for the 2 secondary indexes under each primary index, so as to evaluate the performance of the model in different dimensions (i.e. the dimensions corresponding to each primary index) without dimension emphasis, and calculate the primary score corresponding to each primary index, as follows:
[0127] The primary score of perceptual ability = the secondary score of spatial perception x 0.5 + the secondary score of visual perception x 0.5;
[0128] The primary score of attention = the secondary score of attention concentration x 0.5 + the secondary score of naming recognition x 0.5;
[0129] The first-level score of reasoning ability = the second-level score of analogical reasoning ability x 0.5 + the second-level score of programming ability x 0.5;
[0130] The first-level score of memory ability = the second-level score of short-term visual memory x 0.5 + the second-level score of episodic memory x 0.5.
[0131] As an embodiment of the present application, the analysis module is configured to generate performance evaluation results of the multi-modal large model based on the evaluation scores, specifically comprising:
[0132] determining an evaluation score interval to which the first-level score of each first-level indicator belongs, to obtain evaluation information corresponding to the evaluation score interval, and determining the evaluation information as the performance evaluation result of the first-level score;
[0133] generating performance evaluation results of the multi-modal large model based on the performance evaluation results of the first-level scores.
[0134] In an embodiment, the analysis module generates performance evaluation results of the multi-modal large model in the dimension of each first-level indicator based on the calculated first-level scores of the first-level indicators. In the embodiment of the present application, each first-level indicator has a plurality of evaluation score intervals in the corresponding dimension, and each evaluation score interval corresponds to a performance evaluation result. Optionally, each first-level indicator corresponds to three evaluation score intervals, the first evaluation score interval corresponds to a high performance evaluation result, the second evaluation score interval corresponds to a medium performance evaluation result, and the third evaluation score interval corresponds to a low performance evaluation result. After the first-level score of the multi-modal large model for a first-level indicator is calculated, the first-level score is compared with all evaluation score intervals in the dimension to determine which evaluation score interval the first-level score is in, and the corresponding performance evaluation result is generated based on the determined evaluation score interval. For example, if the first-level score corresponding to the perceptual ability (first-level indicator) is in the first evaluation score interval, the performance evaluation result in the perceptual ability dimension is generated as high.
[0135] Further, the performance evaluation results in the dimensions corresponding to the first-level indicators of all the second-level indicators specified by the user are combined to obtain the performance evaluation results of the multi-modal large model, which include the performance evaluation results in the dimensions corresponding to the first-level indicators of all the second-level indicators specified by the user.
[0136] Optionally, based on the first-level scores corresponding to each first-level indicator and the score weights corresponding to each first-level indicator, the comprehensive score of the multi-modal large model is calculated by weighted summation. Further, based on the comprehensive scores of different models under test, ranking is performed, facilitating the user to select a model with a higher comprehensive score (i.e., better comprehensive performance) to perform a target task, meeting the user's demand. It is worth noting that the sum of the score weights corresponding to all first-level indicators (including: perceptual ability, attention, memory, reasoning ability) is 1. The expression for calculating the comprehensive score is as follows:
[0137] Comprehensive score = first-level score of perceptual ability x score weight + first-level score of attention x score weight + first-level score of memory x score weight + first-level score of reasoning ability x score weight.
[0138] In an embodiment, the multi-modal large model is comprehensively evaluated based on all evaluation indicators, so as to comprehensively evaluate the performance of the model in all dimensions without dimension emphasis. Equal score weights are set for each first-level indicator, i.e., the score weight of perceptual ability is 0.25, the score weight of attention is 0.25, the score weight of memory is 0.25, and the score weight of reasoning ability is 0.25. Based on this, the comprehensive score of the model is calculated, and the expression is as follows:
[0139] Comprehensive score = first-level score of perceptual ability x 0.25 + first-level score of attention x 0.25 + first-level score of memory x 0.25 + first-level score of reasoning ability x 0.25.
[0140] As an embodiment of the present application, the system further comprises an interaction module configured to perform the following steps:
[0141] displaying the selectable first-level indicators and second-level indicators on the user interface and sending the at least one evaluation indicator specified by the user to the evaluation execution module;
[0142] displaying the evaluation questions and the output results of the multi-modal large model on the user interface;
[0143] after the analysis module generates the performance evaluation results, displaying the performance evaluation results and / or the scores corresponding to each evaluation indicator on the user interface.
[0144] In an embodiment, the system further comprises an interaction module for displaying the selectable evaluation indicators (including first-level indicators and second-level indicators) on the user interface. After the user selects the evaluation indicators on the user interface, the interaction module generates an evaluation task based on the multi-modal large model specified by the user and the corresponding evaluation indicators and sends it to the evaluation execution module, so that the evaluation execution model obtains the corresponding evaluation questions from the evaluation question bank based on the evaluation indicators in the evaluation task, and performs performance evaluation on the multi-modal large model.
[0145] In an embodiment of the present application, the interaction module provides a visual interface (i.e., user interface) in a componentized manner through the Vue framework. During the process of processing the test questions by the multi-modal large model, the interaction module displays the test questions and the output results generated by the multi-modal large model in real time on the user interface through the websocket, so that the user can intuitively see the problem-solving process of the multi-modal large model. For example, the interaction module displays the model output results of the current test question and the correct answers on the user interface based on the real-time processing progress of the multi-modal large model, and distinguishes the correct or incorrect answers by different colors, facilitating the user to view.
[0146] After the analysis module generates the performance evaluation results of the multi-modal large model, the interaction module displays the performance evaluation results on the user interface for viewing.
[0147] The analysis module generates multiple types of icons, such as radar charts, column charts, and pie charts, based on the scores of the multi-modal large model for each evaluation index, and displays them on the user interface through the interaction module, so that the user can intuitively understand the distribution of the four-dimensional capabilities (i.e., perceptual ability, attention, memory, and reasoning ability) of the model and the performance comparison of each evaluation index.
[0148] In addition, the interaction module is also used to generate an evaluation report according to the personality evaluation results of the model, which includes the model processing accuracy, time consumption, dimension index scores, and performance evaluation result icons, etc., facilitating the user to view the evaluation results offline and perform subsequent work. The interaction module also provides a ranking list to the user, including a comprehensive ranking list and a dimension-specific ranking list. The comprehensive ranking list shows the user the ranking of the comprehensive capabilities of the mainstream large models currently participating in the evaluation, and the dimension-specific ranking list shows the user the ranking of the dimension-specific capabilities of the mainstream large models currently participating in the evaluation, so as to facilitate the user to select the model for performing the target task and to refer to the subsequent model optimization. For example, the user selects a multi-modal large model with high ranking in the attention dimension from the ranking list based on the short video generation task to be performed, and further optimizes the processing capability of the model in the attention dimension, and uses the optimized model for the short video generation task.
[0149] Figure 2 is the workflow diagram of the multi-modal large model evaluation system in an embodiment of the present application. As shown in Figure 2As shown, the evaluation indicators are determined based on the theory of cognitive psychology and the information processing model, including primary indicators and secondary indicators, forming a hierarchical evaluation framework. Third-party raw data is obtained, and corresponding evaluation questions are constructed in combination with the evaluation indicators, and an evaluation question bank is constructed based on the evaluation questions of each evaluation indicator. The user operation interface is provided through the interactive module (visual interactive terminal), and the user specifies the to-be-evaluated model and the evaluation indicators through the visual interactive terminal. The visual interactive terminal generates an evaluation task and sends it to the evaluation execution module 101. The evaluation execution module 101 determines the number of corresponding questions that need to be extracted according to the evaluation indicators specified by the user, and obtains a corresponding number of evaluation questions from the evaluation question bank in a random extraction manner. In a manner similar to human examination, the evaluation questions are input into the multi-modal large model for processing, and the output results of the multi-modal large model are stored in the Mysql database. The analysis module obtains the output results of the multi-modal large model from the database, calculates the evaluation scores corresponding to each evaluation indicator in combination with the correct answers of each evaluation question, and generates the performance evaluation results of the multi-modal large model based on the evaluation scores. The performance evaluation results are displayed through the visual interactive terminal for the user to view.
[0150] Based on the same inventive concept, an embodiment of the present application provides a multi-modal large model evaluation method based on cognitive psychology. Figure 3 is a flowchart of the multi-modal large model evaluation method based on cognitive psychology according to an embodiment of the present application. As shown in Figure 3 , the method comprises:
[0151] S1: extracting corresponding evaluation questions from the evaluation question bank according to at least one evaluation indicator specified by the user; the evaluation indicators are determined based on the information processing model and the theory of cognitive psychology, and are used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the evaluation questions are constructed based on the evaluation indicators;
[0152] S2: inputting the evaluation questions into the multi-modal large model to be evaluated to obtain output results;
[0153] S3: calculating the evaluation score of the multi-modal large model based on the output results and the correct answers of the evaluation questions;
[0154] S4: generating the performance evaluation results of the multi-modal large model based on the evaluation score.
[0155] As an embodiment of the present application, the method further comprises pre-setting a plurality of evaluation indicators, specifically comprising:
[0156] based on the information processing model and the theory of cognitive psychology, at least one primary indicator is set;
[0157] At least two secondary indexes belonging to the primary index are generated based on each primary index and the Wechsler Intelligence Scale for Children.
[0158] As an embodiment of the present application, the primary indexes include perceptual ability, attention, memory, and reasoning ability; the secondary indexes include visual perception and spatial perception belonging to perceptual ability, attention concentration and naming recognition belonging to attention, short-term visual memory and episodic memory belonging to memory, and analogical reasoning ability and programming ability belonging to reasoning ability; the method further includes:
[0159] The original data of multiple types are obtained, including original multiple-choice questions, original questions and answers, and original fill-in-the-blank questions.
[0160] Based on the secondary indexes, the complete structure of the original multiple-choice questions is retained, and the options of the original multiple-choice questions are rearranged in disorder to obtain first-type evaluation questions.
[0161] Based on the secondary indexes, the stems of the original questions and answers or the original fill-in-the-blank questions are processed to generate multiple-choice questions as second-type evaluation questions.
[0162] Based on each secondary index, the first-type evaluation questions and the second-type evaluation questions are classified and stored to generate an evaluation question bank.
[0163] As an embodiment of the present application, based on the secondary indexes, the stems of the original questions and answers or the original fill-in-the-blank questions are processed to generate multiple-choice questions as second-type evaluation questions, including:
[0164] Based on the stem of the original question and answer or the stem of the original fill-in-the-blank question, one correct option and multiple incorrect options are generated, wherein at least one incorrect option has a similarity greater than or equal to a first threshold value with the correct option.
[0165] According to the order of similarity from high to low with the correct option, multiple incorrect options with high similarity are screened out.
[0166] Based on the correct option and the multiple incorrect options with high similarity screened out, the second-type evaluation questions are generated.
[0167] As an embodiment of the present application, after retaining the complete structure of the original multiple-choice questions and rearranging the options of the original multiple-choice questions in disorder, the method further includes:
[0168] The correct option and the incorrect options of the original multiple-choice questions are obtained.
[0169] The similarity of each incorrect option with the correct option is calculated and compared with a second threshold value.
[0170] In a case where the similarity corresponding to all the incorrect options is less than the second threshold value, generating at least one new incorrect option based on the correct option, the new incorrect option having a similarity greater than or equal to the second threshold value;
[0171] Replacing the incorrect option in the original selection question with the new incorrect option to obtain the first type of test question.
[0172] As an embodiment of the present application, according to at least one test index specified by a user, corresponding test questions are extracted from a test question library, comprising:
[0173] Obtaining the index weight corresponding to each test index specified by the user;
[0174] Based on the index weight of each test index, determining the number of questions to be extracted corresponding to each test index; the number of questions is greater than or equal to a third threshold value;
[0175] Based on the number of questions to be extracted corresponding to each test index, a corresponding number of test questions are obtained from the test question library;
[0176] Mixing and randomly rearranging the test questions corresponding to each test index to determine the order in which the multi-modal large model processes the test questions.
[0177] As an embodiment of the present application, based on the output result and the correct answer of the test question, the test score of the multi-modal large model is calculated, specifically comprising:
[0178] Based on the index weight corresponding to each test index, determining the score weight corresponding to each test index;
[0179] Based on the test question corresponding to each secondary index, the correctness of the output result of the multi-modal large model is calculated, and the corresponding secondary score is determined based on the correctness;
[0180] Based on each secondary score and the corresponding score weight, the primary score corresponding to the primary index to which each secondary index belongs is calculated.
[0181] As an embodiment of the present application, based on the test score, the performance test result of the multi-modal large model is generated, comprising:
[0182] Determining the test score interval to which the primary score of each primary index belongs, to obtain the test information corresponding to the test score interval, and determining the test information as the performance test result of the primary score;
[0183] Based on the performance test result of each primary score, the performance test result of the multi-modal large model is generated.
[0184] As an embodiment of the present application, the method further comprises:
[0185] displaying the selectable primary indicators and secondary indicators on the user interface, and sending the at least one evaluation indicator specified by the user to the evaluation execution module;
[0186] displaying the evaluation questions and the output results of the multi-modal large model on the user interface;
[0187] after the analysis module generates the performance evaluation results, displaying the performance evaluation results and / or the scores corresponding to each evaluation indicator on the user interface.
[0188] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product comprises a computer program which, when executed by a processor, implements the steps in the multi-modal large model evaluation method according to any of the embodiments of the present application.
[0189] Based on the same inventive concept, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the multi-modal large model evaluation method according to any of the embodiments of the present application.
[0190] Based on the same inventive concept, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, which, when executed by the processor, implements the steps in the multi-modal large model evaluation method according to any of the embodiments of the present application.
[0191] As to the method in the above-mentioned embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the system, and will not be described in detail here.
[0192] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0193] For the method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action order described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and components involved are not necessarily necessary for the present application.
[0194] Those skilled in the art will appreciate that embodiments of the application can be supplied as a method, a device, or a computer program product. Thus, embodiments of the application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, embodiments of the application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.
[0195] Embodiments of the application are described herein with reference to the drawings, in which are shown flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the application. It will be understood that each flow and / or block of the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing terminal devices to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal devices, create means for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1
[0196] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1
[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device to cause a series of operational steps to be performed on the computer or other programmable terminal device to produce a computer implemented process such that the instructions which execute on the computer or other programmable terminal device provide steps for implementing the flowcharts and / or block diagrams block or blocks. Figure 1 Figure 1
[0198] While preferred embodiments of the application have been described, those skilled in the art will appreciate that additional modifications and variations to the preferred embodiments are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the application, modifications and variations of the preferred embodiments can be made by those skilled in the art upon the understanding of the teachings of the preferred embodiments, within the spirit and scope of the application as disclosed herein.
[0199] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0200] The above describes in detail the multi-modal large model evaluation system and method based on cognitive psychology provided by the present application. The principles and implementation modes of the present application are described by using specific examples. The above description of the examples is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, the specific implementation modes and application ranges will be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A multi-modal large model evaluation system based on cognitive psychology, characterized in that, Comprise: Theoretical mapping module is configured to set multiple evaluation indexes based on information processing model and cognitive psychology theory, specifically including multiple first-level indexes and multiple second-level indexes belonging to each first-level index; the evaluation dimensions corresponding to the multiple first-level indexes comprehensively cover the key capabilities of the multi-modal large language model; The question bank construction module is configured to obtain multiple types of original data, including: original selection questions, original question and answer questions, and original fill-in-the-blank questions; based on the second-level indexes, the complete structure of the original selection questions is retained, and the options of the original selection questions are rearranged in disorder, obtaining the first type of evaluation questions; based on the second-level indexes, the stems of the original question and answer questions or the original fill-in-the-blank questions are processed to generate selection questions as the second type of evaluation questions; the second type of evaluation questions at least contain one incorrect option with a similarity to the correct option not less than a first threshold; based on each second-level index, the first type of evaluation questions and the second type of evaluation questions are classified and stored to generate an evaluation question bank; The evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation index specified by a user; the evaluation index is determined based on the information processing model and the cognitive psychology theory, and is used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the evaluation questions are constructed based on the evaluation index; the evaluation questions are input into the multi-modal large model to be evaluated to obtain an output result; The analysis module is configured to calculate the evaluation score of the multi-modal large model based on the output result and the correct answer of the evaluation question; and generate a performance evaluation result of the multi-modal large model based on the evaluation score.
2. The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The multiple first-level indexes include perceptual ability, attention, memory, and reasoning ability; the multiple second-level indexes include visual perception and spatial perception belonging to perceptual ability, attention concentration and naming recognition belonging to attention, short-term visual memory and episodic memory belonging to memory, and analogical reasoning ability and programming ability belonging to reasoning ability. 3.The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The question bank construction module is configured to process the stems of the original question and answer questions or the original fill-in-the-blank questions based on the second-level indexes to generate selection questions as the second type of evaluation questions, specifically including: Based on the stems of the original question and answer questions or the stems of the original fill-in-the-blank questions, generate one correct option and multiple incorrect options, at least one of which has a similarity to the correct option greater than or equal to a first threshold; According to the order of similarity from high to low, filter out multiple incorrect options with high similarity; Based on the correct option and the filtered multiple incorrect options with high similarity, generate the second type of evaluation questions.
4. The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The question bank construction module is further configured to, after retaining the complete structure of the original selection questions and rearranging the options of the original selection questions in disorder, perform the following steps: Obtain the correct option and the incorrect option of the original selection question; compute a similarity of each incorrect option to the correct option and compare the similarity to a second threshold value; in a case where the similarity corresponding to each incorrect option is less than the second threshold value, generate at least one new incorrect option having a similarity greater than or equal to the second threshold value based on the correct option; replace the incorrect options in the original multiple-choice question with the new incorrect options to obtain the first type of test question.
5. The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The test execution module is configured to extract corresponding test questions from a test question library according to at least one test index specified by a user, and specifically includes: obtaining an index weight corresponding to each test index specified by the user; determining a number of questions to be extracted corresponding to each test index based on the index weight of the test index; the number of questions is greater than or equal to a third threshold value; obtaining a corresponding number of test questions from the test question library based on the number of questions to be extracted corresponding to each test index; mixing and randomly rearranging the test questions corresponding to each test index to determine an order in which the multi-modal large model processes the test questions.
6. The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The analysis module is configured to calculate a test score of the multi-modal large model based on the output result and a correct answer of the test question, and specifically includes: determining a score weight corresponding to each test index based on the index weight corresponding to each test index; calculating a correctness rate of the output result of the multi-modal large model based on the test questions corresponding to each secondary index, and determining a secondary score corresponding to each secondary index based on the correctness rate; calculating a primary score corresponding to a primary index to which each secondary index belongs based on each secondary score and the score weight corresponding to the secondary score.
7. The cognitive psychology based multi-modal large model evaluation system according to claim 6, wherein, The analysis module is configured to generate a performance evaluation result of the multi-modal large model based on the test score, and specifically includes: determining a test score interval to which each primary score belongs to obtain evaluation information corresponding to the test score interval, and determining the evaluation information as the performance evaluation result of the primary score; generating the performance evaluation result of the multi-modal large model based on the performance evaluation result of each primary score.
8. The cognitive psychology based multi-modal large model evaluation system according to claim 1, wherein, The system further includes an interaction module configured to perform the following steps: displaying selectable primary indexes and secondary indexes on a user interface, and sending at least one test index specified by a user to the test execution module; displaying test questions and output results of the multi-modal large model on a user interface; after the analysis module generates a performance evaluation result, displaying the performance evaluation result and / or a score corresponding to each test index on a user interface.
9. A method for evaluating a multi-modal large model based on cognitive psychology, characterized in that, application to the system of any one of claims 1-8, comprising: extracting corresponding test questions from a test question library according to at least one test index specified by a user; the test index is determined based on an information processing model and cognitive psychology theory, and is used to evaluate the performance of the multi-modal large model in at least one of the following dimensions: perceptual ability, attention, memory, reasoning ability; the test question is constructed based on the test index; inputting the test question into the multi-modal large model to be evaluated to obtain an output result; Based on the output result and the correct answer of the test question, a test score of the multi-modal large model is calculated; Based on the test score, a performance test result of the multi-modal large model is generated.
10. The cognitive psychology-based multi-modal large model evaluation method according to claim 9, characterized in that, Further comprising pre-setting a plurality of test indicators, specifically including: Based on an information processing model and a cognitive psychology theory, at least one primary indicator is set; Based on each primary indicator and a Wechsler Intelligence Scale for Children, at least two secondary indicators belonging to the primary indicator are generated.
Citation Information
Patent Citations
Assessment method and device of large language model, storage medium and computer equipment
CN117291184A
Model-based question and answer test method and system, medium and product
CN120561256A