Multi-modal large model evaluation system and method based on cognitive psychology
By constructing a multi-modal large model evaluation system based on cognitive psychology, multi-dimensional evaluation indicators and question banks are built, which solves the problem of single-dimensional evaluation in traditional evaluation schemes, realizes comprehensive and accurate evaluation of multi-modal large models, and improves the model's task processing performance and user experience.
Patent Information
- Application Number
- CN202511476246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-16
AI Technical Summary
Existing multimodal large model evaluation schemes mainly focus on single-dimensional performance evaluation, lacking consideration of the model's comprehensive multi-dimensional capabilities. This results in an inability to fully measure the model's actual capabilities and application potential, and the task-oriented evaluation results cannot accurately reflect the effectiveness of the model's cognitive mechanisms.
The multimodal large model evaluation system based on cognitive psychology determines multiple evaluation indicators through information processing models and cognitive psychology theories, constructs an evaluation question bank, including evaluation questions for dimensions such as perception, attention, memory, and reasoning, generates multiple-choice questions, and conducts evaluations in conjunction with human cognitive processing processes, thereby achieving objective and comprehensive performance evaluation of the multimodal large model.
It enables multi-dimensional performance evaluation of large multimodal models, accurately reflects the model's true capabilities, helps users clarify optimization directions, and improves task completion quality and user experience.
Smart Images

Figure CN120929795A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a multimodal large model evaluation system and method based on cognitive psychology. Background Technology
[0002] In recent years, with the continuous innovation of artificial intelligence technology, multimodal large models have broken through the limitations of traditional single-modal models (such as text-only or image-only models). They can simultaneously understand, associate, and generate information of multiple modalities (such as images, videos, audio, and text), providing richer and more accurate information representation, and have attracted widespread attention. Faced with complex and highly integrated multimodal large models, how to effectively evaluate their overall performance so that users can debug the model for specific tasks to improve its processing performance has become a problem that needs to be solved.
[0003] However, current performance evaluation schemes for multimodal large models mostly focus on performance evaluation in only one dimension. For example, they only focus on the model's performance in language understanding or image recognition, lacking consideration of the model's comprehensive capabilities across multiple dimensions. This makes it impossible to comprehensively measure the model's actual capabilities and application potential. Furthermore, current evaluation schemes often rely on task-oriented benchmark tests, which can only evaluate the model's performance for a specific task. That is, the evaluation results reflect the model's knowledge reserves in a specific domain, rather than objective and accurate true performance, which can mislead users, affect subsequent model fine-tuning and application, slow down user task progress, and negatively impact user experience. Summary of the Invention
[0004] In view of this, this application aims to propose a multimodal large model evaluation system and method based on cognitive psychology, so as to achieve objective and comprehensive performance evaluation of multimodal large models and accurately reflect the true performance of multimodal large models.
[0005] To achieve the above objectives, the technical solution of this application is as follows: The first aspect of this application provides a multimodal large model evaluation system based on cognitive psychology, the system comprising: The evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank based on at least one evaluation indicator specified by the user; the evaluation indicator is determined based on the information processing model and cognitive psychology theory, and is used to evaluate the performance of the multimodal large model in at least one of the following dimensions: perception, attention, memory, and reasoning; the evaluation questions are constructed based on the evaluation indicator; the evaluation questions are input into the multimodal large model to be evaluated, and the output results are obtained; The analysis module is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions; and to generate the performance evaluation results of the multimodal large model based on the evaluation score.
[0006] Optionally, the system further includes: a theory mapping module and a question bank construction module; The theoretical mapping module is configured to set multiple evaluation indicators based on information processing models and cognitive psychology theories. Specifically, it includes multiple primary indicators and multiple secondary indicators belonging to each primary indicator. The multiple primary indicators include: perception, attention, memory, and reasoning ability. The multiple secondary indicators include: visual perception and spatial perception belonging to perception; attention concentration and name recognition belonging to attention; short-term visual memory and episodic memory belonging to memory; and analogical reasoning ability and programming ability belonging to reasoning ability. The question bank construction module is configured to perform the following steps: Obtain various types of raw data, including: raw multiple-choice questions, raw short-answer questions, and raw fill-in-the-blank questions; Based on the aforementioned secondary indicators, the complete structure of the original multiple-choice questions is preserved, and the options of the original multiple-choice questions are rearranged in a random order to obtain the first type of evaluation questions; Based on the aforementioned secondary indicators, the stems of the original question-and-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions, which serve as the second type of evaluation questions. Based on the various secondary indicators, the first type of evaluation questions and the second type of evaluation questions are classified and stored to generate an evaluation question bank.
[0007] Optionally, the question bank construction module is configured to process the stems of the original open-ended questions or the original fill-in-the-blank questions based on the secondary indicators to generate multiple-choice questions as the second type of assessment questions, specifically including: Based on the stem of the original question-and-answer question or the stem of the original fill-in-the-blank question, generate one correct option and multiple incorrect options, including at least one incorrect option whose similarity to the correct option is greater than or equal to a first threshold. Based on the order of similarity to the correct option from high to low, filter out multiple incorrect options with high similarity. Based on the correct options and several highly similar incorrect options selected, the second type of evaluation questions are generated.
[0008] Optionally, the question bank construction module is further configured to perform the following steps after preserving the complete structure of the original multiple-choice questions and rearranging the options of the original multiple-choice questions in a shuffled order: Obtain the correct and incorrect options for the original multiple-choice questions; Calculate the similarity between each incorrect option and the correct option, and compare it with a second threshold; If the similarity of all incorrect options is less than the second threshold, generate at least one new incorrect option with a similarity greater than or equal to the second threshold based on the correct option; The incorrect options in the original multiple-choice questions are replaced with the new incorrect options to obtain the first type of assessment questions.
[0009] Optionally, the evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation indicator specified by the user, specifically including: Obtain the weights of each evaluation metric specified by the user. Based on the weights of each evaluation indicator, the number of questions to be extracted is determined; the number of questions is greater than or equal to the third threshold. Based on the number of questions to be extracted for each evaluation indicator, a corresponding number of evaluation questions are obtained from the evaluation question bank. The evaluation questions corresponding to each evaluation indicator are mixed and randomly sorted to determine the order in which the multimodal large model processes the evaluation questions.
[0010] Optionally, the analysis module is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions, specifically including: Based on the weight of each evaluation indicator, the score weight of each evaluation indicator is determined. Based on the evaluation questions corresponding to each secondary indicator, the accuracy of the output results of the multimodal large model is calculated, and the corresponding secondary score is determined based on the accuracy. Based on each secondary score and its corresponding score weight, the primary score corresponding to the primary indicator to which each secondary indicator belongs is calculated.
[0011] Optionally, the analysis module is configured to generate performance evaluation results for the multimodal large model based on the evaluation scores, specifically including: Determine the evaluation score range to which the first-level score of each first-level indicator belongs, obtain the evaluation information corresponding to the evaluation score range, and determine the evaluation information as the performance evaluation result of the first-level score; Based on the performance evaluation results of each first-level score, the performance evaluation results of the multimodal large model are generated.
[0012] Optionally, the system further includes an interaction module configured to perform the following steps: The selectable primary and secondary indicators are displayed on the user interface, and at least one evaluation indicator specified by the user is sent to the evaluation execution module. The evaluation questions and the output results of the multimodal large model are displayed on the user interface; After the analysis module generates the performance evaluation results, the performance evaluation results and / or the scores corresponding to each evaluation indicator are displayed on the user interface.
[0013] According to a second aspect of the embodiments of this application, a multimodal large model evaluation method based on cognitive psychology is provided, applied to the system provided in the first aspect of the embodiments of this application, the method comprising: Based on at least one evaluation metric specified by the user, corresponding evaluation questions are extracted from the evaluation question bank; the evaluation metric is determined based on the information processing model and cognitive psychology theory, and is used to evaluate the performance of at least one of the following dimensions of the multimodal large model: perception, attention, memory, and reasoning; the evaluation questions are constructed based on the evaluation metric. Input the evaluation questions into the multimodal large model to be evaluated, and obtain the output results; Based on the output results and the correct answers to the evaluation questions, the evaluation score of the multimodal large model is calculated; The performance evaluation results of the multimodal large model are generated based on the evaluation scores.
[0014] Optionally, the method further includes pre-setting multiple evaluation indicators, specifically including: Based on information processing models and cognitive psychology theories, at least one primary indicator is set. Based on each primary indicator and the Wechsler Intelligence Scale for Children, at least two secondary indicators belonging to the primary indicators are generated.
[0015] The multimodal large model evaluation system based on cognitive psychology provided in this application pre-determines multiple evaluation indicators based on information processing models and cognitive psychology theories, allowing users to select them for performance evaluation of the multimodal large model. Based on these indicators, a corresponding evaluation question bank is constructed. After the user specifies the evaluation indicators for the multimodal large model, the evaluation execution module extracts the corresponding evaluation questions from the question bank based on the user-selected indicators and inputs them into the multimodal large model to be evaluated for processing, obtaining the output results of the multimodal large model for the evaluation questions. Then, the analysis module calculates the corresponding evaluation score based on the correct answers to the evaluation questions and the output results of the multimodal large model. The corresponding performance evaluation results are generated based on the evaluation scores of the multimodal large model.
[0016] This application, based on cognitive psychology theory and the analogy between human cognition and computer information processing in information processing models, evaluates the data processing performance of multimodal large-scale models according to the human cognitive processing process. Specifically, it constructs evaluation indicators and a question bank based on cognitive psychology theory and information processing models for model performance evaluation, thereby establishing the relationship between the model's deep capabilities and human cognitive dimensions. Compared to traditional model performance evaluation schemes that only target specific domain knowledge levels or single dimensions, this application achieves an objective, comprehensive, and accurate evaluation of the cognitive mechanisms and capabilities of multimodal large-scale models. This allows users to fully understand the true performance of the evaluated model in all aspects, facilitating targeted fine-tuning and optimization of the model according to the target task (e.g., text-image question-answering task), thereby improving the model's performance for the target task, enhancing task completion quality, and improving user experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a multimodal large model evaluation system based on cognitive psychology proposed in an embodiment of this application; Figure 2 This is a flowchart of the multimodal large model evaluation system in one embodiment of this application; Figure 3 This is a flowchart of a multimodal large model evaluation method based on cognitive psychology proposed in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0021] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects as detailed in this application.
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0024] Multimodal large models are deep learning models with massive numbers of parameters (typically billions or even trillions). These models automatically learn general patterns and knowledge of language, images, and other modes from unlabeled data through "self-supervised learning." Their capabilities are not only reflected in the scale of parameters; when the model size exceeds a critical point, they exhibit complex capabilities not explicitly programmed into the training data, such as context learning, complex reasoning, and code generation. However, traditional evaluation schemes for multimodal large models have the following problems: (1) Evaluation focuses only on a single dimension, such as the model's performance in language understanding or image recognition, lacking consideration of the model's multi-dimensional comprehensive capabilities. Evaluation benchmarks, such as the widely used VQA (Visual Question Answering) and GQA (Graph-based Question Answering), measure the model's accuracy in basic perception and simple reasoning by pre-setting a single image and manually labeled questions. They also pre-set task datasets (e.g., answering questions based on images, image description generation, multiple-choice knowledge questions) and calculate the model's accuracy, recall, and other metrics on these datasets to measure model performance. This approach can only evaluate performance on a single task in a static scene and cannot determine the model's multi-dimensional processing performance. For example, the COCO dataset focuses on object localization accuracy, and the ImageNet dataset focuses on classification ability; neither addresses the model's ability to handle cross-modal correlations, long contextual dependencies, or dynamic sequences.
[0025] (2) Traditional evaluation schemes rely on task-oriented, knowledge-based tests, neglecting the cognitive mechanisms and capabilities of the model. This results in evaluation results that reflect domain knowledge reserves rather than cognitive mechanism effectiveness. For example, a model that performs well in medical image diagnosis may achieve high scores due to knowledge bias rather than cognitive advantage. This makes it difficult for developers to determine whether the model's problems are perceptual errors (such as misreading image features), inattention (such as ignoring key areas), or reasoning errors (such as misinterpreting pathological associations). Consequently, developers cannot effectively fine-tune the model based on the evaluation results, affecting task processing efficiency and user experience. For example, long text responses actually reflect the attention capabilities of large models. However, traditional evaluation schemes can only evaluate the model's knowledge-based response capabilities at the knowledge level, and the evaluation results cannot accurately reflect the model's attention capabilities.
[0026] To address the shortcomings of traditional evaluation schemes, such as limited evaluation dimensions and neglect of model cognitive process decomposition, this application customizes multi-dimensional evaluation objectives based on information processing models and cognitive psychology theories. This aims to establish the relationship between the model's internal capabilities and human cognitive dimensions, and constructs an evaluation question bank for performance evaluation of multimodal large models. By analyzing the model's processing results of the evaluation questions, the application achieves objective and accurate evaluation of the model's performance in various dimensions, revealing the synergy of large models at the human cognitive level, and enabling users to accurately understand the human-like intelligence level of the tested model.
[0027] The present application will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] Figure 1 This is a schematic diagram of a multimodal large model evaluation system 100 based on cognitive psychology, as proposed in an embodiment of this application. Figure 1 As shown, the system includes: The evaluation execution module 101 is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation indicator specified by the user; the evaluation indicator is determined based on the information processing model and cognitive psychology theory, and is used to evaluate the performance of the multimodal large model in at least one of the following dimensions: perception, attention, memory, and reasoning; the evaluation questions are constructed based on the evaluation indicator; the evaluation questions are input into the multimodal large model to be evaluated, and the output results are obtained; Analysis module 102 is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions; and generate the performance evaluation results of the multimodal large model based on the evaluation score.
[0029] In this embodiment, multiple evaluation indicators are pre-determined based on information processing models and cognitive psychology theories. These indicators are used to measure the performance of the multimodal large model in different dimensions, specifically including: perception, attention, memory, and reasoning. Corresponding evaluation items are constructed based on each evaluation indicator, and an evaluation item bank is generated.
[0030] When evaluating multimodal large models, users select the model to be evaluated and the corresponding evaluation metrics. In practical applications, the multimodal large model to be evaluated can be a model with a mature architecture such as GPT-4o-mini, Qwen-2.5-vl, Gemini-2-flash, GLM-4v-plus, or Doubao-1.5, or it can be a custom model or a model obtained by fine-tuning an existing multimodal large model.
[0031] The evaluation execution module extracts corresponding evaluation questions from the evaluation question bank according to the user-specified evaluation metrics, forming a target question set. Each evaluation question in the target question set is preprocessed, converting it into a format recognizable by the multimodal large-scale model to be evaluated, and inputting it into the multimodal large-scale model for processing. Simultaneously, a timer is started to record the response time, capturing the model output to obtain the output result of the multimodal large-scale model. Optionally, in this embodiment, the constructed evaluation questions are all multiple-choice questions with multiple options. Before the multimodal large-scale model processes the evaluation questions, the output format of the model is constrained based on the number of options in the evaluation questions, forcing the model to return one of the options. Optionally, this system uses the Flask framework's API (Application Programming Interface) to provide model evaluation services to users. The evaluation execution module stores the output result of the multimodal large-scale model in a MySQL database, facilitating subsequent multi-dimensional performance analysis of the model.
[0032] The analysis module calculates the evaluation scores of the multimodal large model for each evaluation metric based on the output of the multimodal large model and the correct answers to each evaluation question in the target question set. It then generates corresponding performance evaluation results based on these scores. Optionally, during the multimodal large model's processing of evaluation questions, the analysis module uses SQLAlchemy to write the model's output for each evaluation question into the database via ORM (Object Relational Mapping). This facilitates comparison and analysis between the model's output and the correct answers to the evaluation questions. Based on the performance evaluation results corresponding to the evaluation metrics, users can accurately understand the model's true processing capabilities in different dimensions. Furthermore, based on the target task to be executed, users can perform targeted fine-tuning of the model to improve its processing performance, thereby enhancing the execution efficiency and completion quality of the target task.
[0033] In this embodiment, specific theories of human cognitive psychology and information processing models are applied to the multidimensional evaluation of the capabilities of multimodal large models. Based on the processing method of human cognitive processing, the learning process of multimodal large models is analogous, thereby objectively, comprehensively and accurately evaluating the cognitive mechanism and real capabilities of the multimodal large models themselves. This makes it easier for users to clarify the direction of subsequent optimization of the model, thereby improving the quality of task completion and enhancing the user experience.
[0034] This system can be used throughout the entire lifecycle of a multimodal large model. Based on the performance evaluation results of the multimodal large model generated by the system, users can perform targeted fine-tuning of the model before, during, and after the execution of the target task. For example, before using a multimodal large model to complete a drone flight mission, drone manufacturers can specify corresponding evaluation indicators (such as spatial perception capability and visual perception capability) based on the necessary conditions for drone flight. Then, they can use the system to obtain the model's performance based on the evaluation indicators, clarify the direction for subsequent targeted fine-tuning of the model, and thus improve the quality of drone flight mission completion.
[0035] As one embodiment of this application, the system further includes: a theory mapping module and a question bank construction module; The theoretical mapping module is configured to set multiple evaluation indicators based on information processing models and cognitive psychology theories. Specifically, it includes multiple primary indicators and multiple secondary indicators belonging to each primary indicator. The multiple primary indicators include: perception, attention, memory, and reasoning ability. The multiple secondary indicators include: visual perception and spatial perception belonging to perception; attention concentration and name recognition belonging to attention; short-term visual memory and episodic memory belonging to memory; and analogical reasoning ability and programming ability belonging to reasoning ability. The question bank construction module is configured to perform the following steps: Obtain various types of raw data, including: raw multiple-choice questions, raw short-answer questions, and raw fill-in-the-blank questions; Based on the aforementioned secondary indicators, the complete structure of the original multiple-choice questions is preserved, and the options of the original multiple-choice questions are rearranged in a random order to obtain the first type of evaluation questions; Based on the aforementioned secondary indicators, the stems of the original question-and-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions, which serve as the second type of evaluation questions. Based on the various secondary indicators, the first type of evaluation questions and the second type of evaluation questions are classified and stored to generate an evaluation question bank.
[0036] In this embodiment, the system pre-constructs a mapping relationship between human cognitive dimensions and the core capabilities of a multimodal large model based on information processing models and cognitive psychology theories through a theoretical mapping module. This transforms the ideas and methods of a specific psychological testing paradigm and applies them to the ability assessment of the cognitive dimensions of a multimodal large model.
[0037] Specifically, a mapping relationship is established based on the core cognitive dimensions of the PASS cognitive model and the Neisser information processing model. Human cognition is analogized to a computer's information processing system. Combining the core dimensions of human cognitive processing—"planning," "attention," "simultaneous processing," and "sequential processing"—and the multiple discrete stages of data processing (including: environmental perception of stimuli, initial interpretation, selective focus of attention to filter irrelevant information, entry into the memory system, and finally, complex operations using input information and long-term memory knowledge in the thinking, reasoning, and decision-making stages), several primary indicators for measuring the multi-dimensional capabilities of a multimodal large model are identified. These include: perception, attention, memory, and reasoning. Specifically, the perception indicator corresponds to the model's visual / auditory feature extraction ability; the attention indicator corresponds to the model's information filtering and focusing ability; the memory indicator corresponds to the model's contextual understanding and sequence maintenance ability; and the reasoning indicator corresponds to the model's cross-modal logical association ability. By setting multiple primary indicators, the system's evaluation dimensions comprehensively cover the model's key capabilities.
[0038] In this embodiment, based on the primary indicators, the Wechsler Intelligence Scale for Children (WISC) is used to further subdivide each primary indicator, identifying multiple secondary indicators and thus forming a hierarchical evaluation framework. These secondary indicators enable a deeper breakdown of the model's key capabilities across various dimensions, improving the accuracy of model performance evaluation. The evaluation results reveal more detailed performance differences across key capability dimensions, allowing users to clearly define the specific direction for model optimization and improving the efficiency of subsequent model training.
[0039] Specifically, the perception index involves the model's ability to analyze and understand input information, including the processing of multimodal information such as speech, images, and text. In this embodiment, based on the perception index combined with the spatial intelligence measurement paradigm of human cognitive block design and visual jigsaw puzzles, visual perception and spatial perception are further divided as subordinate secondary indicators. The visual perception index quantifies the model's ability to extract low-level visual features through fine-grained image recognition, simulating the encoding process of shape and texture by the human retinal-cortical pathway; the spatial perception index evaluates the model's reasoning ability regarding relative position and spatial topology through 3D scene question answering.
[0040] Attention metrics aim to simulate the selective attention process in the human brain during information processing. In this embodiment, based on the attention metrics combined with the human cognitive coding subtest to test the efficiency of "visual-symbol" conversion, attention concentration and naming recognition are further subdivided as secondary metrics to simulate the human selective attention mechanism.
[0041] In this embodiment, based on memory metrics combined with the reproduction test of "digit-space" sequences in the memory span test of human cognition, short-term visual memory and contextual memory are further divided as subordinate secondary metrics. The short-term visual memory metric quantifies the model's ability to retain short-term visual information through multi-round image recognition and analysis tasks; the contextual memory metric uses contextual long-context and multi-round dialogue tasks to evaluate the model's ability to model contextual dependencies.
[0042] In this embodiment, a standardized evaluation paradigm for non-verbal logical reasoning is established based on reasoning ability indicators combined with matrix reasoning tests of human cognition. This paradigm further divides analogical reasoning ability and programming ability into secondary indicators. The analogical reasoning ability indicator uses visual matrix completion to measure the model's ability to abstract and map implicit relationships; the programming ability indicator evaluates the model's ability to decompose and execute structured rules through algorithms and programming tasks.
[0043] Traditional multimodal large model evaluation schemes often suffer from overly simplified evaluation questions, making it difficult to effectively assess model performance. For example, the Multimodal Model Evaluation Benchmark (MME) constructs a binary judgment question bank based on the "perception-cognition" dimension to evaluate object recognition ability and basic logical reasoning. However, binary choices (yes / no) result in a random guess accuracy rate as high as 50%, making it difficult to accurately reflect the model's true understanding level in practical applications. Therefore, this embodiment constructs multiple-choice questions based on various secondary indicators as evaluation questions, thereby avoiding the influence of random guesses on the accuracy of the evaluation results.
[0044] Specifically, based on each secondary indicator, applicable datasets are obtained from different sources as raw data. Evaluation questions corresponding to each evaluation indicator are then constructed based on this raw data. In practical applications, publicly available datasets can be collected from multiple third-party sources as raw data, such as Google, Huggingface, and Kaggle, depending on actual needs. The obtained raw data includes different types of raw questions: multiple choice, open-ended questions, and fill-in-the-blank questions. To construct multiple-choice questions that match the evaluation indicators, adjustments need to be made to the different types of raw questions, as follows: (1) For the original multiple-choice questions, their complete structure is preserved to ensure the integrity of the original context of the data and the task setting. The original options are shuffled and rearranged to change the position of the correct answer and generate the first type of evaluation questions. By shuffling and rearranging the original options, data leakage (i.e., the model has learned the data in advance and therefore answers directly based on memory without thinking and reasoning) is avoided, thereby improving the accuracy of the evaluation results; (2) For original question-and-answer questions or original fill-in-the-blank questions, only the original data is used, without relying directly on the original questions or annotations. The second type of evaluation questions are generated by converting the question stem of the original questions into multiple-choice questions with multiple options that match the evaluation indicators, so as to maintain the same format as the first type of evaluation questions.
[0045] The first and second categories of evaluation questions are merged and stored according to different secondary indicators to form a standardized and normalized evaluation question bank, which facilitates quick querying and retrieval of evaluation questions corresponding to the evaluation indicators when conducting model evaluation.
[0046] In this embodiment, based on multiple primary indicators measuring key capabilities, several secondary indicators are further subdivided to achieve a deep breakdown of the model's key capabilities across various dimensions, forming a hierarchical evaluation framework and improving the accuracy of model performance evaluation. On this basis, the collected raw data is processed based on each secondary indicator to construct a standardized, leak-resistant evaluation question bank. Through multiple-choice questions matching the evaluation indicators, an objective and comprehensive quantitative evaluation of the multimodal large model is achieved, accurately reflecting the model's deep human-like cognitive abilities.
[0047] Optionally, when evaluating a multimodal large model through this system, users can specify primary or secondary indicators as needed. If the user specifies a primary indicator, the evaluation execution module determines all secondary indicators under that primary indicator and then extracts corresponding evaluation questions from the evaluation question bank based on each secondary indicator. If the user specifies a secondary indicator, the evaluation execution module directly extracts evaluation questions from the evaluation question bank based on that secondary indicator.
[0048] As one implementation of this application, the question bank construction module is configured to process the stems of the original question-and-answer questions or the original fill-in-the-blank questions based on the secondary indicators to generate multiple-choice questions as the second type of assessment questions, specifically including: Based on the stem of the original question-and-answer question or the stem of the original fill-in-the-blank question, generate one correct option and multiple incorrect options, including at least one incorrect option whose similarity to the correct option is greater than or equal to a first threshold. Based on the order of similarity to the correct option from high to low, filter out multiple incorrect options with high similarity. Based on the correct options and several highly similar incorrect options selected, the second type of evaluation questions are generated.
[0049] In one embodiment, the question bank construction module constructs multiple-choice questions matching the evaluation metrics as a second type of evaluation questions based on the stems of the original question-and-answer questions or the original fill-in-the-blank questions. Specifically, based on the stems of the original question-and-answer questions or the original fill-in-the-blank questions, multiple options for the multiple-choice questions are generated, including one correct option and multiple incorrect options. During option generation, one or more distractors with high similarity to the correct option are generated, thereby increasing the difficulty of the evaluation questions, reducing guessing during the multimodal large-scale model's problem-solving process, and tapping into the model's deeper understanding, analytical, comparative, and logical judgment abilities, thus improving the accuracy of the test results.
[0050] Specifically, based on the original question stem of the question / fill-in-the-blank question, one correct answer and multiple incorrect answers are generated. At least one incorrect answer must have a similarity to the correct answer that is not less than a first threshold, meaning at least one distractor is generated, and each distractor has a high similarity to the correct answer. All generated incorrect answers are sorted from highest to lowest similarity to the correct answer, and based on the number of options required for the evaluation questions, multiple highly similar incorrect answers are selected. For example, if the evaluation question is a four-option multiple-choice question, in addition to the correct answer, three highly similar incorrect answers need to be selected. Based on the correct answer and the selected incorrect answers, a second type of evaluation question is constructed.
[0051] Compared to traditional evaluation schemes for large multimodal models, this scheme generates distractors that are highly similar to the correct options when constructing the evaluation questions, thereby increasing the complexity of the evaluation questions and enhancing the data leakage resistance of the evaluation question bank. Furthermore, by generating a second type of evaluation question in multiple-choice format based on the original non-multiple-choice questions (including original fill-in-the-blank and original open-ended questions), the sample size of the evaluation questions in the evaluation question bank can be effectively expanded. This allows for a thorough evaluation of the model's performance using a large number of evaluation questions, improving the accuracy of the evaluation results.
[0052] As one embodiment of this application, the question bank construction module is further configured to perform the following steps after preserving the complete structure of the original multiple-choice questions and rearranging the options of the original multiple-choice questions in a shuffled order: Obtain the correct and incorrect options for the original multiple-choice questions; Calculate the similarity between each incorrect option and the correct option, and compare it with a second threshold; If the similarity of all incorrect options is less than the second threshold, generate at least one new incorrect option with a similarity greater than or equal to the second threshold based on the correct option; The incorrect options in the original multiple-choice questions are replaced with the new incorrect options to obtain the first type of assessment questions.
[0053] In one embodiment, after shuffling and rearranging the options of the original multiple-choice question, highly similar distractors are generated based on the correct options of the original question. These distractors are then used to replace the other options besides the correct ones, thus preventing the model from arriving at the answer without analysis due to leakage of the original multiple-choice question, and enhancing the resilience of the evaluation question bank against data leakage. Furthermore, by introducing distractors of the correct options into the original multiple-choice question, the complexity of the original question can be increased, thereby enabling the model to uncover its deeper cognitive process decomposition capabilities during evaluation, rather than merely remaining at the knowledge level.
[0054] Specifically, the correct and incorrect options from the original multiple-choice questions are obtained, and the similarity between each incorrect option and the correct option is calculated. This similarity is then compared to a second threshold. If no incorrect option has a similarity greater than or equal to the second threshold, it is determined that there are no distractors. One or more new incorrect options with a similarity not less than the second threshold are then generated based on the correct option as distractors. The newly generated distractors replace the same number of original incorrect options, generating the first type of evaluation questions. If at least one original incorrect option has a similarity not less than the second threshold with the correct option, it indicates that distractors exist. In this case, all current incorrect options are retained, and no new distractors need to be generated.
[0055] Optionally, in this embodiment, other large models besides the multimodal large model to be evaluated can be selected to generate interference items according to actual needs, and the parameters of the large model used to generate interference items are strictly isolated from the evaluation environment. For example, the Claude 3.5 Sonnet model can be selected to generate interference items for the correct options.
[0056] As one embodiment of this application, the evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation indicator specified by the user, specifically including: Obtain the weights of each evaluation metric specified by the user. Based on the weights of each evaluation indicator, the number of questions to be extracted is determined; the number of questions is greater than or equal to the third threshold. Based on the number of questions to be extracted for each evaluation indicator, a corresponding number of evaluation questions are obtained from the evaluation question bank. The evaluation questions corresponding to each evaluation indicator are mixed and randomly sorted to determine the order in which the multimodal large model processes the evaluation questions.
[0057] In one embodiment, when specifying evaluation metrics, the user also sets corresponding metric weights for each metric. For example, a drone manufacturer sets multiple evaluation metrics, including: visual perception, spatial perception, short-term visual memory, and contextual memory. Considering that visual perception and spatial perception are more important during drone flight, relatively higher metric weights are assigned to these two metrics. Based on the user-specified evaluation metrics and their corresponding weights, the evaluation execution module first determines the number of questions to be extracted according to the metric weights. In practical applications, the mapping relationship between metric weights and the number of questions can be set according to actual needs. It is worth noting that, to ensure the reliability of the evaluation results, in this embodiment, the number of questions to be extracted based on each metric weight is not less than a third threshold. That is, for each evaluation metric, at least the third threshold of evaluation questions needs to be extracted to ensure that the multimodal large model processes a sufficient number of questions and ensures the reliability of the evaluation results.
[0058] After obtaining the user-specified evaluation metrics and their corresponding weights, the number of questions to be extracted is first determined based on each metric weight. Then, according to each evaluation metric and its corresponding number of questions, a corresponding number of evaluation questions are randomly selected from the evaluation question bank. The evaluation questions corresponding to each evaluation metric are mixed and randomly shuffled to obtain the question set for the output multimodal large model. The order of the evaluation questions in this question set is the processing order of the multimodal large model.
[0059] In one embodiment of this application, the analysis module is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions, specifically including: Based on the weight of each evaluation indicator, the score weight of each evaluation indicator is determined. Based on the evaluation questions corresponding to each secondary indicator, the accuracy of the output results of the multimodal large model is calculated, and the corresponding secondary score is determined based on the accuracy. Based on each secondary score and its corresponding score weight, the primary score corresponding to the primary indicator to which each secondary indicator belongs is calculated.
[0060] In one embodiment, an evaluation score is calculated based on the output of the multimodal large model and the correct answers to the evaluation questions. Specifically, a corresponding score weight is determined based on the weight of each evaluation indicator. In this embodiment, a corresponding score weight is set for different indicator weights, thereby amplifying the score differences between different tested models on the evaluation indicators that users focus on. This allows users to more intuitively understand the performance differences of different models in the dimensions they focus on, facilitating model selection and subsequent optimization. Specifically, higher score weights are assigned to evaluation indicators with higher indicator weights. Based on the evaluation questions corresponding to each secondary indicator, the accuracy rate of the multimodal large model's output is calculated, and this accuracy rate is used as the secondary score corresponding to that secondary indicator. Further, based on the secondary scores and corresponding score weights of each secondary indicator, the primary score corresponding to each secondary indicator is calculated. The sum of the score weights of all secondary indicators belonging to the same primary indicator is 1.
[0061] Suppose a user specifies two secondary metrics, metric B and metric C, under primary metric A. To evaluate the performance of a multimodal large model, based on the model's output, the accuracy rate of metric B is calculated to be m, and the accuracy rate of metric C is calculated to be n. The primary score for metric A is calculated using the following expression: The first-level score of indicator A = m × the score weight of indicator B + n × the score weight of indicator C.
[0062] It is worth noting that if the user only specifies some of the secondary indicators under the primary indicator for model performance evaluation, that is, some secondary indicators may not have secondary scores. For the secondary indicators that the user did not specify for evaluation, their secondary scores will be set to 0.
[0063] By calculating the first-level scores of each primary indicator based on different scoring weights, the performance of different models can be more clearly distinguished in the dimensions that users focus on, making it easier for users to select models and perform subsequent optimizations, thereby improving the user experience.
[0064] In one embodiment, equal score weights (i.e., 0.5) are assigned to the two secondary indicators under each primary indicator. This allows for performance evaluation of the model across different dimensions (i.e., the dimensions corresponding to each primary indicator) without dimensional bias, and the primary score for each primary indicator is calculated as follows: The first-level score of perception = the second-level score of spatial perception × 0.5 + the second-level score of visual perception × 0.5; The first-level score for attention = the second-level score for attention concentration × 0.5 + the second-level score for name recognition × 0.5; The Level 1 score for reasoning ability = Level 2 score for analogical reasoning ability × 0.5 + Level 2 score for programming ability × 0.5; The Level 1 score for memory is calculated as follows: Level 2 score for short-term visual memory × 0.5 + Level 2 score for episodic memory × 0.5.
[0065] In one embodiment of this application, the analysis module is configured to generate performance evaluation results of the multimodal large model based on the evaluation score, specifically including: Determine the evaluation score range to which the first-level score of each first-level indicator belongs, obtain the evaluation information corresponding to the evaluation score range, and determine the evaluation information as the performance evaluation result of the first-level score; Based on the performance evaluation results of each first-level score, the performance evaluation results of the multimodal large model are generated.
[0066] In one embodiment, the analysis module generates a performance evaluation result for the multimodal large model in that indicator dimension based on the calculated first-level scores of each first-level indicator. In this embodiment, each first-level indicator has multiple evaluation score intervals corresponding to the dimension, and each evaluation score interval corresponds to a performance evaluation result. Optionally, each first-level indicator corresponds to three evaluation score intervals: the first evaluation score interval corresponds to a high performance evaluation result; the second evaluation score interval corresponds to a medium performance evaluation result; and the third evaluation score interval corresponds to a low performance evaluation result. After calculating the first-level score of the multimodal large model for the first-level indicator, the first-level score is compared with all evaluation score intervals in that dimension to determine which evaluation score interval the first-level score falls into, and the corresponding performance evaluation result is generated based on the determined evaluation score interval. For example, if the first-level score corresponding to perception (a first-level indicator) falls within the first evaluation score interval, then the performance evaluation result for the perception dimension is high.
[0067] Furthermore, the performance evaluation results of each primary indicator's corresponding dimension are combined to obtain the performance evaluation results of the multimodal large model. These performance evaluation results include the performance evaluation results of the primary indicators' corresponding dimensions to which all secondary indicators specified by the user belong.
[0068] Optionally, based on the first-level scores corresponding to each first-level indicator and the score weights corresponding to each first-level indicator, a weighted summation is performed to calculate the comprehensive score of the multimodal large model. Then, the comprehensive scores of different tested models are sorted to facilitate users in selecting models with higher comprehensive scores (i.e., better overall performance) to perform the target task, thus meeting user needs. It is worth noting that the sum of the score weights corresponding to all first-level indicators (including: perception, attention, memory, and reasoning) is 1. The expression for calculating the comprehensive score is as follows: The overall score is calculated as follows: Level 1 score of perception × score weight + Level 1 score of attention × score weight + Level 1 score of memory × score weight + Level 1 score of reasoning × score weight.
[0069] In one embodiment, the multimodal large model is comprehensively evaluated based on all evaluation metrics, thereby comprehensively assessing the model's performance across all dimensions without dimensional bias. Equal weights are assigned to each primary metric: perception (0.25), attention (0.25), memory (0.25), and reasoning (0.25). Based on this, the model's overall score is calculated using the following expression: Overall score = Level 1 score of perception × 0.25 + Level 1 score of attention × 0.25 + Level 1 score of memory × 0.25 + Level 1 score of reasoning × 0.25.
[0070] In one embodiment of this application, the system further includes an interaction module configured to perform the following steps: The selectable primary and secondary indicators are displayed on the user interface, and at least one evaluation indicator specified by the user is sent to the evaluation execution module. The evaluation questions and the output results of the multimodal large model are displayed on the user interface; After the analysis module generates the performance evaluation results, the performance evaluation results and / or the scores corresponding to each evaluation indicator are displayed on the user interface.
[0071] In one embodiment, the system further includes an interaction module for displaying selectable evaluation metrics (including primary and secondary metrics) on the user interface. After the user selects an evaluation metric on the user interface, the interaction module generates an evaluation task based on the user-specified multimodal large model and the corresponding evaluation metric, and sends it to the evaluation execution module. This allows the evaluation execution module to retrieve corresponding evaluation questions from the evaluation question bank based on the evaluation metrics in the evaluation task, and perform performance evaluation on the multimodal large model.
[0072] In this embodiment, the interaction module provides a visual interface (i.e., user interface) in a component-based manner using the Vue framework. During the multimodal large model's processing of evaluation questions, the interaction module displays the evaluation questions and the real-time output results generated by the multimodal large model on the user interface via WebSocket, allowing users to intuitively see the multimodal large model's problem-solving process. For example, based on the real-time processing progress of the multimodal large model, the interaction module compares and displays the model's output result for the current evaluation question with the correct answer on the user interface, using different colors to distinguish between correct and incorrect answers for easy user viewing.
[0073] After the analysis module generates the performance evaluation results of the multimodal large model, the interaction module displays the performance evaluation results on the user interface for viewing.
[0074] The analysis module generates various types of charts, such as radar charts, bar charts, and pie charts, based on the scores of the multimodal large model for each evaluation metric. These charts are then displayed on the user interface through the interactive module, allowing users to intuitively understand the distribution of the model's four-dimensional capabilities (i.e., perception, attention, memory, and reasoning) and the performance comparison of each evaluation metric.
[0075] In addition, the interaction module generates evaluation reports based on the model's personality assessment results. These reports include model processing accuracy, processing time, scores for each dimension, and performance evaluation result icons, allowing users to view the results offline and perform subsequent work. The interaction module also provides users with leaderboards, specifically a comprehensive leaderboard and dimension-specific leaderboards. The comprehensive leaderboard displays the overall capabilities of the mainstream large-scale models currently participating in the evaluation, while the dimension-specific leaderboards display the rankings of the capabilities of the mainstream large-scale models in each dimension. This allows users to refer to the selection of models for their target tasks and subsequent model optimization. For example, based on a short video generation task, a user can select a high-ranking multimodal large-scale model in the attention dimension from the leaderboard, further optimize the model's attention processing capabilities, and then use the optimized model for the short video generation task.
[0076] Figure 2 This is a flowchart of the multimodal large model evaluation system in one embodiment of this application. Figure 2 As shown, evaluation indicators are determined based on cognitive psychology theory and information processing models, including primary and secondary indicators, forming a hierarchical evaluation framework. Third-party raw data is acquired, and corresponding evaluation questions are constructed based on the evaluation indicators. An evaluation question bank is built based on the evaluation questions for each indicator. A user interface is provided through an interactive module (visual interactive terminal). Users specify the model to be evaluated and the evaluation indicators through the visual interactive terminal. The visual interactive terminal generates the evaluation task and sends it to the evaluation execution module 101. The evaluation execution module 101 determines the number of questions to be extracted according to the evaluation indicators specified by the user and obtains the corresponding number of evaluation questions from the evaluation question bank using a random sampling method. The evaluation questions are input into the multimodal large model for processing in a manner similar to a human exam, and the output results of the multimodal large model are stored in a MySQL database. The analysis module retrieves the output results of the multimodal large model from the database, calculates the evaluation score corresponding to each evaluation indicator based on the correct answers to each evaluation question, and generates the performance evaluation result of the multimodal large model based on the evaluation scores. Performance evaluation results are displayed through a visual interactive terminal for users to view.
[0077] Based on the same inventive concept, one embodiment of this application provides a multimodal large model evaluation method based on cognitive psychology. Figure 3 This is a flowchart of a multimodal large model evaluation method based on cognitive psychology proposed in an embodiment of this application. Figure 3 As shown, the method includes: S1: Extract corresponding evaluation questions from the evaluation question bank according to at least one evaluation indicator specified by the user; the evaluation indicator is determined based on the information processing model and cognitive psychology theory, and is used to evaluate the performance of at least one of the following dimensions of the multimodal large model: perception, attention, memory, and reasoning; the evaluation questions are constructed based on the evaluation indicator. S2: Input the evaluation questions into the multimodal large model to be evaluated and obtain the output results; S3: Based on the output results and the correct answers to the evaluation questions, calculate the evaluation score of the multimodal large model; S4: Generate the performance evaluation results of the multimodal large model based on the evaluation scores.
[0078] As one embodiment of this application, the method further includes pre-setting multiple evaluation indicators, specifically including: Based on information processing models and cognitive psychology theories, at least one primary indicator is set. Based on each primary indicator and the Wechsler Intelligence Scale for Children, at least two secondary indicators belonging to the primary indicators are generated.
[0079] As one embodiment of this application, the primary indicators include: perception, attention, memory, and reasoning ability; the secondary indicators include: visual perception and spatial perception belonging to perception, attention concentration and name recognition belonging to attention, short-term visual memory and episodic memory belonging to memory, and analogical reasoning ability and programming ability belonging to reasoning ability; the method further includes: Obtain various types of raw data, including: raw multiple-choice questions, raw short-answer questions, and raw fill-in-the-blank questions; Based on the aforementioned secondary indicators, the complete structure of the original multiple-choice questions is preserved, and the options of the original multiple-choice questions are rearranged in a random order to obtain the first type of evaluation questions; Based on the aforementioned secondary indicators, the stems of the original question-and-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions, which serve as the second type of evaluation questions. Based on the various secondary indicators, the first type of evaluation questions and the second type of evaluation questions are classified and stored to generate an evaluation question bank.
[0080] As one implementation of this application, based on the secondary indicators, the stems of the original question-and-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions as a second type of assessment questions, including: Based on the stem of the original question-and-answer question or the stem of the original fill-in-the-blank question, generate one correct option and multiple incorrect options, including at least one incorrect option whose similarity to the correct option is greater than or equal to a first threshold. Based on the order of similarity to the correct option from high to low, filter out multiple incorrect options with high similarity. Based on the correct options and several highly similar incorrect options selected, the second type of evaluation questions are generated.
[0081] As one embodiment of this application, after preserving the complete structure of the original multiple-choice questions and rearranging the options in a shuffled order, the method further includes: Obtain the correct and incorrect options for the original multiple-choice questions; Calculate the similarity between each incorrect option and the correct option, and compare it with a second threshold; If the similarity of all incorrect options is less than the second threshold, generate at least one new incorrect option with a similarity greater than or equal to the second threshold based on the correct option; The incorrect options in the original multiple-choice questions are replaced with the new incorrect options to obtain the first type of assessment questions.
[0082] As one implementation of this application, according to at least one evaluation indicator specified by the user, corresponding evaluation questions are extracted from the evaluation question bank, including: Obtain the weights of each evaluation metric specified by the user. Based on the weights of each evaluation indicator, the number of questions to be extracted is determined; the number of questions is greater than or equal to the third threshold. Based on the number of questions to be extracted for each evaluation indicator, a corresponding number of evaluation questions are obtained from the evaluation question bank. The evaluation questions corresponding to each evaluation indicator are mixed and randomly sorted to determine the order in which the multimodal large model processes the evaluation questions.
[0083] As one implementation of this application, the evaluation score of the multimodal large model is calculated based on the output results and the correct answers to the evaluation questions, specifically including: Based on the weight of each evaluation indicator, the score weight of each evaluation indicator is determined. Based on the evaluation questions corresponding to each secondary indicator, the accuracy of the output results of the multimodal large model is calculated, and the corresponding secondary score is determined based on the accuracy. Based on each secondary score and its corresponding score weight, the primary score corresponding to the primary indicator to which each secondary indicator belongs is calculated.
[0084] As one embodiment of this application, generating performance evaluation results for the multimodal large model based on the evaluation scores includes: Determine the evaluation score range to which the first-level score of each first-level indicator belongs, obtain the evaluation information corresponding to the evaluation score range, and determine the evaluation information as the performance evaluation result of the first-level score; Based on the performance evaluation results of each first-level score, the performance evaluation results of the multimodal large model are generated.
[0085] As one embodiment of this application, the method further includes: The selectable primary and secondary indicators are displayed on the user interface, and at least one evaluation indicator specified by the user is sent to the evaluation execution module. The evaluation questions and the output results of the multimodal large model are displayed on the user interface; After the analysis module generates the performance evaluation results, the performance evaluation results and / or the scores corresponding to each evaluation indicator are displayed on the user interface.
[0086] Based on the same inventive concept, one embodiment of this application provides a computer program product. This computer program product includes a computer program that, when executed by a processor, implements the steps in the multimodal large model evaluation method as described in any of the above embodiments of this application.
[0087] Based on the same inventive concept, one embodiment of this application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the multimodal large model evaluation method as described in any of the above embodiments of this application.
[0088] Based on the same inventive concept, one embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the steps in the multimodal large model evaluation method as described in any of the above embodiments of this application.
[0089] Regarding the methods in the above embodiments, the specific ways in which each module performs its operations have been described in detail in the embodiments of the system, and will not be elaborated here.
[0090] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0091] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and components involved are not necessarily essential to this application.
[0092] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0093] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0094] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0096] Although preferred embodiments of the embodiments of this application have been described, those skilled in the art, once they understand the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, this application is to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of this application.
[0097] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0098] The above provides a detailed description of the multimodal large model evaluation system and method based on cognitive psychology provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A multimodal large model evaluation system based on cognitive psychology, characterized in that, include: The evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank based on at least one evaluation indicator specified by the user. The evaluation metrics are determined based on information processing models and cognitive psychology theories, and are used to evaluate the performance of at least one of the following dimensions of the multimodal large model: perception, attention, memory, and reasoning. The evaluation questions are constructed based on the evaluation metrics; the evaluation questions are input into the multimodal large model to be evaluated to obtain the output results; The analysis module is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions; The performance evaluation results of the multimodal large model are generated based on the evaluation scores.
2. The multimodal large model evaluation system based on cognitive psychology according to claim 1, characterized in that, Also includes: Theoretical mapping module and question bank construction module; The theoretical mapping module is configured to set multiple evaluation indicators based on information processing models and cognitive psychology theories, specifically including multiple primary indicators and multiple secondary indicators belonging to each primary indicator. The primary indicators include: perception, attention, memory, and reasoning ability; the secondary indicators include: visual perception and spatial perception (belonging to perception), attention concentration and name recognition (belonging to attention), short-term visual memory and episodic memory (belonging to memory), and analogical reasoning ability and programming ability (belonging to reasoning ability). The question bank construction module is configured to perform the following steps: Obtain various types of raw data, including: raw multiple-choice questions, raw short-answer questions, and raw fill-in-the-blank questions; Based on the aforementioned secondary indicators, the complete structure of the original multiple-choice questions is preserved, and the options of the original multiple-choice questions are rearranged in a random order to obtain the first type of evaluation questions; Based on the aforementioned secondary indicators, the stems of the original question-and-answer questions or the original fill-in-the-blank questions are processed to generate multiple-choice questions, which serve as the second type of evaluation questions. Based on the various secondary indicators, the first type of evaluation questions and the second type of evaluation questions are classified and stored to generate an evaluation question bank.
3. The multimodal large model evaluation system based on cognitive psychology according to claim 2, characterized in that, The question bank construction module is configured to process the stems of the original open-ended questions or original fill-in-the-blank questions based on the secondary indicators, generating multiple-choice questions as the second type of assessment questions, specifically including: Based on the stem of the original question-and-answer question or the stem of the original fill-in-the-blank question, generate one correct option and multiple incorrect options, including at least one incorrect option whose similarity to the correct option is greater than or equal to a first threshold. Based on the order of similarity to the correct option from high to low, filter out multiple incorrect options with high similarity. Based on the correct options and several highly similar incorrect options selected, the second type of evaluation questions are generated.
4. The multimodal large model evaluation system based on cognitive psychology according to claim 2, characterized in that, The question bank construction module is further configured to, after preserving the complete structure of the original multiple-choice questions and rearranging the options in a shuffled order, perform the following steps: Obtain the correct and incorrect options for the original multiple-choice questions; Calculate the similarity between each incorrect option and the correct option, and compare it with a second threshold; If the similarity of all incorrect options is less than the second threshold, generate at least one new incorrect option with a similarity greater than or equal to the second threshold based on the correct option; The incorrect options in the original multiple-choice questions are replaced with the new incorrect options to obtain the first type of assessment questions.
5. The multimodal large model evaluation system based on cognitive psychology according to claim 2, characterized in that, The evaluation execution module is configured to extract corresponding evaluation questions from the evaluation question bank based on at least one evaluation indicator specified by the user, specifically including: Obtain the weights of each evaluation metric specified by the user. Based on the weights of each evaluation indicator, the number of questions to be extracted is determined; the number of questions is greater than or equal to the third threshold. Based on the number of questions to be extracted for each evaluation indicator, a corresponding number of evaluation questions are obtained from the evaluation question bank. The evaluation questions corresponding to each evaluation indicator are mixed and randomly sorted to determine the order in which the multimodal large model processes the evaluation questions.
6. The multimodal large model evaluation system based on cognitive psychology according to claim 2, characterized in that, The analysis module is configured to calculate the evaluation score of the multimodal large model based on the output results and the correct answers to the evaluation questions, specifically including: Based on the weight of each evaluation indicator, the score weight of each evaluation indicator is determined. Based on the evaluation questions corresponding to each secondary indicator, the accuracy of the output results of the multimodal large model is calculated, and the corresponding secondary score is determined based on the accuracy. Based on each secondary score and its corresponding score weight, the primary score corresponding to the primary indicator to which each secondary indicator belongs is calculated.
7. The multimodal large model evaluation system based on cognitive psychology according to claim 6, characterized in that, The analysis module is configured to generate performance evaluation results for the multimodal large model based on the evaluation scores, specifically including: Determine the evaluation score range to which the first-level score of each first-level indicator belongs, obtain the evaluation information corresponding to the evaluation score range, and determine the evaluation information as the performance evaluation result of the first-level score; Based on the performance evaluation results of each first-level score, the performance evaluation results of the multimodal large model are generated.
8. The multimodal large model evaluation system based on cognitive psychology according to claim 2, characterized in that, The system also includes an interaction module configured to perform the following steps: The selectable primary and secondary indicators are displayed on the user interface, and at least one evaluation indicator specified by the user is sent to the evaluation execution module. The evaluation questions and the output results of the multimodal large model are displayed on the user interface; After the analysis module generates the performance evaluation results, the performance evaluation results and / or the scores corresponding to each evaluation indicator are displayed on the user interface.
9. A multimodal large model evaluation method based on cognitive psychology, characterized in that, Applied to the system as described in any one of claims 1-8, comprising: Based on at least one evaluation metric specified by the user, corresponding evaluation questions are extracted from the evaluation question bank; the evaluation metric is determined based on the information processing model and cognitive psychology theory, and is used to evaluate the performance of at least one of the following dimensions of the multimodal large model: perception, attention, memory, and reasoning; the evaluation questions are constructed based on the evaluation metric. Input the evaluation questions into the multimodal large model to be evaluated, and obtain the output results; Based on the output results and the correct answers to the evaluation questions, the evaluation score of the multimodal large model is calculated; The performance evaluation results of the multimodal large model are generated based on the evaluation scores.
10. The multimodal large model evaluation method based on cognitive psychology according to claim 9, characterized in that, It also includes pre-setting multiple evaluation indicators, specifically including: Based on information processing models and cognitive psychology theories, at least one primary indicator is set. Based on each primary indicator and the Wechsler Intelligence Scale for Children, at least two secondary indicators belonging to the primary indicators are generated.
Citation Information
Patent Citations
Innovative attainment core cognitive ability combination evaluation system
CN116052880A
Assessment method and device of large language model, storage medium and computer equipment
CN117291184A
Evaluation method and system for large model content security capability
CN118035711A
Assessment method and device of large language model, storage medium and computer equipment
CN118627511A
Large model capability multi-dimensional evaluation method and device
CN118733413A
Cited By
Multi-modal objective scoring method, system and device for cognitive scale
CN121533700A