Multi-modal large language model evaluation method, system and equipment based on auto-reflection mechanism and medium

By constructing a parallel channel evaluation model with a multi-dimensional evaluation dataset and a self-reflective mechanism, the problem of inaccurate evaluation of isolated capabilities and robustness in the evaluation of multimodal large language models is solved, realizing a comprehensive and accurate evaluation of multimodal large language models and improving the robustness and automation of the evaluation.

CN121658367APending Publication Date: 2026-03-13SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing multimodal large language model evaluation methods only focus on isolated capabilities, resulting in inaccurate robustness evaluation and a lack of automation, failing to meet the comprehensiveness, accuracy, and automation requirements of practical applications.

Method used

We construct a multi-dimensional evaluation dataset and achieve comprehensive evaluation of vision, language, and robustness tasks through parallel dual-channel and parallel triple-channel evaluation models with self-reflection mechanisms, combined with data augmentation and model self-stabilization mechanisms.

Benefits of technology

It enables comprehensive and accurate evaluation of multimodal large language models, improves the accuracy and stability of robustness evaluation, reduces evaluation costs, reduces reliance on manual intervention, and enhances evaluation efficiency and objectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658367A_ABST
    Figure CN121658367A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large language model evaluation method, system and device based on an auto-reflection mechanism and a medium, and belongs to the technical field of artificial intelligence and multi-modal large language models. In order to solve the technical problems that an existing multi-modal large language model evaluation method only pays attention to isolation ability, robustness evaluation is inaccurate and automatic evaluation cannot be achieved, the adopted technical scheme includes the steps that a multi-modal evaluation data set is constructed, and multi-modal data in multiple fields are collected and preprocessed; constructing a multi-modal task comprising a visual subtask, a language subtask and a robustness subtask by utilizing the preprocessed multi-dimensional evaluation data, and marking and verifying the difficulty level of the multi-modal task; constructing an evaluation model of the multi-modal large language model based on the auto-reflection mechanism through a context engineering technology; and data summarization and analysis: summarizing the final scores of the various types of subtasks to obtain a comprehensive score of the to-be-tested model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and multimodal large language model technology, specifically a method, system, device and medium for evaluating multimodal large language models based on a self-reflection mechanism. Background Technology

[0002] With the rapid development of artificial intelligence technology, multimodal large language models, integrating multiple information modalities such as vision and language, have shown broad application prospects in various fields. However, existing evaluation methods for multimodal large language models based on self-reflection mechanisms have significant limitations, and traditional evaluation methods struggle to meet the demands of practical applications for comprehensiveness, accuracy, and automation. In recent years, the urgent need for efficient evaluation of multimodal large language models has driven increased interest in evaluation methods based on self-reflection mechanisms. The evaluation of multimodal large language models not only requires examining the basic capabilities of the visual and language modalities separately, but also needs to focus on cross-modal interaction effectiveness. Simultaneously, it must address issues such as data quality fluctuations and model output randomness in real-world scenarios, thereby achieving a comprehensive and efficient evaluation method.

[0003] Existing evaluation methods mostly focus on isolated dimensions of a model's capabilities, such as examining visual recognition accuracy or text generation quality in isolation, neglecting the model's overall performance during multimodal data interaction. This results in evaluation results that fail to reflect the model's actual effectiveness in real-world scenarios. Regarding robustness evaluation, traditional methods often employ single-disturbance forms or limited-sample testing, failing to fully simulate complex real-world scenarios such as fuzzy inputs, factual conflicts, and data noise. This makes it difficult to accurately capture the model's resistance to interference boundaries, leading to inaccurate robustness evaluations. Furthermore, existing evaluation systems heavily rely on human intervention, requiring manual intervention from task design and result judgment to score calibration. This is not only inefficient and costly but also prone to introducing evaluation biases due to subjective judgment differences. In addition, some automated evaluation methods lack dynamic adjustment mechanisms and cannot adaptively verify based on fluctuations in model output, resulting in insufficient stability and reliability of evaluation results. These problems severely restrict the iterative optimization and industrialization of multimodal large language models. Summary of the Invention

[0004] The technical objective of this invention is to provide a method, system, device, and medium for evaluating multimodal large language models based on a self-reflective mechanism, in order to address the problems of existing multimodal large language model evaluation methods that only focus on isolated capabilities, have inaccurate robustness evaluations, and cannot automate the evaluation process.

[0005] The technical objective of this invention is achieved as follows: a multimodal large language model evaluation method based on a self-reflective mechanism, the method being as follows:

[0006] Construct a multimodal evaluation dataset: Collect and preprocess multimodal data from multiple domains, and use the preprocessed multidimensional evaluation data to construct a multimodal task including visual subtasks, language subtasks, and robustness subtasks, and label and verify the difficulty level of the multimodal task;

[0007] An evaluation model for a multimodal large language model based on a self-reflective mechanism is constructed using context engineering techniques: Visual subtasks, language subtasks, and robustness subtasks are extracted from the received multimodal tasks. Representation enhancement processing is performed on the core task types of each multimodal task. Based on the core task examined by each subtask type, corresponding core task-specific prompts are selected, and the original subtask is concatenated with the core task-specific prompts to form the complete task instructions for the corresponding subtask. The received data-enhanced individual subtasks are then evaluated using a parallel dual-channel evaluation process to obtain the final output of the model under test. Finally, the model under test is independently scored using a parallel three-channel scoring process, and the average score is calculated to obtain the final score for each subtask.

[0008] Data aggregation and analysis: The final scores of each type of sub-task are aggregated to obtain the comprehensive score of the model under test.

[0009] As a preferred option, the multi-dimensional evaluation dataset is constructed as follows:

[0010] Collect raw data: Obtain and download multimedia datasets, biological image datasets, publicly available text datasets, and artificially constructed multimodal task data from public network resources;

[0011] Data preprocessing: Image data is cropped, scaled, and enhanced; text data is segmented and stop word removed; multiple-choice, fill-in-the-blank, and subjective question tasks are constructed.

[0012] Constructing multimodal tasks: Divide the preprocessed data into 181 task groups, each task group including at most 1 visual subtask, 1 language subtask and 3 robustness subtasks;

[0013] Core Task Design: Each visual subtask includes one or more core tasks, which fall into five categories: optical character recognition, visual recognition, spatial perception, motion recognition, and environmental understanding. Each language subtask is labeled with one or more core tasks, which fall into four categories: basic common sense, text generation, mathematical reasoning, and logical reasoning. Each robustness subtask is labeled with one core task, which falls into two categories: model illusion and fuzzy input. There are a total of 11 categories of core tasks across the visual, language, and robustness subtasks, each examining the corresponding model capabilities.

[0014] Manually labeled difficulty levels: Based on task complexity, required core capabilities, and modal interaction depth, visual and language subtasks are labeled with three difficulty levels: "high," "medium," and "low," using the following formula:

[0015]

[0016] Among them, D i The value represents the difficulty of the specific task; i represents the i-th sample group; High, Middle, and Low represent the three difficulty levels of "high, medium, and low" respectively; c represents the number of core tasks;

[0017] Difficulty verification: Input all tasks into several different multimodal large language models in sequence, calculate the task accuracy and compare it with the manually labeled difficulty level, and correct any deviations in the task difficulty level.

[0018] As a preferred approach, the evaluation model for a multimodal large language model based on a self-reflective mechanism, constructed using context engineering techniques, is as follows:

[0019] Data augmentation: Extract visual subtasks, language subtasks, and robustness subtasks from the received multimodal task. Perform visual subtask representation augmentation, language subtask representation augmentation, and robustness subtask representation augmentation on the core task types of the multimodal task, respectively. When performing representation augmentation on each type of subtask, select the corresponding core task-specific prompt words according to the core task examined by the corresponding subtask type, and concatenate the original visual subtask, language subtask, or robustness subtask with the corresponding core task-specific prompt words and general prompt words to form the complete task instruction of the corresponding subtask.

[0020] Model self-stabilization: Within the parallel dual channels of the model under test (including channel 1 and channel 2), the model under test receives and responds to the integrity task instructions of the corresponding subtasks. It then obtains the initial outputs of both channel 1 and channel 2, and performs a consistency check on these outputs. If the results are inconsistent, channel 1 and channel 2 are repeatedly designated as abnormal channels, and self-reflection is initiated, performing a limited number of regenerations until the outputs match those of channel 1 and channel 2. If the results match, the initial output of the model under test is taken as its final output.

[0021] A consensus scoring network is constructed: the final output of the model under test is concatenated with the scoring rules, inference hints, and scoring case samples to generate a complete scoring instruction. This complete scoring instruction is then input into three parallel scoring channels: Inference Model Scoring Channel 1, Inference Model Scoring Channel 2, and Inference Model Scoring Channel 3. Each channel independently scores the model, generating three initial scores. The dispersion of these three initial scores is calculated, and it is determined whether the dispersion exceeds a preset threshold. If the dispersion exceeds the preset threshold, a scoring discrepancy is identified, and the Inference Model Scoring Channel with the highest dispersion among the three channels is identified as an abnormal channel. Self-reflection is initiated, and a limited number of scoring corrections are performed until the dispersion does not exceed the preset threshold. If the dispersion does not exceed the preset threshold, the final score is directly output, and the average of the final scores is used as the final evaluation score for the corresponding sub-task.

[0022] More specifically, data augmentation includes the following:

[0023] Visual subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select a visual subtask and read the corresponding visual core task type. By adding general cue words and core task-specific cue words to the original visual subtask, obtain the visual subtask with enhanced representation, as shown in the following formula:

[0024] V enh =V raw +P fix,v +P v,c ;

[0025] Among them, V enh V represents the visual subtask after representation enhancement; raw P represents the input of the original visual subtask; fix,v Fixed cue words for visual subtasks; P v,c These are specific prompts for visual tasks; 'v' indicates a visual subtask; 'c' indicates the type of core visual task corresponding to the visual subtask.

[0026] Language subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select a language subtask and read the corresponding language core task type. By adding general cue words and core task-specific cue words to the original language subtask, obtain the language subtask with enhanced representation, as shown in the following formula:

[0027] L enh =V raw +P fix,l +P l,c ;

[0028] Among them, L enh This represents the language subtask after representation enhancement; V raw P represents the input of the original language subtask; fix,l Representing fixed prompts for language subtasks; P l,c The following are specific prompts for language subtasks: 'l' indicates a language subtask; 'c' indicates the language task type corresponding to the language subtask.

[0029] Robust subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select robust subtasks and read the corresponding robust core task types. By adding general cue words and core task-specific cue words to the original language subtasks, obtain the robust subtasks with enhanced representations, as shown in the following formula:

[0030] R enh =V raw +P fix,r +P r,c ;

[0031] Among them, R enh Represents the robustness of the subtask after representation enhancement; V raw P represents the original robustness subtask input; fix,r Fixed cue words indicating robustness in sub-tasks; P r,c The term "r" indicates a robust task; "r" indicates a robust subtask; and "c" indicates the robust core task type corresponding to the robust subtask.

[0032] More preferably, the model is self-stabilizing as follows:

[0033] Parallel dual-channel evaluation of the model under test: The evaluation task after representation enhancement is input into two independent and identical channels 1 and 2 of the model under test. Each channel contains a model instance of the model under test. The prediction results of the corresponding task are output as the initial outputs of channel 1 and channel 2 of the model under test.

[0034] Consistency verification of results: Input the initial outputs of channel one and channel two of the model under test into the large judge model, and obtain the output of the large judge model, as shown in the following formula:

[0035] J = LLM compare (T task,1 ,T task,2 ,P compare );

[0036] Where J represents the judgment result of the large-scale referee model; LLM compare Represents the large-scale model of the referee; P compare Indicates a contrasting keyword; T task,1 With Ttask,2 These represent the output results of channel one and channel two of the model under test, respectively; task represents the task category, namely visual subtask, language subtask, or robustness subtask.

[0037] Self-reflective regeneration: When the detection results show that the outputs of channel one and channel two of the model under test are inconsistent, the system iteratively selects channel one and channel two of the model under test from the first selected channel and regenerates the results. Using the previous input and output of the selected channel as context, the system uses dialogue commands to regenerate the results of the model under test. The formula is as follows:

[0038]

[0039] in, This represents the output of the test model in the nth round; n represents the number of regeneration rounds. When n is odd, channel one of the test model is selected for self-reflective regeneration; when n is even, channel two of the test model is selected for self-reflective regeneration; LLMs represents the test model; E represents the visual subtask, language subtask, or robustness subtask after representation enhancement. This indicates the output of the previous round; `task` indicates the task category, i.e., visual subtask, language subtask, or robustness subtask; `P` res Indicates the need to regenerate a specific prompt word; C cha Indicates a specific channel; cha represents the channel number;

[0040] If the judgment result of the large-scale model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent, then the self-reflection and regeneration are skipped, and the output of channel one of the test model is taken as the final result T′ of the visual subtask of the test model. task Otherwise, self-reflection and regeneration are performed until the judgment result of the large judge model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent. The output of the test model in the last round is taken as the final result T′ of the test model for the specific task. task If the loop reaches its limit, the output of the last iteration of the test model is taken as the final result T′ of the test model for the specific task. task ;

[0041] In the parallel dual-channel evaluation process of the model under test, the model receives visual sub-tasks, language sub-tasks, and robustness sub-tasks with enhanced representations. Based on the task type, each sub-task is then subjected to the corresponding evaluation, as detailed below:

[0042] Visual subtask evaluation: The visual subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the evaluation task after representation enhancement and outputs the prediction result of the corresponding task as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained, as shown in the following formula:

[0043] T v,cha =LLMs(V enh |C cha );

[0044] Among them, T v,cha V represents the output of the visual subtask of the model under test; enh Represents the visual subtask with enhanced representation; LLMs represent the model to be tested; C cha This indicates a specific channel; cha represents the channel number.

[0045] Language subtask evaluation: The language subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the outputs of the representation-enhanced evaluation task and the visual subtask, and outputs the prediction results of the corresponding tasks as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula:

[0046] T t,cha =LLMs(L enh ,T′ v |C cha );

[0047] Among them, T t,cha T′ represents the output of the language subtask of the model under test; v L represents the final result of the visual subtask of the model under test; enh Represents the language subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel; cha represents the channel number.

[0048] Robustness subtask evaluation: The robustness subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the representation-enhanced evaluation task and outputs the prediction result of the corresponding task as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula:

[0049] T r,cha =LLMs(R enh|C cha )

[0050] Among them, T r,cha R represents the output of the robustness subtask of the model under test; enh Represents the robustness subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel, where cha represents the channel number.

[0051] More optimally, the consensus scoring network is constructed as follows:

[0052] Scoring instruction generation: The final result of the model under test for a specific task is concatenated with inference instruction prompts, scoring rule prompts, and scoring case prompts to obtain the scoring instruction. The formula is as follows:

[0053] P score =P reason +P rule +P example +T′ task ;

[0054] Among them, P score Indicates a scoring instruction; P reason Indicates a clue word for reasoning instructions; P rule Indicates scoring rule prompts; P example Indicates a scoring case prompt; T′ task This represents the final result of the model under test for a specific task; task represents a visual subtask, a language subtask, or a robustness subtask.

[0055] Parallel Three-Channel Scoring: Input the scoring command into three independent and identical large-scale reasoning model scoring channels: Channel 1, Channel 2, and Channel 3. Each channel contains a large-scale reasoning model instance used for scoring. Obtain the preliminary score of the corresponding subtask in each of these channels, as shown in the following formula:

[0056] S cha =LLM reason (P score |C cha )

[0057] Among them, S cha This indicates the initial score of channel cha; cha is the channel number; C cha Indicates a specific channel; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction;

[0058] Scoring Difference Detection: Compare the initial scores of the large inference model scoring channel 1, large inference model scoring channel 2, and large inference model scoring channel 3. When an outlier occurs, the large inference model scoring channel 1 will start a self-reflective regeneration of the score to regenerate the score. After a finite number of regenerations or reaching the loop limit, the final scores S1′, S2′, and S3′ are obtained.

[0059] Self-reflective score regeneration: When the detection results show inconsistencies in the outputs of inference model scoring channels one, two, and three, the previous input and output of the selected abnormal channel are used as context. A special prompt for score regeneration is used to instruct the model under test to regenerate its score. The formula is as follows:

[0060]

[0061] in, This represents the score in round n; n represents the number of rounds to be regenerated; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction; P represents the score in round (n-1). res The score is regenerated using a dedicated prompt; C cha This indicates a specific channel, where cga represents the channel number;

[0062] Calculate the final score: Calculate the average of the three scores that meet the conditions, and use this average as the final evaluation score for the subtask. The formula is as follows:

[0063]

[0064] Wherein, S1′, S2′, and S3′ represent the scores output by the inference model scoring channel one, inference model scoring channel two, and inference model scoring channel three, respectively, which meet the difference requirements. This represents the score of the task in the nth group; task represents a visual subtask, a language subtask, or a robust subtask.

[0065] As a preferred approach, the data summary and analysis are as follows:

[0066] Visual subtask score summary calculation: Obtain the visual subtask scores within all groups, and calculate the overall score of the visual subtasks based on the weight of each subtask, using the following formula:

[0067]

[0068] Among them, S v This represents the overall score of the visual subtask; This represents the score of a single visual subtask across all groups; i represents the group number; v represents the visual subtask class; ω v,i This represents the task difficulty of a single visual subtask in all groups; j represents the group number; m represents the number of task groups in the dataset;

[0069] Language subtask score summary calculation: Obtain the language subtask scores within all groups, and calculate the overall language task score based on the weight of each subtask, using the following formula:

[0070]

[0071] Among them, S l This represents the overall score for the language task; The score for a single visual subtask across all groups; i represents the group number; l represents the language subtask class; ω l,i This represents the task difficulty of a single language subtask in all groups; j represents the group number; m represents the number of task groups in the dataset;

[0072] Robustness subtask score summary calculation: Obtain the robustness subtask scores for all groups, and calculate the overall robustness score based on the weight of each subtask, using the following formula:

[0073]

[0074] Among them, S r The combined score represents the overall score of the language subtasks; This represents the score of a single visual subtask across all groups; i represents the group number; i represents the language task class; m represents the number of task groups in the dataset;

[0075] The overall score of the model under test is calculated by combining the scores of the visual subtask, the language subtask, and the robustness subtask with the number of self-reflective regeneration attempts, according to the weight of each subtask. The formula is as follows:

[0076]

[0077] Where S represents the overall score of the model under test; S v S represents the overall score of the visual subtask; l S represents the overall score of the language subtasks; r denoted by α, z represents the number of self-reflection initiations; Z represents the maximum number of self-reflection initiations; α, β, γ and η are hyperparameters used to measure the proportion of the total score for each task class.

[0078] A multimodal large language model evaluation system based on a self-reflection mechanism is provided. This system implements the aforementioned multimodal large language model evaluation method based on a self-reflection mechanism. The system includes:

[0079] The dataset construction unit is used to collect and preprocess multimodal data from multiple domains, and to construct multimodal tasks including visual subtasks, language subtasks, and robustness subtasks using the preprocessed multidimensional evaluation data, and to label and verify the difficulty level of the multimodal tasks.

[0080] The evaluation model building unit is used to extract visual sub-tasks, language sub-tasks, and robustness sub-tasks from the received multimodal tasks. It performs representation enhancement processing on the core task types of each multimodal task, selects corresponding core task-specific prompts based on the core task examined by each sub-task type, and concatenates the original sub-task with the core task-specific prompts to form the complete task instruction for the corresponding sub-task. Then, the received data-enhanced individual sub-tasks are evaluated through parallel dual-channel testing of the model under test to obtain the final output of the model under test. Finally, the model under test is independently scored through parallel three-channel scoring, and the average score is calculated to obtain the final score for the corresponding sub-task.

[0081] The data aggregation and analysis unit is used to aggregate the final scores of each type of sub-task to obtain the comprehensive score of the model under test.

[0082] An electronic device includes: a memory and at least one processor;

[0083] The memory contains computer programs;

[0084] The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the multimodal large language model evaluation method based on the self-reflection mechanism described above.

[0085] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the self-reflective mechanism-based multimodal large language model evaluation method described above.

[0086] The multimodal large language model evaluation method, system, device, and medium based on self-reflection mechanism of the present invention have the following advantages:

[0087] (I) This invention constructs an evaluation system that includes eleven core tasks in three major categories: vision, language, and robustness. It examines both the basic capabilities of a single modality and focuses on the effectiveness of cross-modal interaction. At the same time, it combines multi-domain datasets and difficulty gradient design to effectively avoid the fragmentation of evaluation dimensions and ensure that the evaluation results can truly reflect the comprehensive performance of the model in real application scenarios. This solves the problem of the disconnect between isolated capability evaluation and real-world scenario requirements in the existing multimodal large language model evaluation based on self-reflection mechanism. Thus, it realizes a systematic evaluation scheme that covers multimodal collaborative capabilities.

[0088] (II) This invention solves the problems of limited robustness testing scenarios and inaccurate anti-interference ability assessment in traditional evaluations by adopting a three-level architecture of "data augmentation module - model self-stabilization module - consensus scoring network". Specifically, the robustness task representation augmentation module can simulate complex real-world interference scenarios such as fuzzy input, factual conflicts, and data noise; the model self-stabilization module corrects output fluctuations through parallel dual-channel verification and self-reflective regeneration; and the consensus scoring network optimizes scoring consistency through multi-channel dispersion detection. Compared to existing evaluation methods that only use a single form of interference and static scoring, this invention achieves a significant improvement in the accuracy and stability of robustness assessment results.

[0089] (III) This invention achieves automatic enhancement of task instructions through context engineering technology, then completes automatic verification and correction of the output of the model under test through the model self-stabilization module, and finally achieves automatic generation and calibration of evaluation scores through consensus scoring network. This framework not only gets rid of the high dependence of traditional evaluation on human intervention, greatly reduces evaluation costs and improves evaluation efficiency, but also avoids human subjective bias through dynamic adjustment mechanism, ensuring the objectivity of evaluation results.

[0090] (iv) This invention uses a task-type-specific representation enhancement strategy to identify the unique evaluation needs of visual, linguistic, and robust tasks. For visual tasks, it supplements the scene context and target attribute association with prompts. For linguistic tasks, it integrates elements such as logical chain guidance and knowledge background supplementation. For robust tasks, it designs strategies such as fuzzy input variants and factual conflict induction. This more meticulously captures the evaluation focus of different task types and ensures that the core assessment objectives of each task are accurately implemented. At the same time, it uses the logic of splicing general prompt words and special prompt words to lay a clear task boundary for subsequent model output verification and scoring calibration.

[0091] (V) This invention proposes to leverage the synergistic advantages of "parallel dual-channel evaluation" and "self-reflective regeneration" to seamlessly integrate model output stability verification and dynamic correction, thereby addressing the unreliability of evaluation results caused by the randomness of multimodal large language model outputs. This mechanism mainly consists of two parts: dual-channel synchronous evaluation and inconsistency self-reflection. The former runs the model under test synchronously through two independent channels to obtain dual outputs under the same input to verify stability. The latter generates guiding prompts based on the context of the preceding output when the dual outputs are inconsistent, driving the model to regenerate a limited number of times until the outputs are consistent (or reaching the loop limit). This mechanism effectively reduces the interference of model output fluctuations on the evaluation results and significantly improves the reliability of the evaluation results.

[0092] (vi) This invention minimizes the subjective bias of a single scoring channel by using parallel three-channel scoring and scoring difference detection. At the same time, it combines self-reflection and regeneration of scores to ensure that the final score not only meets the requirements of the scoring rules, but also accurately matches the core assessment target of the task. This promotes scoring consistency and enhances the ability of the scoring results to distinguish the model's capabilities.

[0093] (vii) This invention solves the problems of existing multimodal large language model evaluation methods based on self-reflection mechanisms that only focus on isolated capabilities, have inaccurate robustness evaluation, and rely on manual evaluation, and achieves a comprehensive and accurate evaluation of multimodal large language models. Attached Figure Description

[0094] The invention will be further described below with reference to the accompanying drawings.

[0095] Appendix Figure 1 The flowchart is shown for a multimodal large language model evaluation method based on a self-reflection mechanism.

[0096] Appendix Figure 2 This is a schematic diagram of the structure of a multimodal large language model evaluation system based on a self-reflection mechanism;

[0097] Appendix Figure 3 A flowchart for constructing an evaluation model based on a self-reflection mechanism;

[0098] Appendix Figure 4 A flowchart for constructing a multi-dimensional evaluation dataset;

[0099] Appendix Figure 5 A flowchart for the data augmentation module;

[0100] Appendix Figure 6 The flowchart is for the model self-stabilizing module;

[0101] Appendix Figure 7 A flowchart for a consensus scoring network. Detailed Implementation

[0102] The following detailed description of the multimodal large language model evaluation method, system, device, and medium based on a self-reflection mechanism of the present invention is provided with reference to the accompanying drawings and specific embodiments.

[0103] Example 1:

[0104] As attached Figure 1 As shown in the figure, this embodiment provides a multimodal large language model evaluation method based on a self-reflection mechanism. The specific method is as follows:

[0105] S1. Construct a multimodal evaluation dataset: Collect and preprocess multimodal data from multiple domains, and use the preprocessed multidimensional evaluation data to construct a multimodal task including visual subtasks, language subtasks, and robustness subtasks. Label and verify the difficulty level of the multimodal task.

[0106] S2. Construct an evaluation model for a multimodal large language model based on a self-reflective mechanism using context engineering techniques: Extract visual subtasks, language subtasks, and robustness subtasks from the received multimodal tasks. Perform representation enhancement processing on the core task types of the multimodal tasks respectively. Based on the core tasks examined by the corresponding subtask types, select corresponding core task-specific prompt words, and concatenate the original subtask with the core task-specific prompt words to form the complete task instructions for the corresponding subtasks. Then, evaluate the received data-enhanced individual subtasks through parallel dual-channel testing of the model under test to obtain the final output of the model under test. Finally, independently score the model under test through parallel three-channel scoring and calculate the average score to obtain the final score of the corresponding subtask.

[0107] S3. Data Summary and Analysis: Summarize the final scores of each type of sub-task to obtain the comprehensive score of the model under test.

[0108] As attached Figure 4 As shown, the construction of the multi-dimensional evaluation dataset in step S1 of this embodiment is as follows:

[0109] S101. Collect raw data: Obtain and download multimedia datasets, biological image datasets, publicly available text datasets, and artificially constructed multimodal task data from public network resources;

[0110] S102. Data preprocessing: Image data is cropped, scaled and enhanced; text data is segmented and stop words are removed; multiple-choice questions, fill-in-the-blank questions and subjective questions are constructed to make the questions basic solvable and more relevant to real-world application scenarios.

[0111] S103. Construct multimodal tasks: Divide the preprocessed data into 181 task groups, each task group including at most 1 visual subtask, 1 language subtask and 3 robustness subtasks;

[0112] S104. Core Task Design: Each visual subtask includes one or more core tasks. The core tasks of the visual subtasks include five categories: optical character recognition, visual recognition, spatial perception, motion recognition, and environmental understanding. Each language subtask is labeled with one or more core tasks. The core tasks of the language subtasks include four categories: basic common sense, text generation, mathematical and logical reasoning. Each robustness subtask is labeled with one core task. The core tasks of the robustness subtasks include two categories: model illusion and fuzzy input. There are a total of 11 categories of core tasks in the visual subtasks, language subtasks, and robustness subtasks, which examine the corresponding model capabilities. The detailed categories will help in the design of dedicated prompts, and the guided task text will help stabilize the scoring process.

[0113] S105. Manually labeled difficulty levels: Based on task complexity, required core capabilities, and modal interaction depth, visual and language subtasks are labeled with three difficulty levels: "high," "medium," and "low," using the following formula:

[0114]

[0115] Among them, D i The difficulty level represents the specific task; i represents the i-th sample group; High, Middle, and Low represent the three levels of difficulty: "high," "medium," and "low," respectively; c represents the number of core tasks; the difficulty level classification can better distinguish the model's capabilities and is closer to the actual application scenario.

[0116] S106. Difficulty Validation: Input all tasks sequentially into 12 different multimodal large language models, calculate the task accuracy and compare it with the manually labeled difficulty levels, and correct any biased task difficulty levels; the different multimodal large language models include Qwen3-VL-32B-Instruct, GPT-4V, Gemini-2.5Pro, Claude3-Opus, Llava-Next-7B, InternVL-2-26B, CogVLM2, Minimax-VL-01, Baichuan-VL, Pixtral-12B, Phi-3-Vision, Step3 and Qwen2-VL-7B-Instruct.

[0117] As attached Figure 3 As shown, the evaluation model for constructing a multimodal large language model based on a self-reflective mechanism through context engineering techniques in step S2 of this embodiment is as follows:

[0118] S201. Data Augmentation: Extract the visual subtask, language subtask, and robustness subtask from the received multimodal task, and perform visual subtask representation augmentation, language subtask representation augmentation, and robustness subtask representation augmentation on the core task type of the multimodal task, respectively. When performing representation augmentation on each type of subtask, select the corresponding core task-specific prompt words according to the core task examined by the corresponding type of subtask, and concatenate the original visual subtask, language subtask, or robustness subtask with the corresponding core task-specific prompt words and general prompt words to form the complete task instruction of the corresponding subtask.

[0119] S202, Model Self-Stabilization: Within the parallel dual channels of the model under test (including channel 1 and channel 2), the model under test receives the integrity task instruction of the corresponding subtask and responds. It obtains the initial output of the model under test from channel 1 and channel 2 respectively, and performs a consistency check on the initial outputs of channel 1 and channel 2. If the results are inconsistent, channel 1 and channel 2 are repeatedly designated as abnormal channels and self-reflection is initiated, performing a limited number of regenerations until the outputs of channel 1 and channel 2 are consistent. If the results are consistent, the initial output of the model under test is taken as the final output of the model under test.

[0120] S203. Construct a consensus scoring network: Concatenate the final output of the model under test with the scoring rules, inference hints, and scoring case samples to generate a complete scoring instruction. Input the complete scoring instruction into three parallel scoring channels: Inference Model Scoring Channel 1, Inference Model Scoring Channel 2, and Inference Model Scoring Channel 3. Each channel independently scores the model, obtaining three initial scores. Calculate the dispersion of the three initial scores and determine if the dispersion exceeds a preset threshold. If the dispersion exceeds the preset threshold, it is determined that there is a scoring discrepancy. The Inference Model Scoring Channel with the largest dispersion among the three channels is identified as an abnormal channel, and self-reflection is initiated to perform a limited number of scoring corrections until the dispersion does not exceed the preset threshold. If the dispersion does not exceed the preset threshold, the final score is directly output, and the average of the final scores is used as the final evaluation score for the corresponding sub-task.

[0121] As attached Figure 5 As shown, the data augmentation in step S201 of this embodiment is as follows:

[0122] S20101, Visual Subtask Representation Enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select the visual subtasks and read the corresponding core visual task types. By adding general cue words and core task-specific cue words to the original visual subtasks, obtain the visual subtasks with enhanced representations. The formula is as follows:

[0123] V enh =V raw +P fix,v +P v,c ;

[0124] Among them, V enh V represents the visual subtask after representation enhancement; raw P represents the input of the original visual subtask; fix,v Fixed cue words for visual subtasks; P v,c These are specific prompts for visual tasks; 'v' indicates a visual subtask; 'c' indicates the type of core visual task corresponding to the visual subtask.

[0125] S20102, Language Subtask Representation Enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select the language subtasks and read the corresponding language core task types. By adding general cue words and core task-specific cue words to the original language subtasks, obtain the language subtasks with enhanced representations. The formula is as follows:

[0126] L enh =V raw +P fix,l +P l,c ;

[0127] Among them, L enh This represents the language subtask after representation enhancement; V raw P represents the input of the original language subtask; fix,l Representing fixed prompts for language subtasks; P l,c The following are specific prompts for language subtasks: 'l' indicates a language subtask; 'c' indicates the language task type corresponding to the language subtask.

[0128] S20203, Robust Subtask Representation Enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select robust subtasks and read the corresponding robust core task types. By adding general cue words and core task-specific cue words to the original language subtasks, obtain the robust subtasks with enhanced representations. The formula is as follows:

[0129] R enh =V raw +P fix,r +P r,c ;

[0130] Among them, R enh Represents the robustness of the subtask after representation enhancement; V raw P represents the original robustness subtask input; fix,r Fixed cue words indicating robustness in sub-tasks; P r,c The term "r" indicates a robust task; "r" indicates a robust subtask; and "c" indicates the robust core task type corresponding to the robust subtask.

[0131] As attached Figure 6 As shown, the model self-stabilization in step S202 of this embodiment is as follows:

[0132] S20201, Parallel Dual-Channel Evaluation of the Model Under Test: The evaluation task after representation enhancement is input into two independent and identical models under test, Channel 1 and Channel 2. Each model under test contains a model instance of the model under test. The prediction results of the corresponding task are output as the initial outputs of Channel 1 and Channel 2. The parallel dual channels constitute a direct comparison, and their differences serve as the basis for initiating self-reflective regeneration.

[0133] S20202, Result Consistency Verification: Input the initial outputs of channel one and channel two of the model under test into the large judge model, and obtain the output of the large judge model, as shown in the following formula:

[0134] J = LLM compare (T task,1 ,T task,2 ,P compare );

[0135] Where J represents the judgment result of the large-scale referee model; LLM compare Represents the large-scale model of the referee; P compare Indicates a contrasting keyword; T task,1 With T task,2 These represent the output results of channel one and channel two of the model under test, respectively; task represents the task category, namely visual subtask, language subtask, or robustness subtask.

[0136] S20203, Self-Reflection and Regeneration: When the detection results show that the outputs of channel one and channel two of the model under test are inconsistent, the system iteratively selects channel one and channel two of the model under test from the first selected channel and regenerates them. Using the previous input and output of the selected channel as context, the system uses dialogue commands to regenerate the results of the model under test. The formula is as follows:

[0137]

[0138] in, This represents the output of the test model in the nth round; n represents the number of regeneration rounds. When n is odd, channel one of the test model is selected for self-reflective regeneration; when n is even, channel two of the test model is selected for self-reflective regeneration; LLMs represents the test model; E represents the visual subtask, language subtask, or robustness subtask after representation enhancement. This indicates the output of the previous round; `task` indicates the task category, i.e., visual subtask, language subtask, or robustness subtask; `P` res Indicates the need to regenerate a specific prompt word; C cha 'cha' indicates a specific channel; 'cha' indicates the channel number; self-reflection and regeneration further promote the thinking of the model under test. This external cues can make the model more confident in its output or more deeply reflect on its mistakes.

[0139] If the judgment result of the large-scale model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent, then the self-reflection and regeneration are skipped, and the output of channel one of the test model is taken as the final result T′ of the visual subtask of the test model. task Otherwise, self-reflection and regeneration are performed until the judgment result of the large judge model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent. The output of the test model in the last round is taken as the final result T′ of the test model for the specific task. task If the loop reaches its limit, the output of the last iteration of the test model is taken as the final result T′ of the test model for the specific task. task .

[0140] In the parallel dual-channel evaluation process of the model under test in step S20201 of this embodiment, the visual subtask, language subtask, and robustness subtask with enhanced representations are received, and the corresponding evaluation is performed according to the task type, as follows:

[0141] S2020101, Visual Subtask Evaluation: The visual subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the evaluation task after representation enhancement and outputs the prediction result of the corresponding task as the initial output of the channel, thus obtaining the outputs of test model channel 1 and test model channel 2, as shown in the following formula:

[0142] T v,cha =LLMs(V enh |C cha );

[0143] Among them, T v,cha V represents the output of the visual subtask of the model under test; enh Represents the visual subtask with enhanced representation; LLMs represent the model to be tested; Ccha This indicates a specific channel; cha represents the channel number.

[0144] S2020102, Language Subtask Evaluation: The language subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the outputs of the representation-enhanced evaluation task and the visual subtask, and outputs the prediction results of the corresponding tasks as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula:

[0145] T t,cha =LLMs(L enh ,T′ v |C cha );

[0146] Among them, T t,cha T′ represents the output of the language subtask of the model under test; v L represents the final result of the visual subtask of the model under test; enh Represents the language subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel; cha represents the channel number.

[0147] S2020103, Robustness Subtask Evaluation: The robustness subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each channel contains a model instance of the test model. The model instance receives the representation-enhanced evaluation task and outputs the prediction result of the corresponding task as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula:

[0148] T r,cha =LLMs(R enh |C cha )

[0149] Among them, T r,cha R represents the output of the robustness subtask of the model under test; enh Represents the robustness subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel, where cha represents the channel number.

[0150] As attached Figure 7 As shown, the construction of the consensus scoring network in step S203 of this embodiment is as follows:

[0151] S20301, Scoring Instruction Generation: The final result of the model under test for a specific task is concatenated with the inference instruction prompts, scoring rule prompts, and scoring case prompts to obtain the scoring instructions. The formula is as follows:

[0152] P score =P reason +P rule +P example +T′ task ;

[0153] Among them, P score Indicates a scoring instruction; P reason Indicates a clue word for reasoning instructions; P rule Indicates scoring rule prompts; P example Indicates a scoring case prompt; T′ task This represents the final result of a specific task of the model under test; task represents a visual subtask, a language subtask, or a robust subtask; the scoring instructions contain a variety of predetermined prompts; the reasoning instructions plan the thinking path for the scoring model's reasoning; the scoring rules are the basis for the scoring model's scoring; and the scoring cases are the guarantee of the stability of the scoring model's output and structure.

[0154] S20302, Parallel Three-Channel Scoring: Input the scoring command into three independent and identical large-scale inference model scoring channels: Channel 1, Channel 2, and Channel 3. Each channel contains a large-scale inference model instance used for scoring. Obtain the preliminary score of the corresponding subtask in each of these channels, as shown in the following formula:

[0155] S cha =LLM reason (P score |C cha )

[0156] Among them, S cha This indicates the initial score of channel cha; cha is the channel number; C cha Indicates a specific channel; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction;

[0157] S20303, Scoring Difference Detection: Compare the initial scores of the large inference model scoring channel 1, the large inference model scoring channel 2, and the large inference model scoring channel 3. When an outlier occurs, the large inference model scoring channel 1 will start a self-reflective regeneration of the score to regenerate the score. After a finite number of regenerations or reaching the loop limit, the final scores S1′, S2′, and S3′ are obtained.

[0158] S20304, Scoring Self-Reflection and Regeneration: When the detection results show that the outputs of the inference model scoring channel 1, inference model scoring channel 2, and inference model scoring channel 3 are inconsistent, the previous input and output of the selected abnormal channel will be used as context. A special prompt for score regeneration will be used to make the model under test regenerate the score. The formula is as follows:

[0159]

[0160] in, This represents the score in round n; n represents the number of rounds to be regenerated; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction; P represents the score in round (n-1). res The score is regenerated using a dedicated prompt; C cha This indicates a specific channel, where cha represents the channel number; the self-reflective regeneration during the scoring phase further promotes the thinking of the scoring model. Considering that scoring itself is more subjective, the error correction effect of external instructions is more obvious.

[0161] S20305. Calculate the final score: Calculate the average of the three scores that meet the conditions, and use it as the final evaluation score for this subtask. The formula is as follows:

[0162]

[0163] Wherein, S1′, S2′, and S3′ represent the scores output by the inference model scoring channel one, inference model scoring channel two, and inference model scoring channel three, respectively, which meet the difference requirements. This represents the score of the task in the nth group; task represents a visual subtask, a language subtask, or a robust subtask.

[0164] The data summarization and analysis in step S3 of this embodiment are as follows:

[0165] S301. Visual Subtask Score Summary Calculation: Obtain the visual subtask scores for all groups, and calculate the overall score for each visual subtask based on its weight. The formula is as follows:

[0166]

[0167] Among them, S v This represents the overall score of the visual subtask; This represents the score of a single visual subtask across all groups; i represents the group number; v represents the visual subtask class; ω v,iThis represents the task difficulty of a single visual subtask in all groups; j represents the group number; m represents the number of task groups in the dataset;

[0168] S302. Language Subtask Score Summary Calculation: Obtain the language subtask scores for all groups, and calculate the overall language task score based on the weight of each subtask, using the following formula:

[0169]

[0170] Among them, S l This represents the overall score for the language task; The score for a single visual subtask across all groups; i represents the group number; l represents the language subtask class; ω l,i This represents the task difficulty of a single language subtask in all groups; j represents the group number; m represents the number of task groups in the dataset;

[0171] S303. Robustness Subtask Score Summary Calculation: Obtain the robustness subtask scores for all groups, and calculate the overall robustness score based on the weight of each subtask, using the following formula:

[0172]

[0173] Among them, S r The combined score represents the overall score of the language subtasks; This represents the score of a single visual subtask across all groups; i represents the group number; i represents the language task class; m represents the number of task groups in the dataset;

[0174] S304. Overall Score of the Model Under Test: Obtain the scores of the visual subtask, language subtask, and robustness subtask, and combine them with the number of self-reflective regeneration startups. Calculate the overall score of the model under test based on the weight of each subtask, using the following formula:

[0175]

[0176] Where S represents the overall score of the model under test; S v S represents the overall score of the visual subtask; l S represents the overall score of the language subtasks; r denoted by α, z represents the number of self-reflection initiations; Z represents the maximum number of self-reflection initiations; α, β, γ and η are hyperparameters used to measure the proportion of the total score for each task class.

[0177] Example 2:

[0178] This embodiment provides a multimodal large language model evaluation system based on a self-reflection mechanism. This system is used to implement the multimodal large language model evaluation method based on a self-reflection mechanism as described in Embodiment 1. The system includes:

[0179] The dataset construction unit is used to collect and preprocess multimodal data from multiple domains, and to construct multimodal tasks including visual subtasks, language subtasks, and robustness subtasks using the preprocessed multidimensional evaluation data, and to label and verify the difficulty level of the multimodal tasks.

[0180] The evaluation model building unit is used to extract visual sub-tasks, language sub-tasks, and robustness sub-tasks from the received multimodal tasks. It performs representation enhancement processing on the core task types of each multimodal task, selects corresponding core task-specific prompts based on the core task examined by each sub-task type, and concatenates the original sub-task with the core task-specific prompts to form the complete task instruction for the corresponding sub-task. Then, the received data-enhanced individual sub-tasks are evaluated through parallel dual-channel testing of the model under test to obtain the final output of the model under test. Finally, the model under test is independently scored through parallel three-channel scoring, and the average score is calculated to obtain the final score for the corresponding sub-task.

[0181] The data aggregation and analysis unit is used to aggregate the final scores of each type of sub-task to obtain the comprehensive score of the model under test.

[0182] As attached Figure 2 As shown, the evaluation model construction unit in this embodiment includes:

[0183] A data augmentation module is constructed to enhance the representation of the original problem based on the problem type and core task type. Specifically, when a set of data is received, visual tasks, language tasks, and robustness sub-tasks are extracted, and visual task representation enhancement, language task representation enhancement, and robustness task representation enhancement are performed respectively based on the core task type of the problem. When performing representation enhancement for each task, corresponding core task-specific prompt words are selected according to the core task examined by the task. Then, the original task is concatenated with the core task-specific prompt words and general prompt words to form the complete task instruction for that sub-task.

[0184] A model self-stabilization module is constructed to evaluate the model under test to obtain its prediction results and to stabilize the output results of the model under test using self-reflection. Specifically, the parallel dual-channel evaluation receives a single sub-question after data augmentation. Within a single channel, the model under test receives the task instruction and responds to obtain the initial output of the model under test. The initial outputs of the two models under test in the parallel dual-channel evaluation are judged for consistency. If the conclusions are different, the abnormal channel initiates self-reflection and performs a limited number of regenerations. Otherwise, the final output of the model under test is directly output.

[0185] A consensus scoring network is constructed to score the final output of the model under test. Self-reflection is used to stabilize the output of the scoring model. Specifically, the final output of the model under test is concatenated with the scoring rules, thought chain prompts, and scoring case samples to generate a complete scoring instruction input in three parallel channels. Each channel of the inference model within the network performs independent scoring to obtain three initial scores. The scoring difference detection module calculates the dispersion of the three initial scores. If the dispersion exceeds a preset threshold, it is determined that there is a scoring difference, and the abnormal channel initiates self-reflection to perform a limited number of scoring corrections. Otherwise, the final score is directly output. After obtaining the final score, its average score is calculated as the final evaluation score for this subtask.

[0186] Example 3:

[0187] This embodiment also provides an electronic device, including: a memory and a processor;

[0188] The memory stores the instructions executed by the computer.

[0189] The processor executes computer execution instructions stored in the memory, causing the processor to execute the multimodal large language model evaluation method based on a self-reflective mechanism in any embodiment of the present invention.

[0190] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0191] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0192] Example 4:

[0193] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the multimodal large language model evaluation method based on a self-reflective mechanism according to any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.

[0194] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0195] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0196] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0197] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for evaluating multimodal large language models based on a self-reflection mechanism, characterized in that, The method is as follows: Construct a multimodal evaluation dataset: Collect and preprocess multimodal data from multiple domains, and use the preprocessed multidimensional evaluation data to construct a multimodal task including visual subtasks, language subtasks, and robustness subtasks, and label and verify the difficulty level of the multimodal task; An evaluation model for a multimodal large language model based on a self-reflection mechanism is constructed using context engineering techniques: visual subtasks, language subtasks, and robustness subtasks are extracted from the received multimodal tasks. Representation enhancement processing is performed on the core task types of the multimodal tasks respectively. Based on the core tasks examined by the corresponding subtask types, corresponding core task-specific prompt words are selected, and the original subtasks and core task-specific prompt words are concatenated to form the complete task instructions for the corresponding subtasks. The received data-enhanced individual subtasks are then evaluated using parallel dual-channel testing of the model under test to obtain the final output of the model under test. Finally, the model under test is independently scored using parallel three-channel scoring, and the average score is calculated to obtain the final score of the corresponding subtask. Data aggregation and analysis: The final scores of each type of sub-task are aggregated to obtain the comprehensive score of the model under test.

2. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 1, characterized in that, The multi-dimensional evaluation dataset is constructed as follows: Collect raw data: Obtain and download multimedia datasets, biological image datasets, publicly available text datasets, and artificially constructed multimodal task data from public network resources; Data preprocessing: Image data is cropped, scaled, and enhanced; text data is segmented and stop word removed; multiple-choice, fill-in-the-blank, and subjective question tasks are constructed. Constructing multimodal tasks: Divide the preprocessed data into 181 task groups, each task group including at most 1 visual subtask, 1 language subtask and 3 robustness subtasks; Core task design: Each visual subtask includes one or more core tasks, and the core tasks of the visual subtasks include five categories of tasks: optical character recognition, visual recognition, spatial perception, motion recognition, and environmental understanding; each language subtask is labeled with one or more core tasks, and the core tasks of the language subtasks include four categories of tasks: basic common sense, text generation, mathematical and logical reasoning; each robustness subtask is labeled with one core task, and the core tasks of the robustness subtasks include two categories of tasks: model illusion and fuzzy input. There are a total of 11 core tasks in the visual subtask, language subtask, and robustness subtask, which examine the corresponding model capabilities. Manually labeled difficulty levels: Based on task complexity, required core capabilities, and modal interaction depth, visual and language subtasks are labeled with three difficulty levels: "high," "medium," and "low," using the following formula: Among them, D i The value represents the difficulty of the specific task; i represents the i-th sample group; High, Middle, and Low represent the three difficulty levels of "high, medium, and low" respectively; c represents the number of core tasks; Difficulty verification: Input all tasks into several different multimodal large language models in sequence, calculate the task accuracy and compare it with the manually labeled difficulty level, and correct any deviations in the task difficulty level.

3. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 1, characterized in that, The evaluation model for a multimodal large language model based on a self-reflective mechanism, constructed using context engineering techniques, is as follows: Data augmentation: Extract the visual subtask, language subtask, and robustness subtask from the received multimodal task, and perform visual subtask representation augmentation, language subtask representation augmentation, and robustness subtask representation augmentation on the core task types of the multimodal task, respectively. When performing representation enhancement for each type of subtask, select the corresponding core task-specific prompt words according to the core task examined by the corresponding type of subtask, and concatenate the original visual subtask, language subtask, or robustness subtask with the corresponding core task-specific prompt words and general prompt words to form the complete task instruction for the corresponding subtask. Model self-stabilization: The model under test in the parallel dual channels of the model under test, including channel 1 and channel 2, receives the integrity task instruction of the corresponding subtask and responds. It obtains the initial output of the model under test in channel 1 and the initial output of the model under test in channel 2, respectively. It performs a consistency judgment on the initial output of the model under test in channel 1 and the initial output of the model under test in channel 2. If the results are inconsistent, it cyclically sets channel 1 and channel 2 as abnormal channels and starts self-reflection, performing a limited number of regenerations until they are consistent with the output of channel 1 and channel 2. If the results are consistent, the initial output of the model under test will be used as the final output of the model under test. Construct a consensus scoring network: The final output of the model under test is concatenated with the scoring rules, inference hints, and scoring case samples to generate a complete scoring instruction. The complete scoring instruction is then input into three parallel scoring channels: Inference Model Scoring Channel 1, Inference Model Scoring Channel 2, and Inference Model Scoring Channel 3. Each channel independently scores the model, obtaining three initial scores. The dispersion of the three initial scores is calculated, and it is determined whether the dispersion exceeds a preset threshold. If the dispersion exceeds the preset threshold, it is determined that there is a scoring difference. The Inference Model Scoring Channel with the largest dispersion among the three channels is identified as an abnormal channel, and self-reflection is initiated to perform a limited number of scoring corrections until the dispersion does not exceed the preset threshold. If the dispersion does not exceed the preset threshold, the final score is output directly, and the average of the final scores is used as the final evaluation score of the corresponding sub-task.

4. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 3, characterized in that, The data augmentation is as follows: Visual subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select a visual subtask and read the corresponding visual core task type. By adding general cue words and core task-specific cue words to the original visual subtask, obtain the visual subtask with enhanced representation, as shown in the following formula: In enh =V raw +P fix,v +P v,c ; Among them, V enh V represents the visual subtask after representation enhancement; raw P represents the input of the original visual subtask; fix,v Fixed cue words for visual subtasks; P v,c These are specific prompts for visual tasks; 'v' indicates a visual subtask; 'c' indicates the type of core visual task corresponding to the visual subtask. Language subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select a language subtask and read the corresponding language core task type. By adding general cue words and core task-specific cue words to the original language subtask, obtain the language subtask with enhanced representation, as shown in the following formula: L enh =V raw +P fix,l +P l,c ; Among them, L enh This represents the language subtask after representation enhancement; V raw P represents the input of the original language subtask; fix,l Representing fixed prompts for language subtasks; P l,c The following are specific prompts for language subtasks: 'l' indicates a language subtask; 'c' indicates the language task type corresponding to the language subtask. Robust subtask representation enhancement: Taking a set of tasks from the multimodal evaluation dataset as input, select robust subtasks and read the corresponding robust core task types. By adding general cue words and core task-specific cue words to the original language subtasks, obtain the robust subtasks with enhanced representations, as shown in the following formula: R enh =V raw +P fix,r +P r,c ; Among them, R enh Represents the robustness of the subtask after representation enhancement; V raw P represents the original robustness subtask input; fix,r Fixed cue words indicating robustness in sub-tasks; P r,c The term "r" indicates a robust task; "r" indicates a robust subtask; and "c" indicates the robust core task type corresponding to the robust subtask.

5. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 3, characterized in that, The model is self-stabilizing as follows: Parallel dual-channel evaluation of the model under test: The evaluation task after representation enhancement is input into two independent and identical channels 1 and 2 of the model under test. Each channel contains a model instance of the model under test. The prediction results of the corresponding task are output as the initial outputs of channel 1 and channel 2 of the model under test. Consistency verification of results: Input the initial outputs of channel one and channel two of the model under test into the large judge model, and obtain the output of the large judge model, as shown in the following formula: J=LLM compare (T task,1 ,T task,2 ,P compare ); Where J represents the judgment result of the large-scale referee model; LLM compare Represents the large-scale model of the referee; P compare Indicates a contrasting keyword; T task,1 With T task,2 These represent the output results of channel one and channel two of the model under test, respectively; task represents the task category, namely visual subtask, language subtask, or robustness subtask. Self-reflective regeneration: When the detection results show that the outputs of channel one and channel two of the model under test are inconsistent, the system iteratively selects channel one and channel two of the model under test from the first selected channel and regenerates the results. Using the previous input and output of the selected channel as context, the system uses dialogue commands to regenerate the results of the model under test. The formula is as follows: in, This represents the output of the test model in the nth round; n represents the number of regeneration rounds. When n is odd, channel one of the test model is selected for self-reflective regeneration; when n is even, channel two of the test model is selected for self-reflective regeneration; LLMs represents the test model; E represents the visual subtask, language subtask, or robustness subtask after representation enhancement. This indicates the output of the previous round; `task` indicates the task category, i.e., visual subtask, language subtask, or robustness subtask; `P` res Indicates the need to regenerate a specific prompt word; C cha Indicates a specific channel; cha represents the channel number; If the judgment result of the large-scale model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent, then the self-reflection and regeneration are skipped, and the output of channel one of the test model is taken as the final result T′ of the visual subtask of the test model. task Otherwise, self-reflection and regeneration are performed until the judgment result of the large judge model shows that the outputs of the two independent channels, channel one and channel two, of the test model are consistent. The output of the test model in the last round is taken as the final result T′ of the test model for the specific task. task If the loop reaches its limit, the output of the last iteration of the test model is taken as the final result T′ of the test model for the specific task. task ; In the parallel dual-channel evaluation process of the model under test, the model receives visual sub-tasks, language sub-tasks, and robustness sub-tasks with enhanced representations. Based on the task type, each sub-task is then subjected to the corresponding evaluation, as detailed below: Visual subtask evaluation: The visual subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the evaluation task after representation enhancement and outputs the prediction result of the corresponding task as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained, as shown in the following formula: T v,cha =LLMs(V enh |C cha ); Among them, T v,cha V represents the output of the visual subtask of the model under test; enh Represents the visual subtask with enhanced representation; LLMs represent the model to be tested; C cha This indicates a specific channel; cha represents the channel number. Language subtask evaluation: The language subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the outputs of the representation-enhanced evaluation task and the visual subtask, and outputs the prediction results of the corresponding tasks as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula: T t,cha =LLMs(L enh ,T′ v ∣C cha ); Among them, T t,cha T′ represents the output of the language subtask of the model under test; v L represents the final result of the visual subtask of the model under test; enh Represents the language subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel; cha represents the channel number. Robustness subtask evaluation: The robustness subtask, after representation enhancement, is input into two independent and identical test model channels 1 and 2. Each test model channel contains a model instance of the test model. The model instance receives the representation-enhanced evaluation task and outputs the prediction result of the corresponding task as the initial output of the channel. The outputs of test model channel 1 and test model channel 2 are obtained using the following formula: T r,cha =LLMs(R enh |C cha ) Among them, T r,cha R represents the output of the robustness subtask of the model under test; enh Represents the robustness subtask after representation enhancement; LLMs represent the model to be tested; C cha This indicates a specific channel, where cha represents the channel number.

6. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 3, characterized in that, The consensus scoring network is constructed as follows: Scoring instruction generation: The final result of the model under test for a specific task is concatenated with inference instruction prompts, scoring rule prompts, and scoring case prompts to obtain the scoring instruction. The formula is as follows: P score =P reason +P rule +P example +T′ task ; Among them, P score Indicates a scoring instruction; P reason Indicates a clue word for reasoning instructions; P rule Indicates scoring rule prompts; P example Indicates a scoring case prompt; T′ task This represents the final result of the model under test for a specific task; task represents a visual subtask, a language subtask, or a robustness subtask. Parallel Three-Channel Scoring: Input the scoring command into three independent and identical large-scale reasoning model scoring channels: Channel 1, Channel 2, and Channel 3. Each channel contains a large-scale reasoning model instance used for scoring. Obtain the preliminary score of the corresponding subtask in each of these channels, as shown in the following formula: S cha =LLM reason (P score |C cha ) Among them, S cha This indicates the initial score of channel cha; cha is the channel number; C cha Indicates a specific channel; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction; Scoring Difference Detection: Compare the initial scores of the large inference model scoring channel 1, large inference model scoring channel 2, and large inference model scoring channel 3. When an outlier occurs, the large inference model scoring channel 1 will start a self-reflective regeneration of the score to regenerate the score. After a finite number of regenerations or reaching the loop limit, the final scores S1′, S2′, and S3′ are obtained. Self-reflective score regeneration: When the detection results show inconsistencies in the outputs of inference model scoring channels one, two, and three, the previous input and output of the selected abnormal channel are used as context. A special prompt for score regeneration is used to instruct the model under test to regenerate its score. The formula is as follows: in, This represents the score in round n; n represents the number of rounds to be regenerated; LLM reason This represents the large inference model used for scoring; P score Indicates a scoring instruction; P represents the score in round (n-1). res The score is regenerated using a dedicated prompt; C cha This indicates a specific channel, where cha represents the channel number; Calculate the final score: Calculate the average of the three scores that meet the conditions, and use this average as the final evaluation score for the subtask. The formula is as follows: Wherein, S1′, S2′, and S3′ represent the scores output by the inference model scoring channel one, inference model scoring channel two, and inference model scoring channel three, respectively, which meet the difference requirements. This represents the score of the task in the nth group; task represents a visual subtask, a language subtask, or a robust subtask.

7. The multimodal large language model evaluation method based on self-reflection mechanism according to claim 1, characterized in that, The data summary and analysis are as follows: Visual subtask score summary calculation: Obtain the visual subtask scores within all groups, and calculate the overall score of the visual subtasks based on the weight of each subtask, using the following formula: Among them, S v This represents the overall score of the visual subtask; This represents the score of a single visual subtask across all groups; i represents the group number; v represents the visual subtask class; ω v,i This represents the task difficulty of a single visual subtask in all groups; j represents the group number; m represents the number of task groups in the dataset; Language subtask score summary calculation: Obtain the language subtask scores within all groups, and calculate the overall language task score based on the weight of each subtask, using the following formula: Among them, S l This represents the overall score for the language task; The score for a single visual subtask across all groups; i represents the group number; l represents the language subtask class; ω l,i This represents the task difficulty of a single language subtask in all groups; j represents the group number; m represents the number of task groups in the dataset; Robustness subtask score summary calculation: Obtain the robustness subtask scores for all groups, and calculate the overall robustness score based on the weight of each subtask, using the following formula: Among them, S r The combined score represents the overall score of the language subtasks; This represents the score of a single visual subtask across all groups; i represents the group number; i represents the language task class; m represents the number of task groups in the dataset; The overall score of the model under test is calculated by combining the scores of the visual subtask, the language subtask, and the robustness subtask with the number of self-reflective regeneration attempts, according to the weight of each subtask. The formula is as follows: Where S represents the overall score of the model under test; S v S represents the overall score of the visual subtask; l S represents the overall score of the language subtasks; r denoted by α, z represents the number of self-reflection initiations; Z represents the maximum number of self-reflection initiations; α, β, γ and η are hyperparameters used to measure the proportion of the total score for each task class.

8. A multimodal large language model evaluation system based on a self-reflection mechanism, characterized in that, This system is used to implement the multimodal large language model evaluation method based on a self-reflective mechanism as described in any one of claims 1 to 7; the system comprises: The dataset construction unit is used to collect and preprocess multimodal data from multiple domains, and to construct multimodal tasks including visual subtasks, language subtasks, and robustness subtasks using the preprocessed multidimensional evaluation data, and to label and verify the difficulty level of the multimodal tasks. The evaluation model building unit is used to extract visual sub-tasks, language sub-tasks, and robustness sub-tasks from the received multimodal tasks. It performs representation enhancement processing on the core task types of each multimodal task, selects corresponding core task-specific prompts based on the core task examined by each sub-task type, and concatenates the original sub-task with the core task-specific prompts to form the complete task instruction for the corresponding sub-task. Then, the received data-enhanced individual sub-tasks are evaluated through parallel dual-channel testing of the model under test to obtain the final output of the model under test. Finally, the model under test is independently scored through parallel three-channel scoring, and the average score is calculated to obtain the final score for the corresponding sub-task. The data aggregation and analysis unit is used to aggregate the final scores of each type of sub-task to obtain the comprehensive score of the model under test.

9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the multimodal large language model evaluation method based on a self-reflective mechanism as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the multimodal large language model evaluation method based on a self-reflective mechanism as described in any one of claims 1 to 7.

Citation Information

Cited By

  • General agent reflection system based on large-model multi-path fusion and gating mechanism

    CN121882219A