Multi-modal large language model evaluation method and related device

By acquiring user behavior signals, determining user behavior patterns, and adjusting the weights of evaluation dimensional indicators, this approach solves the problem that traditional evaluation methods cannot meet the evaluation requirements of multimodal large language models. It enables comprehensive evaluation of multimodal large language models, adapts to different scenario needs, and improves the accuracy and comprehensiveness of the evaluation.

CN121960780APending Publication Date: 2026-05-01IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional large language model evaluation methods mainly focus on text generation quality, which cannot meet the evaluation needs of multimodal large language models in different modalities such as text, images/icons, and tables.

Method used

By acquiring user behavior signals, user behavior patterns are determined, and the basic weights of evaluation dimension indicators are adjusted based on user behavior signals, including text modality, image modality, table modality, and cross-modal consistency evaluation. The weights are dynamically adjusted to evaluate the performance of the multimodal large language model.

Benefits of technology

It enables comprehensive evaluation of multimodal large language models, dynamically adapts to different scenario requirements, and improves the accuracy and comprehensiveness of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960780A_ABST
    Figure CN121960780A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large language model evaluation method and a related device, and relates to the technical field of large language models.The method comprises the steps that firstly, a user behavior signal corresponding to a to-be-evaluated multi-modal large language model is acquired; then, based on the user behavior signal, determining a user behavior mode; then, determining an evaluation dimension index basic weight corresponding to the user behavior mode; the evaluation dimension indexes corresponding to the user behavior mode are a plurality of evaluation dimension indexes in preset multi-modal evaluation dimension indexes; adjusting an evaluation dimension index basic weight corresponding to the user behavior mode based on the user behavior signal to obtain an evaluation dimension index new weight corresponding to the user behavior mode; and finally, the to-be-evaluated multi-modal large language model is evaluated based on the new weight of the evaluation dimension index corresponding to the user behavior pattern, an evaluation score of the to-be-evaluated multi-modal large language model is obtained, and evaluation of the multi-modal large language model can be realized based on the scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large language model technology, and in particular to a multimodal large language model evaluation method and related apparatus. Background Technology

[0002] The evaluation of Large Language Models (LLMs) is a crucial bridge connecting technological research and development with the realization of their value. It not only drives the evolution of the models themselves but also helps society establish reasonable expectations of their capabilities and ensures that their development aligns with the overall goals of safety, ethics, and practicality.

[0003] With the widespread application of large language models in finance, healthcare, education, and other fields, large language models have evolved into multimodal large language models (MLLMs). Traditional evaluation methods for large language models mainly focus on the quality of text generation, while multimodal large language models need to simultaneously evaluate the model's performance across different modalities such as text, images / icons, and tables. Therefore, traditional evaluation methods for large language models can no longer meet the evaluation requirements of multimodal large language models.

[0004] Therefore, how to provide an evaluation scheme for multimodal large language models to achieve the evaluation of multimodal large language models has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of the above problems, this application provides a method and related apparatus for evaluating multimodal large language models, so as to achieve the purpose of evaluating multimodal large language models. The specific solution is as follows:

[0006] The first aspect of this application provides a method for evaluating multimodal large language models, including:

[0007] Obtain user behavior signals corresponding to the multimodal large language model to be evaluated;

[0008] Based on the user behavior signals, determine the user behavior pattern;

[0009] Determine the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators in the preset multimodal evaluation dimension indicators.

[0010] Based on the user behavior signals, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0011] The multimodal language model to be evaluated is evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, and the evaluation score of the multimodal language model to be evaluated is obtained.

[0012] In one possible implementation, determining the user behavior pattern based on the user behavior signal includes:

[0013] The user behavior signal is input into the user behavior pattern prediction model. The user behavior pattern prediction model performs feature processing on the user behavior signal to obtain a user behavior feature vector, and predicts the user behavior pattern based on the user behavior feature vector.

[0014] In one possible implementation, the preset multimodal evaluation dimension indicators include text modality evaluation dimension indicators, image modality evaluation dimension indicators, table modality evaluation dimension indicators, and cross-modal coordination evaluation dimension indicators.

[0015] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0016] Obtain the mapping relationship between preset user behavior signals and evaluation dimension indicator weight adjustment rules;

[0017] Select the target evaluation dimension indicator weight adjustment rule that has a mapping relationship with the user behavior signal from the preset mapping relationship between user behavior signal and evaluation dimension indicator weight adjustment rule;

[0018] The basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted using the target evaluation dimension indicator weight adjustment rules to obtain the new weights of the evaluation dimension corresponding to the user behavior pattern.

[0019] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0020] Based on cross-modal consistency, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0021] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on cross-modal consistency to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0022] Calculate a cross-modal consistency score, which is used to indicate the consistency between the modal element pointed to by the user behavior signal and other relevant modal information;

[0023] Based on the cross-modal consistency score, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0024] In one possible implementation, the evaluation of the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, to obtain the evaluation score of the multimodal large language model to be evaluated, includes:

[0025] The evaluation dimension indicators corresponding to the user behavior pattern are scored to obtain the scores of the evaluation dimension indicators corresponding to the user behavior pattern.

[0026] Obtain the preset cross-modal consistency reward base score and consistency reward coefficient;

[0027] The evaluation score of the multimodal large language model to be evaluated is obtained based on the evaluation dimension index score corresponding to the user behavior pattern, the basic score of the cross-modal consistency reward, and the consistency reward coefficient.

[0028] A second aspect of this application provides a multimodal large language model evaluation device, comprising:

[0029] The acquisition unit is used to acquire user behavior signals corresponding to the multimodal large language model to be evaluated.

[0030] The user behavior pattern determination unit is used to determine the user behavior pattern based on the user behavior signal.

[0031] The basic weight determination unit is used to determine the basic weight of the evaluation dimension indicators corresponding to the user behavior pattern; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators in the preset multimodal evaluation dimension indicators.

[0032] The weight adjustment unit is used to adjust the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal, so as to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0033] The evaluation unit is used to evaluate the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, and to obtain the evaluation score of the multimodal large language model to be evaluated.

[0034] In one possible implementation, the user behavior pattern determination unit is specifically used for:

[0035] The user behavior signal is input into the user behavior pattern prediction model. The user behavior pattern prediction model performs feature processing on the user behavior signal to obtain a user behavior feature vector, and predicts the user behavior pattern based on the user behavior feature vector.

[0036] In one possible implementation, the preset multimodal evaluation dimension indicators include text modality evaluation dimension indicators, image modality evaluation dimension indicators, table modality evaluation dimension indicators, and cross-modal coordination evaluation dimension indicators.

[0037] In one possible implementation, the weight adjustment unit is specifically used for:

[0038] Obtain the mapping relationship between preset user behavior signals and evaluation dimension indicator weight adjustment rules;

[0039] Select the target evaluation dimension indicator weight adjustment rule that has a mapping relationship with the user behavior signal from the preset mapping relationship between user behavior signal and evaluation dimension indicator weight adjustment rule;

[0040] The basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted using the target evaluation dimension indicator weight adjustment rules to obtain the new weights of the evaluation dimension corresponding to the user behavior pattern.

[0041] In one possible implementation, the weight adjustment unit is specifically used for:

[0042] Based on cross-modal consistency, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0043] In one possible implementation, the weight adjustment unit is specifically used for:

[0044] Calculate a cross-modal consistency score, which is used to indicate the consistency between the modal element pointed to by the user behavior signal and other relevant modal information;

[0045] Based on the cross-modal consistency score, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0046] In one possible implementation, the evaluation unit is specifically used for:

[0047] The evaluation dimension indicators corresponding to the user behavior pattern are scored to obtain the scores of the evaluation dimension indicators corresponding to the user behavior pattern.

[0048] Obtain the preset cross-modal consistency reward base score and consistency reward coefficient;

[0049] The evaluation score of the multimodal large language model to be evaluated is obtained based on the evaluation dimension index score corresponding to the user behavior pattern, the basic score of the cross-modal consistency reward, and the consistency reward coefficient.

[0050] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the multimodal large language model evaluation method described in the first aspect or any implementation thereof.

[0051] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0052] The memory is used to store computer programs;

[0053] The processor is used to execute the computer program so that the electronic device can implement the multimodal large language model evaluation method of the first aspect or any implementation thereof.

[0054] The fifth aspect of this application provides a computer-readable storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform a multimodal large language model evaluation method as described in the first aspect or any implementation thereof.

[0055] By employing the above technical solution, the multimodal large language model evaluation method and related apparatus provided in this application first acquire user behavior signals corresponding to the multimodal large language model to be evaluated; then, based on the user behavior signals, determine the user behavior pattern; subsequently, determine the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators among the preset multimodal evaluation dimension indicators; and adjust the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signals to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern; finally, evaluate the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern to obtain the evaluation score of the multimodal large language model to be evaluated. Based on this scheme, the evaluation of multimodal large language models can be realized. Attached Figure Description

[0056] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0057] Figure 1 A flowchart illustrating a multimodal large language model evaluation method provided in this application embodiment;

[0058] Figure 2 A schematic diagram of the structure of a multimodal large language model evaluation device provided in this application embodiment;

[0059] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0060] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0061] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0062] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0063] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0064] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application program, server, or storage medium executing the operation of this invention, based on the prompt message.

[0065] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0066] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of the present invention. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present invention.

[0067] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0068] The evaluation of Large Language Models (LLMs) is a crucial bridge connecting technological research and development with the realization of their value. It not only drives the evolution of the models themselves but also helps society establish reasonable expectations of their capabilities and ensures that their development aligns with the overall goals of safety, ethics, and practicality.

[0069] Traditional evaluation methods for large language models include those based on automatic metrics (such as BLEU, ROUGE, and Perplexity), those based on standard test sets (such as benchmarks), methods where the large model acts as the judge, human evaluation, security and robustness testing, and simulated testing with real people. However, traditional large language model evaluation methods primarily focus on text generation quality, while multimodal large language models require simultaneous evaluation of the model's performance across different modalities such as text, images / icons, and tables. Therefore, traditional large language model evaluation methods are no longer sufficient to meet the evaluation needs of multimodal large language models.

[0070] To address the aforementioned problems, this application provides a method for evaluating multimodal large language models, which enables the evaluation of multimodal large language models. The method described below with reference to the accompanying drawings provides a detailed explanation of the large language model evaluation method of this application.

[0071] Reference Figure 1 , Figure 1 A flowchart illustrating a multimodal large language model evaluation method provided in this application embodiment is shown below. Figure 1 As shown in the embodiments of this application, a multimodal large language model evaluation method may include the following steps, which are described in detail below.

[0072] S101: Obtain user behavior signals corresponding to the multimodal large language model to be evaluated;

[0073] In this application, user behavior signals are used to indicate the time series of user behavior data in a certain page or session, including clicks, scrolling, dwell time, voice, facial expressions, etc.

[0074] S102: Determine the user behavior pattern based on the user behavior signal;

[0075] In this application, user behavior patterns are used to characterize user types. Different types of users have different user behavior patterns, and different user behavior patterns generate different user behavior signals. In this application, specific user behavior patterns can be set based on scenario requirements, and this application does not impose any limitations on this. In one possible implementation, the user behavior patterns include deep readers, fast browsers, error-correcting users, decision-making users, and verification users.

[0076] S103: Determine the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators in the preset multimodal evaluation dimension indicators.

[0077] In this application, the basic weights of the evaluation dimension indicators corresponding to each user behavior pattern can be preset. For ease of understanding, it is assumed that the user behavior patterns include in-depth readers, fast browsers, error-correcting users, decision-making users, and verification users. The embodiments of this application provide an example of the correspondence between each user behavior pattern and the basic weights of the evaluation dimension indicators, as follows:

[0078] User Behavior Patterns Evaluation Dimension Indicators Basic Weights Deep Reader Text accuracy (0.35), logical consistency (0.3), cross-modal consistency (0.25) Quick Browser Chart readability (0.4), summary clarity (0.3), key information prominence (0.3) Error-correcting users Data integrity (0.3), numerical accuracy (0.3), security compliance (0.4) Decision-making users Trend matching degree (0.4), data reliability (0.3), decision support (0.3) Validated users Semantic consistency (0.4), information complementarity (0.3), conflict detection capability (0.3)

[0079] S104: Adjust the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0080] In this application, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted based on the user behavior signals, and the new weights of the evaluation dimension indicators corresponding to the user behavior pattern can be consistent with the dynamic user experience of the multimodal large language model to be evaluated.

[0081] S105: Evaluate the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, and obtain the evaluation score of the multimodal large language model to be evaluated.

[0082] For example, in a financial analysis scenario, when generating financial reports, the system dynamically adjusts weights by monitoring user behavior (such as clicking on chart areas or providing voice feedback like "This data has a problem").

[0083] Basic weights: text accuracy (0.3), chart trend matching (0.25), numerical consistency (0.2), and multimodal fusion naturalness (0.25);

[0084] When users click on the Q2 revenue chart area, the weights are adjusted to text accuracy (0.3), chart trend matching degree (0.35), numerical consistency (0.25), and multimodal fusion naturalness (0.1);

[0085] When a user voice feedback message reads "There is a problem with this data," the weights are further adjusted to: text accuracy (0.45), chart trend matching (0.3), numerical consistency (0.15), and multimodal fusion naturalness (0.1).

[0086] Final assessment score: 0.85.

[0087] In medical diagnosis scenarios: When generating medical reports, the system dynamically adjusts weights by monitoring user behavior (such as lingering on a symptom description for an extended period or displaying a confused expression).

[0088] Basic weights: accuracy (0.4), security (0.3), readability (0.2), cross-modal consistency (0.1);

[0089] For users who spend a long time looking at a symptom description, the weights are adjusted to accuracy (0.5), security (0.35), readability (0.1), and cross-modal consistency (0.05).

[0090] High user facial expression confusion: The weights were further adjusted to accuracy (0.6), security (0.3), readability (0.05), and cross-modal consistency (0.05);

[0091] Final assessment score: 0.68.

[0092] Educational content scenario: When generating teaching materials, the system dynamically adjusts weights by monitoring user behavior (such as repeatedly viewing a knowledge point or scrolling slowly).

[0093] Basic weights: readability (0.4), logical consistency (0.3), information integrity (0.2), cross-modal consistency (0.1);

[0094] When a user repeatedly views a certain knowledge point, the weights are adjusted to readability (0.35), logical consistency (0.4), information integrity (0.15), and cross-modal consistency (0.1).

[0095] Slow user scrolling speed: The weights were further adjusted to readability (0.5), logical consistency (0.25), information integrity (0.15), and cross-modal coordination (0.1);

[0096] Final assessment score: 0.92.

[0097] The multimodal large language model evaluation method provided in this embodiment first obtains the user behavior signal corresponding to the multimodal large language model to be evaluated; then, based on the user behavior signal, the user behavior pattern is determined; next, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are determined; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators in the preset multimodal evaluation dimension indicators; and the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted based on the user behavior signal to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern; finally, the multimodal large language model to be evaluated is evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern to obtain the evaluation score of the multimodal large language model to be evaluated. Based on this scheme, the evaluation of multimodal large language models can be realized.

[0098] In one possible implementation, determining the user behavior pattern based on the user behavior signal includes:

[0099] The user behavior signal is input into the user behavior pattern prediction model. The user behavior pattern prediction model performs feature processing on the user behavior signal to obtain a user behavior feature vector, and predicts the user behavior pattern based on the user behavior feature vector.

[0100] In this application, the user behavior pattern prediction model includes a user behavior feature processing network and a prediction network. The user behavior feature processing network can be designed as a lightweight feature processing network (such as MobileNet, BiLSTM, etc.), which efficiently encodes and fuses user behavior signals to obtain user behavior feature vectors. The prediction network can implement prediction based on an attention mechanism. In one possible implementation, the prediction network can include an attention mechanism layer and a fully connected layer. The fully connected layer outputs the prediction results of each user behavior pattern. The probability of each user behavior pattern can be obtained by normalizing it to the 0-1 range using Softmax. Based on the probabilities of each user behavior pattern, the user behavior pattern corresponding to the user behavior signal can be determined.

[0101] After feature processing of user behavior signals, each time step (e.g., per second, per interaction event, or fixed time window) corresponds to a feature vector, which is used to characterize click density, dwell time, scrolling speed, etc. The user behavior feature vector is the feature matrix of the time series, and its shape is usually [T,D], where T is the number of time steps (i.e., the length of the time series), and D is the number of feature dimensions for each time step.

[0102] For ease of understanding, let's take an example and assume that we record the user's behavior in four consecutive windows on a page, with each window lasting 5 seconds. The feature dimension for each time step is 7. Then, the user behavior feature vector, a feature matrix of [4,7], can be represented as follows:

[0103] [0.4,2.1,80,150,0.9,2,0.1],#First 5-second window;

[0104] [0.1,0.5,-20,30,0.6,1,0.3],#Second 5-second window;

[0105] [0.6,3.0,200,300,0.8,3,0.5],#The third 5-second window;

[0106] [0.0,1.2,0,10,0.4,1,0.6],#The fourth 5-second window.

[0107] In one possible implementation, the preset multimodal evaluation dimension indicators include text modality evaluation dimension indicators, image modality evaluation dimension indicators, table modality evaluation dimension indicators, and cross-modal coordination evaluation dimension indicators.

[0108] In this application, the multimodal evaluation dimensions not only cover the basic evaluation dimensions of each modality, but also place special emphasis on cross-modal consistency, which is key to the evaluation of multimodal large language models. By systematically organizing these evaluation dimensions, a foundation can be laid for subsequent weight mapping and dynamic adjustment.

[0109] In this application, the specific content of the text modality evaluation dimension metrics, image / chart modality evaluation dimension metrics, table modality evaluation dimension metrics, and cross-modal compatibility evaluation dimension metrics can be set based on the actual scenario. This application does not impose any limitations. For ease of understanding, the following examples are provided:

[0110] Text modality assessment dimensions and metrics include:

[0111] Textual accuracy is used to characterize the correctness of facts and statistical results (such as the accuracy of medical diagnoses).

[0112] Textual consistency is used to characterize logical self-consistency and whether there are contradictions between the preceding and following texts (such as the logical coherence of financial reports).

[0113] Text security is used to characterize the presence of harmful or misleading content (such as the compliance of financial advice).

[0114] Text readability is used to characterize sentence complexity and summary clarity (such as the comprehensibility of educational content).

[0115] Image modality evaluation dimensions and metrics include:

[0116] Visual realism is used to characterize the plausibility of details in a generated image (such as the clarity of medical images).

[0117] Information readability is used to characterize the clarity and color contrast of images or charts (such as the readability of financial charts).

[0118] Trend matching degree is used to characterize whether the text description is consistent with the trend of the chart (such as the visualization of economic forecasts).

[0119] Labeling accuracy is used to characterize the location of key data points and numerical labeling errors (such as the accuracy of scientific charts).

[0120] The table-based modal assessment dimensions and metrics include:

[0121] Data integrity is used to characterize the handling of missing values ​​and the consistency of units (such as the completeness of financial statement tables).

[0122] Formatting standards are used to characterize row and column alignment and border rationality (such as the standardization of legal document forms).

[0123] Numerical accuracy is used to characterize the degree of matching between the calculation results and the original data (such as the accuracy of medical data tables).

[0124] Statistical reliability is used to characterize the reliability of data source labeling and confidence interval descriptions (such as the reliability of research report tables).

[0125] Cross-modal compatibility assessment dimensions and metrics include:

[0126] Semantic consistency is used to characterize the semantic matching degree between text and images / charts (such as CLIPScore).

[0127] Information complementarity is used to characterize the details covered by images in addition to text (such as the combination of text and images in product introductions).

[0128] Conflict detection capability, used to identify and label cross-modal contradictions (such as contradictions between charts and text values ​​in medical reports).

[0129] Multimodal fusion naturalness is used to characterize the smoothness of modality switching (e.g., " Figure 1 Display…and the correlation with the corresponding image).

[0130] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0131] Obtain the preset mapping relationship between user behavior signals and evaluation dimension indicator weight adjustment rules; select the target evaluation dimension indicator weight adjustment rule that has a mapping relationship with the user behavior signals from the preset mapping relationship between user behavior signals and evaluation dimension indicator weight adjustment rules; adjust the basic weight of the evaluation dimension indicator corresponding to the user behavior pattern using the target evaluation dimension indicator weight adjustment rule to obtain the new weight of the evaluation dimension corresponding to the user behavior pattern.

[0132] In this application, a mapping mechanism can be established between user behavior signals and the weight adjustment rules of evaluation dimension indicators.

[0133] In one possible implementation, specific user behavior signals can be directly mapped to the weight adjustments of the corresponding evaluation dimension indicators by predefined evaluation dimension indicator weight adjustment rules.

[0134] For ease of understanding, this application provides examples of mapping mechanisms between user behavior signals and evaluation dimension indicator weight adjustment rules, as shown in the table below:

[0135] User behavior signals Triggering conditions Evaluation Dimension Indicator Weight Adjustment Rules Example application scenarios Click on the chart area Click-through rate > 60% "Trend Matching Degree" Weight += α Revenue charts that users focus on in financial analysis reports Voice negative feedback Emotional score < 0.3 "Text accuracy" weight * = 1.5 User denies symptom description in medical diagnosis High level of confusion in facial expression Confusion level > 0.7 and dwell time > 5 seconds "Textual consistency" weight += β Users are confused about the derivation process in educational content. Long-term text segment Average dwell time > 3 seconds and accounting for > 60% "Text readability" weight += γ Users repeatedly review the terms in legal documents

[0136] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0137] Based on cross-modal consistency, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0138] For multimodal large language models, cross-modal consistency includes text-image consistency, text-table consistency, and image-table consistency.

[0139] In this application, when different modalities are highly consistent, the weight of the corresponding evaluation dimension index is increased; conversely, the weight is appropriately weakened. When different modalities show significant contradictions or inconsistencies on a certain evaluation dimension index, a negative weight adjustment term is applied to that evaluation dimension index to weaken its influence and prevent erroneous information from dominating decision-making.

[0140] In this application, the applicable target cross-modal consistency can be determined based on user behavior signals, and the basic weights of the evaluation dimensions corresponding to the user behavior patterns can be adjusted based on the target cross-modal consistency to obtain new weights for the evaluation dimensions corresponding to the user behavior patterns.

[0141] In one possible implementation, adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on cross-modal consistency to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes:

[0142] S201: Calculate the cross-modal consistency score, which is used to indicate the consistency between the modal element pointed to by the user behavior signal and other relevant modal information;

[0143] In this application, cross-modal consistency scores include text-image consistency scores, text-table consistency scores, and image-table consistency scores.

[0144] In this application, a differentiated calculation method is used to calculate the cross-modal consistency score for different modal combinations, wherein:

[0145] Text-image consistency scores can be calculated using the CLIP model to determine semantic similarity, with the following formula:

[0146] ;

[0147] in, and A visual and text encoder for CLIP.

[0148] The accuracy of text table consistency scores can be verified using statistical methods, such as:

[0149] ;

[0150] Image-table consistency scores can be obtained through coordinate matching and numerical verification, for example:

[0151] ;

[0152] S202: Based on the cross-modal consistency score, adjust the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0153] In this application, the basic weights of the evaluation dimensions corresponding to the user behavior patterns can be adjusted with reference to the following formula:

[0154] ;

[0155] in, : New weights for evaluation dimension indicator i corresponding to user behavior patterns;

[0156] The basic weights of the evaluation dimension indicator i corresponding to the user behavior pattern;

[0157] Consistency enhancement factor (default 0.15, configurable);

[0158] : Conflict penalty coefficient (default 0.1, configurable);

[0159] : Modal elements (such as clicked chart data points) that user behavior information points to;

[0160] : Confidence level of other relevant modal information;

[0161] It should be noted that the consistency enhancement coefficient can be configured based on scenario requirements to improve scenario adaptability.

[0162] For example:

[0163] In the financial context: α1 = 0.2 (emphasizing accuracy);

[0164] In the medical setting: α1 = 0.3 (emphasizing safety and accuracy);

[0165] Educational scenario: α1=0.1 (emphasizing readability and logic);

[0166] Entertainment scenario: α1=0.05 (emphasizing attractiveness).

[0167] It should also be noted that conflict penalties can be set based on scenario requirements, and this application does not impose any restrictions on this.

[0168] In one possible implementation, the evaluation of the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, to obtain the evaluation score of the multimodal large language model to be evaluated, includes:

[0169] The evaluation dimension indicators corresponding to the user behavior pattern are scored to obtain the evaluation dimension indicator scores corresponding to the user behavior pattern; the preset cross-modal consistency reward base score and consistency reward coefficient are obtained; based on the evaluation dimension indicator scores corresponding to the user behavior pattern, the cross-modal consistency reward base score and the consistency reward coefficient, the evaluation score of the multimodal large language model to be evaluated is obtained.

[0170] In this application, the scores of each evaluation dimension indicator can be weighted and summed based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern to obtain the evaluation dimension indicator score; then, the cross-modal consistency reward score can be calculated using the preset cross-modal consistency reward base score and consistency reward coefficient; finally, the evaluation dimension indicator score and the cross-modal consistency reward score can be summed to obtain the evaluation score of the multimodal large language model to be evaluated.

[0171] For ease of understanding, the evaluation score of the multimodal large language model to be evaluated can be calculated using the following formula:

[0172] ;

[0173] in:

[0174] : Sub-score (range 0-1) of evaluation dimension indicator i corresponding to user behavior pattern;

[0175] Cross-modal consistency reward base score (range 0-0.2);

[0176] Consistency reward coefficient (dynamically adjusted according to the scenario).

[0177] The above describes a multimodal large language model evaluation method provided by the embodiments of this application. The following describes the apparatus for performing the above multimodal large language model evaluation method.

[0178] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a multimodal large language model evaluation device provided in an embodiment of this application. Figure 2 As shown, the multimodal large language model evaluation device includes:

[0179] Acquisition unit 11 is used to acquire user behavior signals corresponding to the multimodal large language model to be evaluated;

[0180] User behavior pattern determination unit 12 is used to determine user behavior patterns based on the user behavior signals;

[0181] The basic weight determination unit 13 is used to determine the basic weight of the evaluation dimension index corresponding to the user behavior pattern; the evaluation dimension index corresponding to the user behavior pattern is a plurality of evaluation dimension indexes in the preset multimodal evaluation dimension indexes.

[0182] The weight adjustment unit 14 is used to adjust the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal, so as to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0183] Evaluation unit 15 is used to evaluate the multimodal large language model to be evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, and obtain the evaluation score of the multimodal large language model to be evaluated.

[0184] In one possible implementation, the user behavior pattern determination unit is specifically used for:

[0185] The user behavior signal is input into the user behavior pattern prediction model. The user behavior pattern prediction model performs feature processing on the user behavior signal to obtain a user behavior feature vector, and predicts the user behavior pattern based on the user behavior feature vector.

[0186] In one possible implementation, the preset multimodal evaluation dimension indicators include text modality evaluation dimension indicators, image modality evaluation dimension indicators, table modality evaluation dimension indicators, and cross-modal coordination evaluation dimension indicators.

[0187] In one possible implementation, the weight adjustment unit is specifically used for:

[0188] Obtain the mapping relationship between preset user behavior signals and evaluation dimension indicator weight adjustment rules;

[0189] Select the target evaluation dimension indicator weight adjustment rule that has a mapping relationship with the user behavior signal from the preset mapping relationship between user behavior signal and evaluation dimension indicator weight adjustment rule;

[0190] The basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted using the target evaluation dimension indicator weight adjustment rules to obtain the new weights of the evaluation dimension corresponding to the user behavior pattern.

[0191] In one possible implementation, the weight adjustment unit is specifically used for:

[0192] Based on cross-modal consistency, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0193] In one possible implementation, the weight adjustment unit is specifically used for:

[0194] Calculate a cross-modal consistency score, which is used to indicate the consistency between the modal element pointed to by the user behavior signal and other relevant modal information;

[0195] Based on the cross-modal consistency score, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

[0196] In one possible implementation, the evaluation unit is specifically used for:

[0197] The evaluation dimension indicators corresponding to the user behavior pattern are scored to obtain the scores of the evaluation dimension indicators corresponding to the user behavior pattern.

[0198] Obtain the preset cross-modal consistency reward base score and consistency reward coefficient;

[0199] The evaluation score of the multimodal large language model to be evaluated is obtained based on the evaluation dimension index score corresponding to the user behavior pattern, the basic score of the cross-modal consistency reward, and the consistency reward coefficient.

[0200] Each unit in the aforementioned multimodal large language model evaluation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each unit.

[0201] This application also provides an electronic device in its embodiments. (See reference...) Figure 3 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0202] like Figure 3 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0203] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0204] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the multimodal large language model evaluation methods provided in this application.

[0205] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the multimodal large language model evaluation methods provided in this application.

[0206] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0208] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0209] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for evaluating multimodal large language models, characterized in that, include: Obtain user behavior signals corresponding to the multimodal large language model to be evaluated; Based on the user behavior signals, determine the user behavior pattern; Determine the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern; the evaluation dimension indicators corresponding to the user behavior pattern are multiple evaluation dimension indicators in the preset multimodal evaluation dimension indicators. Based on the user behavior signals, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern. The multimodal language model to be evaluated is evaluated based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, and the evaluation score of the multimodal language model to be evaluated is obtained.

2. The method according to claim 1, characterized in that, The step of determining user behavior patterns based on the user behavior signals includes: The user behavior signal is input into the user behavior pattern prediction model. The user behavior pattern prediction model performs feature processing on the user behavior signal to obtain a user behavior feature vector, and predicts the user behavior pattern based on the user behavior feature vector.

3. The method according to claim 1, characterized in that, The preset multimodal evaluation dimensions include text modality evaluation dimensions, image modality evaluation dimensions, table modality evaluation dimensions, and cross-modal coordination evaluation dimensions.

4. The method according to claim 1, characterized in that, The step of adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes: Obtain the mapping relationship between preset user behavior signals and evaluation dimension indicator weight adjustment rules; Select the target evaluation dimension indicator weight adjustment rule that has a mapping relationship with the user behavior signal from the preset mapping relationship between user behavior signal and evaluation dimension indicator weight adjustment rule; The basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted using the target evaluation dimension indicator weight adjustment rules to obtain the new weights of the evaluation dimension corresponding to the user behavior pattern.

5. The method according to claim 1, characterized in that, The step of adjusting the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on the user behavior signal to obtain new weights for the evaluation dimension indicators corresponding to the user behavior pattern includes: Based on cross-modal consistency, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

6. The method according to claim 5, characterized in that, The adjustment of the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern based on cross-modal consistency to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern includes: Calculate a cross-modal consistency score, which is used to indicate the consistency between the modal element pointed to by the user behavior signal and other relevant modal information; Based on the cross-modal consistency score, the basic weights of the evaluation dimension indicators corresponding to the user behavior pattern are adjusted to obtain the new weights of the evaluation dimension indicators corresponding to the user behavior pattern.

7. The method according to claim 1, characterized in that, The evaluation of the multimodal large language model to be evaluated is performed based on the new weights of the evaluation dimension indicators corresponding to the user behavior pattern, resulting in an evaluation score for the multimodal large language model to be evaluated, including: The evaluation dimension indicators corresponding to the user behavior pattern are scored to obtain the scores of the evaluation dimension indicators corresponding to the user behavior pattern. Obtain the preset cross-modal consistency reward base score and consistency reward coefficient; The evaluation score of the multimodal large language model to be evaluated is obtained based on the evaluation dimension index score corresponding to the user behavior pattern, the basic score of the cross-modal consistency reward, and the consistency reward coefficient.

8. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the multimodal large language model evaluation method as described in any one of claims 1 to 7.

9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the multimodal large language model evaluation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the multimodal large language model evaluation method as described in any one of claims 1 to 7.