A multi-modal data quality evaluation method based on an agent model mapping relationship

By constructing a mapping relationship model, a mapping from the performance gain of the surrogate model to the performance gain of the target model is established, solving the problem of converting the evaluation results of the surrogate model to the actual performance gain of the target model. This achieves efficient and accurate multimodal data quality assessment, which is applicable to standardized quality evaluation of artificial intelligence training datasets such as text, images, and audio.

CN122365015APending Publication Date: 2026-07-10YUNZENG TECHNOLOGY (JIANGSU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNZENG TECHNOLOGY (JIANGSU) CO LTD
Filing Date
2026-06-03
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In third-party certification assessment scenarios, existing technologies have failed to effectively establish a mapping relationship between the assessment results of proxy models and the actual performance gains of target models, resulting in low accuracy of data quality assessment, high assessment costs, and an inability to meet the core requirements of certification assessment.

Method used

We construct a general calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm. We collect paired data of surrogate model performance gain and target model performance gain, establish a model of the actual mapping relationship from surrogate model performance gain to target model performance gain, evaluate multimodal datasets through lightweight surrogate models, calculate the equivalent target model performance gain prediction value, and finally determine the certification level.

Benefits of technology

Significantly reduce the computational power consumption and time cost of assessment, improve the accuracy and credibility of assessment conclusions, realize efficient and large-scale operation of certification assessment business, adapt to potential changes in target model or data distribution, and ensure the traceability and reproducibility of assessment conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365015A_ABST
    Figure CN122365015A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal data quality assessment method based on a surrogate model mapping relationship. By pre-establishing an actual mapping relationship model from the performance gain of the surrogate model to the performance gain of the target model, this invention, when actually assessing the quality of a specific multimodal dataset, uses a lightweight surrogate model to replace the target model to obtain the predicted performance gain value of the target model. Based on the actual mapping relationship model, the predicted performance gain value of the target model is converted into an equivalent predicted performance gain value of the target model. This significantly reduces the computational power and time costs of the assessment. Furthermore, through the calibration of the actual mapping relationship model, the output equivalent predicted performance gain value of the target model can accurately approximate the true target model gain. Thus, while ensuring the accuracy of the assessment conclusions, this method achieves efficient and scalable operation of certification assessment services, and is applicable to business scenarios involving standardized quality evaluation of various artificial intelligence training datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data quality assessment technology, and specifically to a multimodal data quality assessment method based on proxy model mapping relationships. Background Technology

[0002] In the development of large-scale language models and multimodal artificial intelligence models, the quality of pre-training data has a decisive impact on the final performance of the model. In third-party certification and evaluation scenarios, certification bodies need to conduct quality evaluations on the multimodal training datasets submitted by clients to determine whether the dataset can effectively improve the performance of the target artificial intelligence model. However, certification bodies face real-world constraints such as limited time, computing power, and venue resources in actual operation, making it difficult to conduct complete large-scale model training and evaluation for each submitted dataset. How to efficiently and reliably complete dataset quality evaluation under limited resource conditions has become a core technical challenge in certification and evaluation business. Currently, using lightweight proxy models to replace large models for data evaluation is an important approach to reduce evaluation costs.

[0003] For example, the Meta-rater method proposed by the Shanghai Artificial Intelligence Laboratory is aimed at the internal data screening scenario of large model training teams. This method first scores the quality of the dataset from multiple dimensions, and then trains a small surrogate model to fit the regression relationship between the weight combination of different quality score dimensions and the validation loss of the surrogate model, thereby searching for the optimal score weight combination. Its experiments show that this method can effectively screen high-quality training data and improve the training efficiency of large models. However, this method does not establish a mapping relationship between the evaluation results of the surrogate model and the actual performance gain of the target model, and cannot directly output the core indicators required for certification evaluation.

[0004] Secondly, a Chinese invention patent with publication number CN121303389B, titled "A Method for Measuring and Screening Training Data of Large Models that Balances Generality and Security Performance," also addresses the internal data screening scenario for large model training teams. It uses multiple datasets to be evaluated to supervise and fine-tune a large language model, assesses the fine-tuned model on general performance and security performance benchmarks, and calculates a comprehensive quality score as a label. Then, it extracts multi-dimensional feature vectors from each dataset as input to train a machine learning model as a proxy model. For new datasets, only feature vectors need to be extracted to predict their comprehensive quality score. However, this method establishes a mapping relationship between "dataset feature vectors" and "comprehensive quality score," and the proxy model outputs the predicted quality score of the dataset. It does not establish a conversion relationship between the proxy model's evaluation results and the actual performance gain of the target large model.

[0005] It is evident that current data quality assessment methods, when applied in actual third-party certification assessment scenarios, fail to establish a mapping relationship from the performance gain of the proxy model to the performance gain of the target model. They also lack a mechanism to convert the assessment results of the proxy model into predicted values ​​of the performance gain of the target model. This results in low accuracy of data quality assessment, high computational and time costs, and an inability to meet the core needs of certification assessment scenarios.

[0006] Therefore, it is necessary to invent a multimodal data quality assessment method based on the mapping relationship of the proxy model to solve the above problems. Summary of the Invention

[0007] The purpose of this invention is to provide a multimodal data quality assessment method based on the mapping relationship of the proxy model. In the context of certification assessment, this method can clearly establish the actual mapping relationship model from the performance gain of the proxy model to the performance gain of the target model. This reduces the computational power consumption of the assessment while ensuring that the certification conclusion has a verifiable scientific basis, thus addressing the aforementioned shortcomings in the technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multimodal data quality assessment method based on proxy model mapping relationships, comprising the following steps:

[0009] Step 1: Construct a general calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm;

[0010] Step 2: Acquire paired data of the performance gain of the proxy model and the performance gain of the target model through dual-channel acquisition to obtain the initial mapping calibration dataset;

[0011] Step 3: Based on the initial mapping calibration dataset, establish a model of the actual mapping relationship from the performance gain of the surrogate model to the performance gain of the target model, and output the corresponding predicted value of the performance gain of the target model.

[0012] Step 4: Perform a lightweight surrogate model evaluation on the multimodal dataset to be evaluated, and calculate the predicted performance gain of the equivalent target model;

[0013] Step 5: Compare the predicted performance gain values ​​of the equivalent target model with the certification level threshold to determine the certification level of the evaluation dataset;

[0014] Step 6: As the platform continues to operate and accumulates new matching data, the actual mapping relationship model is updated for subsequent evaluation.

[0015] The aforementioned multimodal data quality assessment method based on proxy model mapping relationships, in step 1, constructs a general calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm. The specific steps are as follows:

[0016] 1.1 Construct a general calibration dataset pool The specific formula is as follows:

[0017] ;

[0018] in, To calibrate the first in the dataset pool A sample dataset, To calibrate the index number of the dataset; The total number of sample datasets, And satisfy , This is the preset minimum sample size threshold;

[0019] 1.2 Determine the standardized anchoring evaluation task set system The specific formula is as follows:

[0020] ;

[0021] in, Each set of evaluation tasks corresponds to a pre-defined training paradigm, as follows:

[0022] The evaluation task set corresponding to the autoregressive language modeling paradigm is selected from the public knowledge question answering evaluation task set; The evaluation task set corresponding to the semantic embedding contrastive learning paradigm is selected from the publicly available semantic similarity evaluation task set; The evaluation task set corresponding to the discriminative classification or sequence labeling paradigm is selected from the public classification or sequence labeling evaluation task set; The evaluation task set corresponding to the cross-modal contrastive learning paradigm is selected from the publicly available cross-modal understanding evaluation task set; The evaluation task set corresponding to the edge-side reasoning paradigm is selected from the publicly available edge-side standard test task set; For evaluation task sets corresponding to other paradigms besides the five mentioned above, a set of public test tasks matching the corresponding paradigm is selected.

[0023] 1.3. Based on the target model being evaluated Sample dataset and standardized anchoring evaluation task sets Define a general performance gain measurement paradigm The specific formula is as follows:

[0024] ;

[0025] For from A standardized set of anchored evaluation tasks is provided, which satisfies the following conditions: it is publicly available and accessible, allowing any third party to fully reproduce the evaluation process based on publicly available information; the evaluation indicators are numerical, and the evaluation results are expressed in real number form; and ;

[0026] Before using the sample dataset, the target model is evaluated on a standardized anchoring evaluation task set. The benchmark performance index value is returned as a real number. After using the sample dataset, the target model is evaluated on the standardized anchored evaluation task set. The benchmark performance index value is returned as a real number.

[0027] The aforementioned multimodal data quality assessment method based on surrogate model mapping relationships, in step 2, involves dual-channel acquisition of paired data of surrogate model performance gain and target model performance gain to obtain an initial mapping calibration dataset. The specific steps are as follows:

[0028] 2.1 For the target model Let its training paradigm be The specific formula is as follows:

[0029] ;

[0030] 2.2. Standardizing the evaluation task set system Selection and training paradigm Corresponding evaluation task set And according to the training paradigm Choose a lightweight proxy model ;

[0031] 2.3. Calibration Data Pool Each sample dataset The performance gain measurements for the proxy model channel and the target model channel were performed separately, as follows:

[0032] Computational proxy model through sample dataset Performance gain of proxy model The specific formula is as follows:

[0033] ;

[0034] The target model is computed using the sample dataset. Target model performance gain The specific formula is as follows:

[0035] ;

[0036] 2.4 Pair the performance gain values ​​of the surrogate model and the target model to obtain a set of paired data. traverse all From a sample dataset, we obtain the initial mapping calibration dataset. The specific formula is as follows:

[0037] ;

[0038] in, Each pair of data in the dataset comes from the same sample dataset and the same evaluation task set.

[0039] The aforementioned multimodal data quality assessment method based on the mapping relationship of the surrogate model, in step 3, establishes an actual mapping relationship model from the performance gain of the surrogate model to the performance gain of the target model based on the initial mapping calibration dataset, and outputs the corresponding predicted value of the performance gain of the target model. The specific steps are as follows:

[0040] 3.1. Based on the predicted values ​​of the equivalent target model Based on this, establish an initial mapping relationship model. The specific formula is as follows:

[0041] ;

[0042] Among them, the equivalent target model predicted value It is for the sample dataset The predicted values ​​are used for training or validation;

[0043] 3.2. Assume an initial mapping relationship model. Selected from model space Through the initial mapping relationship model The actual mapping relationship model is calculated. The specific formula is as follows:

[0044] ;

[0045] in, For model space Any candidate model or candidate function in the list; To solve for the objective functional Operators for candidate models that obtain the minimum value;

[0046] 3.3. The determined actual mapping relationship model The actual mapping relationship model is solidified within the authentication platform and can be invoked, resulting in performance gains for the proxy model with any new input. , Output the corresponding target model performance gain prediction value The specific formula is as follows:

[0047] .

[0048] The aforementioned multimodal data quality assessment method based on surrogate model mapping relationships, in step 4, performs a lightweight surrogate model evaluation on the multimodal dataset to be evaluated and calculates the equivalent target model performance gain prediction value. The specific steps are as follows:

[0049] 4.1. Use the evaluation task set Performance gain measurement paradigm Measuring lightweight agent models exist Performance gain value The specific formula is as follows:

[0050] ;

[0051] in, The multimodal dataset to be evaluated submitted by the client; and The value is a real number;

[0052] 4.2, will Input actual mapping relationship model In the process, the predicted performance gain of the equivalent target model is calculated. The specific formula is as follows:

[0053] ;

[0054] Among them, the predicted performance gain of the equivalent target model It is for the dataset to be evaluated The predicted value is used for final authentication.

[0055] In the aforementioned multimodal data quality assessment method based on surrogate model mapping relationships, step 4.2 involves obtaining information simultaneously when the actual mapping relationship model supports outputting prediction uncertainty information. exist The confidence interval below is denoted as ;

[0056] in, The preset confidence level, and ; This is the lower bound of the confidence interval. This is the upper bound of the confidence interval. , All are real numbers.

[0057] The aforementioned multimodal data quality assessment method based on proxy model mapping relationships determines the authentication level of the assessment dataset in step 5 by comparing the predicted performance gain value of the equivalent target model with the authentication level threshold. The specific steps are as follows:

[0058] 5.1. Preset a set of authentication level thresholds, specifically denoted as... , ,..., And satisfy ;

[0059] in, The total number of certification levels, and , ; It is a set of positive integers;

[0060] It is the first Each certification level threshold, , Furthermore, the certification level thresholds can be updated according to industry standards;

[0061] 5.2. Predict the performance gain value of the equivalent target model The certification level is determined by comparing it with the certification level threshold and following these rules. The details are as follows:

[0062] ;

[0063] in, ;

[0064] The index number is the threshold interval, and ;

[0065] 5.3 When step 4 outputs the confidence interval simultaneously In this case, the certification level will be determined according to the following rules. The details are as follows:

[0066] .

[0067] In the aforementioned multimodal data quality assessment method based on the proxy model mapping relationship, step 6 involves updating the actual mapping relationship model as the platform continues to operate and accumulates new paired data for subsequent evaluation. The specific steps are as follows:

[0068] 6.1. Initial mapping calibration dataset Record Calculate the current mapping calibration dataset The specific formula is as follows:

[0069] ;

[0070] in, The counting symbol for the number of elements in a set; This indicates the total number of paired data sets contained in the initially established mapping calibration dataset;

[0071] 6.2. Let the cumulative number of newly added paired data groups be... Based on the current mapping calibration dataset To determine whether the actual mapping relationship model needs to be updated, follow these steps:

[0072] When the cumulative number of newly added paired data groups meets the requirement At that time, the actual mapping relationship model is updated;

[0073] in, ; A single-element set containing only zero; The preset update ratio threshold, and ;

[0074] This represents the size of the current mapping calibration dataset before this update, initially set to... ;

[0075] 6.3 After the actual mapping relationship model is triggered to update, the following steps are executed:

[0076] 6.3.1 Merge the newly added paired data sets with the current mapping calibration dataset to form the expanded target calibration dataset. The specific formula is as follows:

[0077] ;

[0078] in, This is the version number after this update, and Meanwhile, the initial, unupdated version number was Corresponding to the initial mapping calibration dataset ;

[0079] For the current mapping calibration dataset before updating the actual mapping relationship model, when At the same time, it also serves as the initial mapping calibration dataset. ;

[0080] The set consisting of newly added paired data contains Group data;

[0081] 6.3.2 Based on the target calibration dataset Repeat step 3 to obtain the updated actual mapping relationship model. ;

[0082] 6.3.3 Adopting a new actual mapping relationship model Replace the original actual mapping relationship model This is for subsequent evaluation.

[0083] In the aforementioned multimodal data quality assessment method based on proxy model mapping relationships, step 6.3 involves the platform recording the version traceability chain of the mapping relationship and recording relevant information for each version, specifically including:

[0084] Version number Initial version After each update Incrementing; version update timestamp recorded The size of the target calibration dataset is denoted as . ,and .

[0085] Compared with the prior art, the beneficial effects of the present invention are:

[0086] 1. This invention establishes a pre-built actual mapping relationship model from the performance gain of the proxy model to the performance gain of the target model. When performing quality assessment on a specific multimodal dataset, a lightweight proxy model is used to replace the target model to obtain the predicted performance gain value of the target model. Based on the actual mapping relationship model, the predicted performance gain value of the target model is converted into an equivalent predicted performance gain value of the target model. This significantly reduces the computational power consumption and time cost of the assessment. Furthermore, through the calibration of the actual mapping relationship model, the output equivalent predicted performance gain value of the target model can accurately approximate the real target model gain. Thus, while ensuring the accuracy of the assessment conclusions, this invention achieves efficient and large-scale operation of the certification assessment business.

[0087] 2. This invention determines the final certification level of the dataset by comparing the predicted performance gain of the equivalent target model or the lower bound of the confidence interval with the certification level threshold, thereby enhancing the credibility of the evaluation conclusion. By recording version traceability information, the actual mapping relationship model can continuously optimize its prediction accuracy as business data accumulates and adapt to potential changes in the target model or data distribution, ensuring the long-term effectiveness of the evaluation framework. This provides the evaluation conclusion with traceable and reproducible scientific evidence and is applicable to business scenarios that conduct standardized quality evaluation of various artificial intelligence training datasets such as text, images, and audio. Attached Figure Description

[0088] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0089] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0090] This invention provides the following: Figure 1 The multimodal data quality assessment method shown includes the following steps: (The method is based on a proxy model mapping relationship.)

[0091] Step 1: Construct a general calibration dataset pool Standardized Anchored Evaluation Task Set System and performance gain measurement paradigm The specific steps are as follows:

[0092] 1.1 Construct a general calibration dataset pool The specific formula is as follows:

[0093] ;

[0094] in, To calibrate the first in the dataset pool A sample dataset, To calibrate the index number of the dataset; The total number of sample datasets, And satisfy , The minimum sample size threshold is set; each sample dataset covers at least one data modality, including text, image, audio, and hybrid modalities, and the modality types can be expanded as new modalities emerge; each sample dataset is associated with a known quality level label. ,and , The preset highest quality level is used; the calibration dataset pool has sufficient coverage of sample datasets of different quality levels.

[0095] 1.2 Determine the standardized anchoring evaluation task set system The specific formula is as follows:

[0096] ;

[0097] in, Each set of evaluation tasks corresponds to a pre-defined training paradigm, as follows:

[0098] The evaluation task set corresponding to the autoregressive language modeling paradigm is selected from the public knowledge question answering evaluation task set; The evaluation task set corresponding to the semantic embedding contrastive learning paradigm is selected from the publicly available semantic similarity evaluation task set; The evaluation task set corresponding to the discriminative classification or sequence labeling paradigm is selected from the public classification or sequence labeling evaluation task set; The evaluation task set corresponding to the cross-modal contrastive learning paradigm is selected from the publicly available cross-modal understanding evaluation task set; The evaluation task set corresponding to the edge-side reasoning paradigm is selected from the publicly available edge-side standard test task set; For evaluation task sets corresponding to other paradigms besides the five mentioned above, a set of public test tasks matching the corresponding paradigm is selected.

[0099] 1.3. Based on the target model being evaluated Sample dataset and standardized anchoring evaluation task sets Define a general performance gain measurement paradigm The specific formula is as follows:

[0100] ;

[0101] For from A standardized set of anchored evaluation tasks is provided, which satisfies the following conditions: it is publicly available and accessible, allowing any third party to fully reproduce the evaluation process based on publicly available information; the evaluation indicators are numerical, and the evaluation results are expressed in real number form; and ;

[0102] Before using the sample dataset, the target model is evaluated on a standardized anchoring evaluation task set. The benchmark performance index value is returned as a real number.

[0103] After using the sample dataset, the target model is evaluated on the standardized anchored evaluation task set. The benchmark performance index value is returned as a real number.

[0104] The target model uses sample datasets in the following ways: lightweight fine-tuning (for...) use Perform parameter updates), context learning (and) The example in the text serves as context input. ) or zero-sample evaluation (not updated) Parameters, direct measurement exist Performance on related tasks), and other ways of using sample datasets for various target models.

[0105] In step 1, by constructing a calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm, a reproducible and comparable foundation is provided for the subsequent collection of paired data on the performance gains of the surrogate model and the target model. This ensures that the actual mapping relationship model can be established in the future. It can be based on the same measurement rules, avoiding systematic biases introduced by inconsistencies in data sources, evaluation tasks, or gain calculation methods.

[0106] Step 2: Performance gain of dual-channel acquisition proxy model Performance gain of the target model Pairing data The initial mapping calibration dataset is obtained. The specific steps are as follows:

[0107] 2.1 For the target model Let its training paradigm be The specific formula is as follows:

[0108] ;

[0109] 2.2. Standardizing the evaluation task set system Selection and training paradigm Corresponding evaluation task set And according to the training paradigm Choose a lightweight proxy model ;

[0110] in, Furthermore, the lightweight proxy model satisfies:

[0111] The training task type and loss function form belong to the same training paradigm category as the target model;

[0112] The number of parameters is less than that of the target model, i.e. ,in For the parameter functions of the target model and the lightweight surrogate model;

[0113] 2.3. Calibration Data Pool Each sample dataset The performance gain measurements for the proxy model channel and the target model channel were performed separately, as follows:

[0114] Computational proxy model through sample dataset Performance gain of proxy model The specific formula is as follows:

[0115] ;

[0116] The target model is computed using the sample dataset. Target model performance gain The specific formula is as follows:

[0117] ;

[0118] 2.4 Pair the performance gain values ​​of the surrogate model and the target model to obtain a set of paired data. traverse all From a sample dataset, we obtain the initial mapping calibration dataset. The specific formula is as follows:

[0119] ;

[0120] in, Each pair of data in the dataset comes from the same sample dataset and the same evaluation task set, providing a foundation for establishing the mapping relationship between the performance gain of the surrogate model and the performance gain of the target model.

[0121] In step 2, the performance gain of the proxy model is acquired through dual-channel acquisition. Paired data with the target model's performance gain form an initial mapping calibration dataset, which can provide supervised paired training samples for the subsequent establishment of actual mapping relationship models. Moreover, each pair of paired data comes from the same sample dataset and the same evaluation task set, ensuring the consistency and reliability of mapping learning and avoiding confusion caused by different data sources or evaluation tasks.

[0122] Step 3: Calibrate the dataset based on the initial mapping Establish performance gains from the proxy model Performance gain of the target model Actual mapping relationship model Output the corresponding target model performance gain prediction value. The specific steps are as follows:

[0123] 3.1. Based on the predicted values ​​of the equivalent target model Based on this, establish an initial mapping relationship model. The specific formula is as follows:

[0124] ;

[0125] Among them, the equivalent target model predicted value It is for the sample dataset The predicted values ​​are used for training or validation;

[0126] When establishing the initial mapping relationship model, the paired data of each group in the initial mapping calibration dataset should satisfy the following:

[0127] a. The performance gains of the surrogate model and the target model are derived from the same sample dataset. ;

[0128] b. The performance gains of the surrogate model and the target model are compared on the same standardized anchored evaluation task set. The above measurements were obtained;

[0129] c. The performance gain of the surrogate model and the performance gain of the target model are measured using the same performance gain paradigm. Calculated;

[0130] 3.2. Assume an initial mapping relationship model. Selected from model space Through the initial mapping relationship model The actual mapping relationship model is calculated. The specific formula is as follows:

[0131] ;

[0132] in, To from model space The instance of the optimal initial mapping relationship model selected in the process, i.e., the actual mapping relationship model; For model space Any candidate model or candidate function in the list; To solve for the objective functional Operators for candidate models that obtain the minimum value;

[0133] Among them, the objective functional For In calibration dataset The accuracy of the up-mapping includes mean squared error, mean absolute error, or other scalar functions that measure prediction bias;

[0134] The model space is the set of all mapping models or functions that can learn the correspondence between input and output from paired data. Examples of initial mapping relationship models include regression models and neural network models.

[0135] 3.3. The determined actual mapping relationship model The actual mapping relationship model is solidified within the authentication platform and can be invoked, resulting in performance gains for the proxy model with any new input. , Output the corresponding target model performance gain prediction value The specific formula is as follows:

[0136] ;

[0137] in, no It is not a subscripted variable in the text, but an unsubscripted general representation of the variable (proxy model performance gain). The purpose is to emphasize that the actual mapping relationship model described in step 3.3 can accept proxy model performance gain inputs from any source, and the two are related as "type" and "instance".

[0138] In step 3, an actual mapping relationship model from the performance gain of the proxy model to the performance gain of the target model is established based on the initial mapping calibration dataset. This actual mapping relationship model is then embedded within the certification platform. This allows the evaluation of any subsequent new sample dataset to be obtained simply by running a lightweight proxy model. The equivalent target model performance gain can be quickly predicted, thereby obtaining an evaluation result close to the real target model gain with extremely low computing power consumption.

[0139] Step 4: Perform lightweight surrogate model evaluation on the multimodal dataset to be evaluated, and calculate the predicted performance gain of the equivalent target model. The specific steps are as follows:

[0140] 4.1. Use the evaluation task set Performance gain measurement paradigm Measuring lightweight agent models exist Performance gain value The specific formula is as follows:

[0141] ;

[0142] in, The multimodal dataset to be evaluated submitted to the client, i.e., the dataset to be evaluated, covers at least one data modality; and the The value is a real number.

[0143] 4.2, will Input actual mapping relationship model In the process, the predicted performance gain of the equivalent target model is calculated. The specific formula is as follows:

[0144] ;

[0145] Among them, the performance gain prediction value of the equivalent target model It is for the dataset to be evaluated The predicted value is used for final authentication;

[0146] Specifically, when the actual mapping relationship model supports outputting prediction uncertainty information, it simultaneously obtains... exist The confidence interval below is denoted as ;

[0147] in, The preset confidence level, and ; This is the lower bound of the confidence interval. This is the upper bound of the confidence interval. , All are real numbers:

[0148] In step 4, the lightweight proxy model is used to evaluate the dataset to be evaluated, and the measured performance gain value is input into the fixed actual mapping relationship model to directly calculate the equivalent target model's performance gain prediction value. This avoids the time-consuming complete training and evaluation using the target model. The computational power consumption and time cost of each evaluation are significantly lower than those of the lightweight proxy model. The results are comparable to, and far lower than, the full-scale large-scale model evaluation. At the same time, through the calibration of the actual mapping relationship model, the output equivalent target model performance gain prediction value can closely approximate the real target model gain with high accuracy. Thus, while ensuring the accuracy of the evaluation conclusion, the certification evaluation business has achieved efficient and large-scale operation.

[0149] Step 5: Predict performance gain based on the equivalent target model The authentication level of the evaluation dataset is determined by comparing it with an authentication level threshold. The specific steps are as follows:

[0150] 5.1. Preset a set of authentication level thresholds, specifically denoted as... , ,..., And satisfy ;

[0151] in, The total number of certification levels, and , ; It is a set of positive integers;

[0152] It is the first Each certification level threshold, , Furthermore, the certification level thresholds can be updated according to industry standards;

[0153] 5.2. Predict the performance gain value of the equivalent target model The certification level is determined by comparing it with the certification level threshold and following these rules. The details are as follows:

[0154] ;

[0155] in, ;

[0156] The index number is the threshold interval, and ;

[0157] 5.3 When step 4 outputs the confidence interval simultaneously In this case, the certification level will be determined according to the following rules. The details are as follows:

[0158] ;

[0159] Therefore, at the preset confidence level The actual performance gain of the dataset is no less than [amount missing]. The evaluation dataset must reach at least the grade. :

[0160] In step 5, the final certification level of the evaluation dataset is determined by comparing the predicted value of the performance gain of the equivalent target model or the lower bound of the confidence interval with the certification level threshold. This transforms the continuous predicted value into a discrete, interpretable certification conclusion. The lower bound of the confidence interval is used for conservative judgment, which ensures the authenticity and reliability of the certification level at a given confidence level and enhances the credibility of the evaluation conclusion.

[0161] Step 6: As the platform continues to operate and accumulates new matching data, the actual mapping relationship model is updated. To perform the update, follow these steps:

[0162] 6.1. Initial mapping calibration dataset Record Calculate the current mapping calibration dataset The specific formula is as follows:

[0163] ;

[0164] in, The counting symbol for the number of elements in a set; This indicates the total number of paired data sets contained in the initially established mapping calibration dataset;

[0165] Among them, with the development of certification business, the actual performance gain value reported by customers after actually training the target model, together with the measured gain of the corresponding proxy model, constitutes new paired data;

[0166] 6.2. Let the cumulative number of newly added paired data groups be... Based on the current mapping calibration dataset To determine whether the actual mapping relationship model needs to be updated, follow these steps:

[0167] When the cumulative number of newly added paired data groups meets the requirement At that time, the actual mapping relationship model is updated;

[0168] in, ; A single-element set containing only zero; The preset update ratio threshold, and This indicates the percentage of the initial mapping calibration dataset size at which the update of the actual mapping relationship model is initiated;

[0169] This represents the size of the current mapping calibration dataset before this update, initially set to... ;

[0170] 6.3 After the actual mapping relationship model is triggered to update, the following steps are executed:

[0171] 6.3.1 Merge the newly added paired data sets with the current mapping calibration dataset to form the expanded target calibration dataset. The specific formula is as follows:

[0172] ;

[0173] in, This is the version number after this update, and Meanwhile, the initial, unupdated version number was Corresponding to the initial mapping calibration dataset ;

[0174] For the current mapping calibration dataset before updating the actual mapping relationship model, when At the same time, it also serves as the initial mapping calibration dataset. ;

[0175] The set consisting of newly added paired data contains Group data;

[0176] 6.3.2 Based on the target calibration dataset Repeat step 3 to obtain the updated actual mapping relationship model. ;

[0177] 6.3.3 Adopting a new actual mapping relationship model Replace the original actual mapping relationship model For subsequent evaluation;

[0178] The platform records a version traceability chain of mapping relationships, recording relevant information for each version, specifically including:

[0179] (1) Version number Initial version After each update Increasing;

[0180] (2) Version update timestamp recorded as ;

[0181] (3) The size of the target calibration dataset is denoted as ,and ;

[0182] In step 6, an adaptive update mechanism for accumulating newly added paired data groups is introduced. When the triggering condition is met, the number of newly added paired data groups is merged, and step 3 is re-executed to update the new actual mapping relationship model. At the same time, version traceability information is recorded, so that the actual mapping relationship model can continuously optimize its prediction accuracy as business data accumulates and adapt to potential changes in the target model or data distribution, ensuring the long-term effectiveness of the evaluation framework. Meanwhile, traceable version management provides a basis for the reproduction and verification of certification conclusions.

[0183] Meanwhile, when there are significant changes to the version of the target model or when the standardized anchored evaluation task set is updated, the certification platform proactively triggers the recalibration of the mapping relationship model to ensure that the evaluation framework keeps pace with technological changes.

[0184] Verification Experiment

[0185] To verify the technical effectiveness of this invention, the following specific experimental data illustrates the advantages of this method over existing technologies in terms of evaluation accuracy, computing power consumption, and time cost:

[0186] I. Experimental Setup

[0187] Calibration Dataset Pool: 20 text datasets of known quality levels were collected, ranging in size from 500M tokens to 38B tokens, covering three levels: high quality (academic literature, Wikipedia), medium quality (news corpus, forum posts), and low quality (noisy web page text). Among them, the high-quality datasets averaged about 30B tokens in size, and full training required approximately 480 to 550 hours. Each dataset was labeled with performance gain values ​​obtained from actual testing using a fully trained 7B parameter large language model as ground truth values.

[0188] Proxy model: A lightweight decoder architecture model with 18M parameters is selected, with the number of parameters being approximately one-four-hundredth of the target large model.

[0189] Target Model: The open-source large language model with 7B parameters is selected as the representative target model.

[0190] Standardized anchored assessment task set: The open knowledge question-and-answer assessment task set is selected, which includes three sub-items: language comprehension, common sense reasoning, and reading comprehension.

[0191] Comparison method:

[0192] Comparative Example 1 (Direct Evaluation of Naked Proxy Model): Only the proxy model is used to train and evaluate the dataset to be evaluated. The performance gain value of the proxy model is directly used as the authentication conclusion output without establishing any mapping relationship.

[0193] Comparative Example 2 (Full Large Model Evaluation): The target large model is fully trained and evaluated for each dataset to be evaluated, and the measured performance gain is used as the certification conclusion. This method yields the most accurate results but has the highest computational cost.

[0194] The method of this invention is as follows: Steps 1 to 6 are executed. First, a mapping relationship model is established using 15 sample datasets, and then the model is used to evaluate the remaining 5 datasets.

[0195] II. Experimental Results

[0196] (1) Comparison of assessment accuracy

[0197] The evaluation results of the five test datasets were compared with the actual performance gain values ​​obtained by training the full large model. The average prediction error of each method was calculated, and the experimental data were recorded in Table 1.

[0198] Table 1

[0199]

[0200] As shown in Table 1, the results directly evaluated by the naked proxy model deviate significantly from the actual values, with a maximum error of 1.68. However, this invention reduces the average error by about two-thirds by pre-establishing a mapping relationship model, and increases the correlation coefficient between the predicted and actual values ​​from 0.61 to 0.87, significantly improving the accuracy of the evaluation conclusions.

[0201] (2) Comparison of computing power consumption

[0202] The specific experimental data on computing power consumption are shown in Table 2;

[0203] Table 2

[0204]

[0205] As shown in Table 2, the computational power consumption of the method of this invention for a single evaluation is only about one two-thousandth of that for a full large model evaluation. The one-time investment mainly includes dual-channel acquisition of proxy models and target models for 15 calibration datasets. After calibration, each subsequent certification evaluation only requires very low computational power. When the number of evaluations exceeds 5, this method begins to save total computational power. The more evaluations, the more significant the advantage. In actual certification operations, the number of annual evaluations can reach dozens or even hundreds, and the total computational power saved is extremely considerable.

[0206] (3) Comparison of time costs

[0207] Using a typical hardware configuration of a certification body (a single 8-GPU server) as the test environment, the time required to complete one dataset evaluation was statistically analyzed, and the specific results are shown in Table 3.

[0208] Table 3

[0209]

[0210] As shown in Table 3, the full-scale large-scale model evaluation requires continuous operation for more than three weeks (about 21 days). In business scenarios where certification bodies need to evaluate multiple customer datasets, this time cost is completely impractical. However, the present invention only requires about half a day for a single evaluation, and can complete the evaluation and issue a certification report within the service period promised by the certification body, which has a decisive advantage from a business operation perspective.

[0211] (4) Discrimination ability of datasets of different quality levels

[0212] Five test datasets were grouped according to their actual quality levels to observe whether the present invention could effectively distinguish datasets of different quality levels.

[0213] Experimental results show that the method of this invention provides equivalent gain predictions for high-quality datasets in the higher range, predictions for low-quality datasets in the lower range, and predictions for medium-quality datasets in between. Furthermore, the confidence interval width is positively correlated with the degree of dataset quality fluctuation. This indicates that the present invention can effectively identify the true quality differences in datasets and will not lose its discriminative ability due to the use of a proxy model.

[0214] III. Experimental Conclusions

[0215] The above experimental comparisons have verified that the present invention has the following technical effects:

[0216] First, the evaluation accuracy is significantly better than that of the naked proxy model. The predicted values ​​are highly correlated with the actual performance gain of the target large model. The mapping relationship model effectively makes up for the evaluation bias between the proxy model and the large model.

[0217] Second, the computational power and time consumption of a single evaluation are comparable to those of a bare proxy model evaluation, and far lower than those of a full-scale large-scale model evaluation. The one-time investment in establishing the mapping relationship can be quickly amortized in large-scale operations, making the commercial operation of the certification service feasible.

[0218] Third, this method can effectively distinguish datasets of different quality levels, and the reliability of the conclusions can be quantitatively assessed through confidence intervals, providing technical assurance for certification bodies to issue authoritative conclusions to the public.

[0219] In summary, this invention establishes a pre-defined mapping relationship model from the performance gain of the proxy model to the performance gain of the target model. When performing quality assessment on specific multimodal datasets, a lightweight proxy model is used to replace the target model to obtain the predicted performance gain of the target model. Based on the actual mapping relationship model, the predicted performance gain of the target model is converted into an equivalent predicted performance gain of the target model. This significantly reduces the computational power consumption and time cost of the assessment. Furthermore, through the calibration of the actual mapping relationship model, the output equivalent predicted performance gain of the target model can accurately approximate the real target model gain. Thus, while ensuring the accuracy of the assessment conclusions, this invention achieves efficient and large-scale operation of the certification assessment business.

[0220] By comparing the predicted performance gain of the equivalent target model or the lower bound of the confidence interval with the certification level threshold, the final certification level of the dataset is determined, which enhances the credibility of the evaluation conclusion. By recording version traceability information, the actual mapping relationship model can continuously optimize its prediction accuracy as business data accumulates and adapt to potential changes in the target model or data distribution, ensuring the long-term effectiveness of the evaluation framework. This provides the evaluation conclusion with traceable and reproducible scientific evidence, and is suitable for business scenarios that conduct standardized quality evaluation of various artificial intelligence training datasets such as text, images, and audio.

[0221] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used above are only some embodiments described in this invention. Obviously, those skilled in the art can obtain other drawings based on these drawings.

[0222] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A multimodal data quality assessment method based on surrogate model mapping relationships, characterized in that: Includes the following steps: Step 1: Construct a general calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm; Step 2: Acquire paired data of the performance gain of the proxy model and the performance gain of the target model through dual-channel acquisition to obtain the initial mapping calibration dataset; Step 3: Based on the initial mapping calibration dataset, establish a model of the actual mapping relationship from the performance gain of the surrogate model to the performance gain of the target model, and output the corresponding predicted value of the performance gain of the target model. Step 4: Perform a lightweight surrogate model evaluation on the multimodal dataset to be evaluated, and calculate the predicted performance gain of the equivalent target model; Step 5: Compare the predicted performance gain values ​​of the equivalent target model with the certification level threshold to determine the certification level of the evaluation dataset; Step 6: As the platform continues to operate and accumulates new matching data, the actual mapping relationship model is updated for subsequent evaluation.

2. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 1, characterized in that: In step 1, a general calibration dataset pool, a standardized anchoring evaluation task set system, and a performance gain measurement paradigm are constructed. The specific steps are as follows: 1.1 Construct a general calibration dataset pool The specific formula is as follows: ; in, To calibrate the first in the dataset pool A sample dataset, To calibrate the index number of the dataset; The total number of sample datasets, And satisfy , This is the preset minimum sample size threshold; 1.2 Determine the standardized anchoring evaluation task set system The specific formula is as follows: ; in, Each set of evaluation tasks corresponds to a pre-defined training paradigm, as follows: The evaluation task set corresponding to the autoregressive language modeling paradigm is selected from the public knowledge question answering evaluation task set; The evaluation task set corresponding to the semantic embedding contrastive learning paradigm is selected from the publicly available semantic similarity evaluation task set; The evaluation task set corresponding to the discriminative classification or sequence labeling paradigm is selected from the public classification or sequence labeling evaluation task set; The evaluation task set corresponding to the cross-modal contrastive learning paradigm is selected from the publicly available cross-modal understanding evaluation task set; The evaluation task set corresponding to the edge-side reasoning paradigm is selected from the publicly available edge-side standard test task set; For evaluation task sets corresponding to other paradigms besides the five mentioned above, a set of public test tasks matching the corresponding paradigm is selected. 1.

3. Based on the target model being evaluated Sample dataset and standardized anchoring evaluation task sets Define a general performance gain measurement paradigm The specific formula is as follows: ; For from A standardized set of anchored evaluation tasks is provided, which satisfies the following conditions: it is publicly available and accessible, allowing any third party to fully reproduce the evaluation process based on publicly available information; the evaluation indicators are numerical, and the evaluation results are expressed in real number form; and ; Before using the sample dataset, the target model is evaluated on a standardized anchoring evaluation task set. The benchmark performance index value is returned as a real number. After using the sample dataset, the target model is evaluated on the standardized anchored evaluation task set. The benchmark performance index value is returned as a real number.

3. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 1, characterized in that: In step 2, paired data of the performance gain of the surrogate model and the performance gain of the target model are acquired through dual-channel acquisition to obtain the initial mapping calibration dataset. The specific steps are as follows: 2.1 For the target model Let its training paradigm be The specific formula is as follows: ; 2.

2. Standardizing the evaluation task set system Selection and training paradigm Corresponding evaluation task set And according to the training paradigm Choose a lightweight proxy model ; 2.

3. Calibration Data Pool Each sample dataset The performance gain measurements for the proxy model channel and the target model channel were performed separately, as follows: Computational proxy model through sample dataset Performance gain of proxy model The specific formula is as follows: ; The target model is computed using the sample dataset. Target model performance gain The specific formula is as follows: ; 2.4 Pair the performance gain values ​​of the surrogate model and the target model to obtain a set of paired data. traverse all From a sample dataset, we obtain the initial mapping calibration dataset. The specific formula is as follows: ; in, Each pair of data in the dataset comes from the same sample dataset and the same evaluation task set.

4. The multimodal data quality assessment method based on surrogate model mapping relationship as described in claim 1, characterized in that: In step 3, based on the initial mapping calibration dataset, a model is established to represent the actual mapping relationship from the performance gain of the surrogate model to the performance gain of the target model, and the corresponding predicted value of the performance gain of the target model is output. The specific steps are as follows: 3.

1. Based on the predicted values ​​of the equivalent target model Based on this, establish an initial mapping relationship model. The specific formula is as follows: ; Among them, the equivalent target model predicted value It is for the sample dataset The predicted values ​​are used for training or validation; 3.

2. Assume an initial mapping relationship model. Selected from model space Through the initial mapping relationship model The actual mapping relationship model is calculated. The specific formula is as follows: ; in, For model space Any candidate model or candidate function in the list; To solve for the objective functional Operators for candidate models that obtain the minimum value; 3.

3. The determined actual mapping relationship model The actual mapping relationship model is solidified within the authentication platform and can be invoked, resulting in performance gains for the proxy model with any new input. , Output the corresponding target model performance gain prediction value The specific formula is as follows: 。 5. The multimodal data quality assessment method based on surrogate model mapping relationship as described in claim 1, characterized in that: In step 4, a lightweight surrogate model evaluation is performed on the multimodal dataset to be evaluated, and the predicted performance gain of the equivalent target model is calculated. The specific steps are as follows: 4.

1. Use the evaluation task set Performance gain measurement paradigm Measuring lightweight agent models exist Performance gain value The specific formula is as follows: ; in, The multimodal dataset to be evaluated submitted by the client; and The value is a real number; 4.2, will Input actual mapping relationship model In the process, the predicted performance gain of the equivalent target model is calculated. The specific formula is as follows: ; Among them, the predicted performance gain of the equivalent target model It is for the dataset to be evaluated The predicted value is used for final authentication.

6. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 5, characterized in that: In step 4.2, when the actual mapping relationship model supports outputting prediction uncertainty information, it simultaneously obtains... exist The confidence interval below is denoted as ; in, The preset confidence level, and ; This is the lower bound of the confidence interval. This is the upper bound of the confidence interval. , All are real numbers.

7. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 1, characterized in that: In step 5, the authentication level of the evaluation dataset is determined by comparing the predicted performance gain value of the equivalent target model with the authentication level threshold. The specific steps are as follows: 5.

1. Preset a set of authentication level thresholds, specifically denoted as... , ,..., And satisfy ; in, The total number of certification levels, and , ; It is a set of positive integers; It is the first Each certification level threshold, , Furthermore, the certification level thresholds can be updated according to industry standards; 5.

2. Predict the performance gain value of the equivalent target model The certification level is determined by comparing it with the certification level threshold and following these rules. The details are as follows: ; in, ; The index number is the threshold interval, and ; 5.3 When step 4 outputs the confidence interval simultaneously In this case, the certification level will be determined according to the following rules. The details are as follows: 。 8. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 1, characterized in that: In step 6, as the platform continues to operate and accumulates new pairing data, the actual mapping relationship model is updated for subsequent evaluation. The specific steps are as follows: 6.

1. Initial mapping calibration dataset Record Calculate the current mapping calibration dataset The specific formula is as follows: ; in, The counting symbol for the number of elements in a set; This indicates the total number of paired data sets contained in the initially established mapping calibration dataset; 6.

2. Let the cumulative number of newly added paired data groups be... Based on the current mapping calibration dataset To determine whether the actual mapping relationship model needs to be updated, follow these steps: When the cumulative number of newly added paired data groups meets the requirement At that time, the actual mapping relationship model is updated; in, ; A single-element set containing only zero; The preset update ratio threshold, and ; This represents the size of the current mapping calibration dataset before this update, initially set to... ; 6.3 After the actual mapping relationship model is triggered to update, the following steps are executed: 6.3.1 Merge the newly added paired data sets with the current mapping calibration dataset to form the expanded target calibration dataset. The specific formula is as follows: ; in, This is the version number after this update, and Meanwhile, the initial, unupdated version number was Corresponding to the initial mapping calibration dataset ; For the current mapping calibration dataset before updating the actual mapping relationship model, when At the same time, it also serves as the initial mapping calibration dataset. ; The set consisting of newly added paired data contains Group data; 6.3.2 Based on the target calibration dataset Repeat step 3 to obtain the updated actual mapping relationship model. ; 6.3.3 Adopting a new actual mapping relationship model Replace the original actual mapping relationship model This is for subsequent evaluation.

9. The multimodal data quality assessment method based on proxy model mapping relationship as described in claim 8, characterized in that: In step 6.3, the platform records the version traceability chain of the mapping relationship, recording relevant information for each version, specifically including: Version number Initial version After each update Incrementing; version update timestamp recorded The size of the target calibration dataset is denoted as . ,and .

Citation Information

Patent Citations

  • A general-purpose and security-performance-considered large model training data metric and screening method

    CN121303389B