Large model evaluation method and system, dynamic evaluation data set construction method and system thereof, medium and equipment

By constructing a parallel evaluation method for dynamic evaluation data sets, the randomness and incompleteness of gender fairness evaluation of large models are solved, and the true reliability and stability of gender fairness scores are achieved, which reduces application risks and improves the experience of choosing to use large models.

CN120163253AActive Publication Date: 2025-06-17DATA SPACE RES INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510340119.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-17
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The lack of clear definition of gender fairness indicators in the existing technology, and the evaluation data set is small and fixed, which leads to the incomplete and fairness of the evaluation results, and the evaluation process is random, making it difficult to reflect the actual performance of the big model in terms of gender fairness, and there are legal and regulatory and social moral risks.

Method used

Construct a dynamic evaluation data set, and use dynamic evaluation text data sets to generate image sets, perform data quality control, generate problem options, and use multiple dynamic evaluation data sets to perform parallel evaluation to calculate gender fairness scores.

Benefits of technology

An objective evaluation of gender fairness in big models is achieved, the authenticity and stability of gender fairness scores is ensured, the laws and regulations and social moral risks in the application process are reduced, and the experience of choosing to use big models is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163253A_ABST
    Figure CN120163253A_ABST
Patent Text Reader

Abstract

The invention discloses a large model evaluation method and system, a dynamic evaluation data set construction method and system thereof, a medium and equipment, and relates to the field of large model testing, and the method comprises the following steps: obtaining a dynamic evaluation text data set; performing graph generation on the dynamic evaluation text data set to obtain a dynamic evaluation picture set; performing data quality control on the dynamic evaluation picture set to obtain an intermediate dynamic evaluation picture set; and generating problem options for the pictures in the middle dynamic evaluation picture set to obtain a dynamic evaluation data set. The gender fairness score is ensured to be true and reliable. According to the method, the high quality and the randomness of the constructed dynamic evaluation data set are ensured, and the situation that the large model to be evaluated is subjected to targeted training according to the static evaluation data set or the evaluation result is unstable due to the randomness of the data and the model is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model evaluation technology, and in particular to a large model evaluation method and system, and a dynamic evaluation data set construction method and system, medium and equipment. Background Art

[0002] In today's society, gender equality constantly appears in various hot news. When the big model outputs answers based on user questions, it should also maintain basic gender equality, not give answers that are gender-discriminatory or gender-biased, and comply with the constraints of existing laws, regulations, and social moral systems. Therefore, evaluating the big model in terms of gender equality is one of the basic conditions to ensure that the big model can be widely used in reality.

[0003] However, the current methods for evaluating gender fairness of large models still have the following problems:

[0004] (1) Lack of clear definition of gender equity indicators;

[0005] (2) Evaluation datasets are generally manually screened and labeled, and the process of generating datasets consumes a lot of manpower and material resources. Therefore, the scale of evaluation datasets is generally small. Once the evaluation dataset is manually labeled, it will remain fixed during the evaluation of large models. This allows evaluation models that are trained and adjusted based on known fixed datasets to obtain higher false scores when being evaluated. However, limited datasets have certain limitations in the breadth of evaluation, and the coverage of gender equality evaluation is not comprehensive enough, so the evaluation results are inevitably unfair.

[0006] (3) The evaluation process is random: Since the underlying logic of the large model is deep learning, its output is related to probability, and the output may be slightly different when the same data is input.

[0007] Therefore, a single evaluation on a fixed and limited data set is not sufficient to reflect the actual performance of the model in terms of gender fairness. The large models selected by researchers, companies or the general public based on this may run the risk of violating laws, regulations and social ethics during actual application. Summary of the invention

[0008] In order to solve the technical problems existing in the background technology, the present invention proposes a large model evaluation method and system and a dynamic evaluation data set construction method and system, medium and equipment.

[0009] In a first aspect, the present invention proposes a method for constructing a dynamic evaluation data set for large model evaluation, comprising:

[0010] Get dynamic evaluation text dataset;

[0011] Generate images from the dynamic evaluation text dataset to obtain a dynamic evaluation image set;

[0012] Perform data quality control on the dynamic evaluation image set to obtain an intermediate dynamic evaluation image set;

[0013] Generate question options for the images in the intermediate dynamic evaluation image set to obtain a dynamic evaluation dataset.

[0014] Preferably, before obtaining the dynamic evaluation text dataset, it further includes:

[0015] Construct a professional text library; where the professional text library includes multiple professions, and each profession includes two descriptive text lists, which respectively correspond to men and women; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the profession and gender information of the person;

[0016] Among them, obtaining the dynamic evaluation text dataset specifically includes: randomly selecting a certain number of professions from the professional text library; using all the descriptive texts of the selected professions as a dynamic evaluation text dataset.

[0017] Preferably, performing data quality control on the dynamic evaluation image set to obtain an intermediate dynamic evaluation image set specifically includes:

[0018] Input all the images in the dynamic evaluation image set into the referee model for data quality review;

[0019] During the data quality review process, when a certain image does not meet the review requirements of the referee model, delete the image from the corresponding dynamic evaluation image set;

[0020] When the referee model traverses all the images in a certain dynamic evaluation image set, it means that the data quality review of this dynamic evaluation image set has been completed, and use this dynamic evaluation image set that has completed the data quality review as the intermediate dynamic evaluation image set.

[0021] Preferably, each image in each dynamic evaluation dataset corresponds to four question options, one of which is the correct profession, and the other three options are similar professions with high, medium, and low similarities to the correct profession respectively.

[0022] In a second aspect, the present invention also proposes a dynamic evaluation dataset construction system for large model evaluation, including:

[0023] An acquisition module for obtaining a dynamic evaluation text dataset;

[0024] An image generation module for generating images from the dynamic evaluation text dataset to obtain a dynamic evaluation image set;

[0025] A data quality control module for performing data quality control on a dynamic evaluation picture set to obtain an intermediate dynamic evaluation picture set;

[0026] A question option generation module for generating question options for the pictures in the intermediate dynamic evaluation picture set to obtain a dynamic evaluation data set.

[0027] Thirdly, a large model evaluation method proposed by the present invention includes:

[0028] Constructing a plurality of dynamic evaluation data sets by using the dynamic evaluation data set construction method described in any item of the first aspect;

[0029] Parallelly evaluating the large model to be evaluated by using a plurality of dynamic evaluation data sets to obtain a plurality of answer result sets;

[0030] Calculating a gender fairness score of the large model according to a plurality of answer result sets.

[0031] Preferably, parallelly evaluating the large model to be evaluated by using a plurality of dynamic evaluation data sets to obtain a plurality of answer result sets, specifically including:

[0032] Deploying the large model to be evaluated into a plurality of large model replicas;

[0033] Inputting a plurality of dynamic evaluation data sets into a plurality of large model replicas one by one for evaluation to obtain a plurality of answer result sets; wherein, the evaluations of the plurality of large model replicas are executed in parallel.

[0034] Preferably, each answer result set includes the options and answer result scores of the occupations of the characters in each picture in the corresponding dynamic evaluation data set; wherein, the answer result scores of the four options are S1, S2, S3, and S4 in the order of the correct occupation, high similarity, medium similarity, and low similarity, and S1 > S2 > S3 > S4.

[0035] Preferably,

[0036] In the formula, Score represents the gender fairness score of the large model; score i represents the gender fairness score of the large model under the i-th dynamic evaluation data set, i = 1, 2,..., n; n represents the number of dynamic evaluation data sets.

[0037] Preferably,

[0038] In the formula, score i represents the gender fairness score of the large model under the i-th dynamic evaluation data set, i = 1, 2,..., n; n represents the number of dynamic evaluation data sets; m irepresents the number of occupations randomly selected from the i-th evaluation dataset; question i,j,男 represents the score of the large model's answer to the male question corresponding to the j-th occupation in the i-th dynamic evaluation dataset; question i,j,女 is the score of the large model's answer to the female question corresponding to the j-th occupation in the i-th dynamic evaluation dataset.

[0039] Preferably, after calculating the gender fairness score of the large model according to multiple answer result sets, it further includes:

[0040] Calculating the evaluation stability of the large model according to the gender fairness score.

[0041] Preferably,

[0042] In the formula, represents the evaluation stability, n represents the number of dynamic evaluation datasets, score i represents the gender fairness score of the large model under the i-th dynamic evaluation dataset, μ represents the average value of the gender fairness scores of the large model under all dynamic evaluation datasets, and i = 1, 2,..., n.

[0043] Fourthly, the present invention also proposes a large model evaluation system, including:

[0044] A dataset construction module for constructing a plurality of dynamic evaluation datasets by using the dynamic evaluation dataset construction method described in any item of the first aspect;

[0045] An evaluation module for parallelly evaluating the large model to be evaluated by using a plurality of dynamic evaluation datasets to obtain a plurality of answer result sets;

[0046] A calculation module for calculating the gender fairness score of the large model according to a plurality of answer result sets.

[0047] Fifthly, the present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the dynamic evaluation dataset construction method described in any item of the first aspect.

[0048] Sixthly, the present invention also proposes an electronic device, including: a processor and a memory, the memory is used to store one or more programs; when one or more programs are executed by the processor, it implements the dynamic evaluation dataset construction method described in any item of the first aspect.

[0049] In a seventh aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the large model evaluation method described in any one of the third aspects is implemented.

[0050] In an eighth aspect, the present invention further provides an electronic device, including: a processor and a memory, where the memory is used to store one or more programs; when the one or more programs are executed by the processor, the large model evaluation method described in any one of the third aspects is implemented.

[0051] In the present invention, the proposed large model evaluation method and system, as well as its dynamic evaluation dataset construction method and system, medium and device, obtain a dynamic evaluation text dataset, dynamically generate a dynamic evaluation picture dataset according to the dynamic evaluation text dataset, and perform data quality control on the dynamic evaluation picture dataset respectively to obtain an intermediate dynamic evaluation picture dataset; generate question options for the pictures in the intermediate dynamic evaluation picture dataset to obtain a dynamic evaluation dataset. The present invention ensures the high quality and randomness of the constructed dynamic evaluation dataset, effectively avoids the targeted training of the large model to be evaluated according to the static evaluation dataset or the instability of the gender fairness score caused by the randomness of data and the model, thereby ensuring the true and reliable gender fairness score. Moreover, the present invention performs multiple rounds of parallel evaluation on the large model to be evaluated through multiple dynamic evaluation datasets, obtains multiple answer result sets, and then calculates the gender fairness score of the large model according to the multiple answer result sets, realizing an objective evaluation of the large model in terms of gender fairness, reducing the randomness of the gender fairness score while significantly shortening the evaluation time.

[0052] The present invention realizes an objective evaluation of the large model in terms of gender fairness, facilitating subsequent selection and use of suitable large models by researchers, companies or the general public in actual application scenarios, effectively reducing the risk of violating laws, regulations and social ethics during the application process of the large model, and effectively improving the experience of researchers, companies or the general public in the process of selecting and using large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic flowchart of the dynamic evaluation dataset construction method in an embodiment proposed by the present invention.

[0054] Figure 2 It is a schematic diagram of parallel evaluation in an embodiment proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0056] In a first aspect, as Figure 1 shown, the present invention proposes a method for constructing a dynamic evaluation dataset for large model evaluation, including:

[0057] Obtain a dynamic evaluation text dataset;

[0058] Generate images from the dynamic evaluation text dataset to obtain a dynamic evaluation image dataset;

[0059] Conduct data quality control on the dynamic evaluation image dataset to obtain an intermediate dynamic evaluation image dataset;

[0060] Generate question options for the images in the intermediate dynamic evaluation image dataset to obtain a dynamic evaluation dataset.

[0061] By obtaining a dynamic evaluation text dataset, dynamically generating a dynamic evaluation image dataset based on the dynamic evaluation text dataset, and respectively conducting data quality control on the dynamic evaluation image dataset to obtain an intermediate dynamic evaluation image dataset; generating question options for the images in the intermediate dynamic evaluation image dataset to obtain a dynamic evaluation dataset, the present invention ensures the high quality and randomness of the constructed multiple dynamic evaluation datasets, effectively avoiding the subsequent large model to be evaluated from conducting targeted training based on a static evaluation dataset or the instability of the gender fairness score caused by the randomness of data and the model, thereby ensuring the true reliability of the gender fairness score.

[0062] In order to obtain a dynamic evaluation text dataset, in this embodiment, before obtaining the dynamic evaluation text dataset, it further includes: constructing a professional text library; wherein, the professional text library includes multiple occupations, and each occupation includes two descriptive text lists, which respectively correspond to men and women; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the person;

[0063] Among them, obtaining the dynamic evaluation text dataset specifically includes: randomly selecting a certain number of occupations from the professional text library; using all the descriptive texts of the selected occupations as a dynamic evaluation text dataset.

[0064] Among them, the data format of the professional text library is as follows.

[0065]

[0066]

[0067] In this embodiment, generating images from the dynamic evaluation text dataset to obtain a dynamic evaluation image dataset specifically includes: inputting all the descriptive texts of all occupations in the dynamic evaluation text dataset into an image generation model to obtain a dynamic evaluation image dataset.

[0068] In this embodiment, the text-to-image model dynamically generates a dynamic evaluation image set according to the dynamic evaluation text data set, ensuring the randomness of the subsequent obtained dynamic evaluation data set, and effectively avoiding the targeted training of the large model to be evaluated according to the static evaluation data set or the instability of the gender fairness score caused by the randomness of the data and the model. Thus, the authenticity and reliability of the gender fairness score are effectively guaranteed.

[0069] In one specific embodiment, the text-to-image model is the stable-diffusion3 model.

[0070] In another specific embodiment, the text-to-image model is the Janus-Pro multimodal large model.

[0071] In another specific embodiment, the text-to-image model is the Emu3-Gen model.

[0072] In this embodiment, data quality control is performed on the dynamic evaluation image set to obtain an intermediate dynamic evaluation image set, which specifically includes:

[0073] All images in the dynamic evaluation image set are input into the referee model for data quality review;

[0074] During the data quality review process, when a certain image does not meet the review requirements of the referee model, the image is deleted from the dynamic evaluation image set to which it belongs;

[0075] When the referee model traverses all images of a dynamic evaluation image set, the dynamic evaluation image set has completed data quality review, and the dynamic evaluation image set that has completed data quality review is used as the intermediate dynamic evaluation image set.

[0076] This embodiment uses the referee model to control the quality of the generated data, which can ensure the randomness of the subsequent dynamic evaluation data set while having high quality, and further avoid the targeted training of the large model according to the static evaluation data set or the instability of the gender fairness score caused by the randomness of the data and the model. Thus, the authenticity and reliability of the gender fairness score are further guaranteed.

[0077] In one specific embodiment, the referee model is the Grounded-Segment-Anything model.

[0078] In specific implementation, all images in the dynamic evaluation image set are input into the Grounded-Segment-Anything model for data quality review;

[0079] During the data quality review process, if the Grounded-Segment-Anything model detects that there are multiple people in a certain picture, then the picture will be deleted from the dynamic evaluation picture set to which it belongs;

[0080] When the Grounded-Segment-Anything model traverses all the pictures in a certain dynamic evaluation picture set, then the data quality review of this dynamic evaluation picture set has been completed;

[0081] Take this dynamic evaluation picture set that has completed the data quality review as the intermediate dynamic evaluation picture set.

[0082] In another specific embodiment, the referee model is the Qianwen large model.

[0083] In specific implementation, all the pictures in the dynamic evaluation picture set are input into the Qianwen large model for data quality review;

[0084] If the Qianwen large model identifies that the gender and occupation of a person in a certain picture are inconsistent with the gender and occupation information in the corresponding descriptive text, then the picture will be deleted from the dynamic evaluation picture set to which it belongs;

[0085] When the Qianwen large model traverses all the pictures in a certain dynamic evaluation picture set, then the data quality review of this dynamic evaluation picture set has been completed;

[0086] Take this dynamic evaluation picture set that has completed the data quality review as the intermediate dynamic evaluation picture set.

[0087] In another embodiment, the referee model is the Grounded-Segment-Anything model and the Qianwen large model. Among them, the Qianwen large model is responsible for judging whether the gender and occupation of the person in the picture are consistent with the gender and occupation information in the corresponding descriptive text, and the Grounded-Segment-Anything is responsible for judging whether the number of people in the picture exceeds one.

[0088] In specific implementation, first input all the pictures in the dynamic evaluation picture set into the Grounded-Segment-Anything model for data quality review, and then input the pictures in the dynamic evaluation picture set that have passed the review by the Grounded-Segment-Anything model into the Qianwen large model for data quality review, effectively improving the high quality of the intermediate dynamic evaluation picture set and the subsequent dynamic evaluation data set.

[0089] In this embodiment, there are four question options corresponding to each picture in each dynamic evaluation data set, one of which is the correct occupation, and the other three options are similar occupations with high, medium, and low similarities to the correct occupation respectively.

[0090] In a second aspect, the present invention also provides a dynamic evaluation dataset construction system for large model evaluation, including:

[0091] An acquisition module, configured to acquire a dynamic evaluation text dataset;

[0092] An image generation from text module, configured to perform image generation from text on the dynamic evaluation text dataset to obtain a dynamic evaluation image set;

[0093] A data quality control module, configured to perform data quality control on the dynamic evaluation image set to obtain an intermediate dynamic evaluation image set;

[0094] A question option generation module, configured to generate question options for the images in the intermediate dynamic evaluation image set to obtain a dynamic evaluation dataset.

[0095] In this embodiment, during the process of acquiring the dynamic evaluation text dataset by the acquisition module, a certain number of occupations are randomly selected from a pre-constructed occupational text library;

[0096] All the descriptive texts of the selected occupations are used as a dynamic evaluation text dataset.

[0097] In this embodiment, during the process of image generation from text by the image generation from text module, all the descriptive texts of all the occupations in the dynamic evaluation text dataset are input into the image generation from text model to obtain a dynamic evaluation image set.

[0098] In this embodiment, during the process of data quality control by the data quality control module, all the images in the dynamic evaluation image set are input into the referee model for data quality review;

[0099] During the data quality review process, when a certain image does not meet the review requirements of the referee model, the image is deleted from the dynamic evaluation image set to which it belongs;

[0100] When the referee model traverses all the images of a certain dynamic evaluation image set, the data quality review of the dynamic evaluation image set is completed, and the dynamic evaluation image set that has completed the data quality review is used as the intermediate dynamic evaluation image set.

[0101] In this embodiment, each image in each dynamic evaluation dataset corresponds to four question options, one of which is the correct occupation, and the other three options are all similar occupations with high, medium, and low similarities to the correct occupation respectively.

[0102] In a third aspect, a large model evaluation method proposed by the present invention includes:

[0103] Acquire multiple dynamic evaluation text datasets;

[0104] Generate images from text for multiple dynamic evaluation text datasets respectively to obtain multiple dynamic evaluation image sets;

[0105] Perform data quality control on multiple dynamic evaluation image sets respectively to obtain multiple intermediate dynamic evaluation image sets;

[0106] Generate question options for the images in multiple intermediate dynamic evaluation image sets respectively to obtain multiple dynamic evaluation datasets;

[0107] Use multiple dynamic evaluation datasets to perform parallel evaluation on the large model to be evaluated to obtain multiple sets of answer results;

[0108] Calculate the gender fairness score of the large model based on multiple sets of answer results.

[0109] In the present invention, by obtaining multiple dynamic evaluation text datasets, dynamically generating multiple dynamic evaluation image sets according to the multiple dynamic evaluation text datasets, and performing data quality control on the multiple dynamic evaluation image sets respectively, the high quality and randomness of the multiple subsequent obtained dynamic evaluation datasets are ensured, effectively avoiding the targeted training of the large model to be evaluated according to static evaluation datasets or the instability of the gender fairness score caused by the randomness of data and the model, thereby ensuring the authenticity and reliability of the gender fairness score. Moreover, in the present invention, multiple dynamic evaluation datasets are used to perform multiple rounds of parallel evaluation on the large model to be evaluated to obtain multiple sets of answer results, and then the gender fairness score of the large model is calculated based on the multiple sets of answer results, realizing the objective evaluation of the large model in terms of gender fairness, reducing the randomness of the gender fairness score while greatly shortening the evaluation time.

[0110] A large model evaluation method proposed by the present invention realizes the objective evaluation of the large model in terms of gender fairness, facilitates researchers, companies or the general public to select and use appropriate large models in actual application scenarios, effectively reduces the risk of violating laws, regulations and social ethics during the application process of the large model, and effectively improves the experience of researchers, companies or the general public in the process of selecting and using large models.

[0111] In this embodiment, before obtaining multiple dynamic evaluation text datasets, it further includes: constructing a professional text library; wherein, the professional text library includes multiple occupations, and each occupation includes two descriptive text lists, which respectively correspond to men and women; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the person.

[0112] Among them, the data format of the professional text library is as follows.

[0113]

[0114] Therefore, obtaining multiple dynamic evaluation text datasets in this embodiment specifically includes: randomly selecting a certain number of occupations from the occupational text library; using all the descriptive texts of the selected occupations as the dynamic evaluation text datasets; repeating the above steps until multiple dynamic evaluation text datasets are obtained.

[0115] In this embodiment, generating images from text for multiple dynamic evaluation text datasets respectively to obtain multiple dynamic evaluation image sets specifically includes: inputting all the descriptive texts of all occupations in multiple dynamic evaluation text datasets into the image generation from text model respectively to obtain multiple dynamic evaluation image sets.

[0116] With such a setting in this embodiment, multiple dynamic evaluation image sets are dynamically generated according to multiple dynamic evaluation text datasets through the image generation from text model, ensuring the randomness of the subsequent obtained dynamic evaluation datasets, and effectively avoiding the targeted training of the large model to be evaluated according to static evaluation datasets or the instability of the gender fairness score caused by the randomness of data and the model, thus effectively ensuring the authenticity and reliability of the gender fairness score.

[0117] In one specific embodiment, the image generation from text model is the stable-diffusion3 model.

[0118] In another specific embodiment, the image generation from text model is the Janus-Pro multi-modal large model.

[0119] In another specific embodiment, the image generation from text model is the Emu3-Gen model.

[0120] In this embodiment, performing data quality control on multiple dynamic evaluation image sets respectively to obtain multiple intermediate dynamic evaluation image sets specifically includes:

[0121] Inputting all the images in multiple dynamic evaluation image sets into the referee model for data quality review respectively;

[0122] During the data quality review process, when a certain image does not meet the review requirements of the referee model, deleting the image from the dynamic evaluation image set to which it belongs;

[0123] When the referee model traverses all the images in a certain dynamic evaluation image set, the data quality review of this dynamic evaluation image set is completed;

[0124] Regarding this dynamic evaluation image set that has completed data quality review as the intermediate dynamic evaluation image set;

[0125] When the data quality reviews of multiple dynamic evaluation image sets are all completed, multiple intermediate dynamic evaluation image sets are obtained.

[0126] In this embodiment, the referee model is used to control the quality of the generated data, which can ensure high quality while guaranteeing the randomness of the subsequent dynamic evaluation data set, further avoiding the targeted training of the large model based on the static evaluation data set or the instability of the gender fairness score caused by the randomness of the data and the model, thus further ensuring the authenticity and reliability of the gender fairness score.

[0127] In one specific embodiment, the referee model is the Grounded-Segment-Anything model.

[0128] During specific implementation, all the pictures in multiple dynamic evaluation picture sets are respectively input into the Grounded-Segment-Anything model for data quality review;

[0129] During the data quality review process, if the Grounded-Segment-Anything model detects that there are multiple people in a certain picture, then this picture is deleted from the corresponding dynamic evaluation picture set;

[0130] When the Grounded-Segment-Anything model traverses all the pictures in a certain dynamic evaluation picture set, then this dynamic evaluation picture set has completed the data quality review;

[0131] The dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.

[0132] In another specific embodiment, the referee model is the Qianwen large model.

[0133] During specific implementation, all the pictures in multiple dynamic evaluation picture sets are respectively input into the Qianwen large model for data quality review;

[0134] If the Qianwen large model identifies that the gender and occupation of the person in a certain picture are inconsistent with the gender and occupation information in the corresponding descriptive text, then this picture is deleted from the corresponding dynamic evaluation picture set;

[0135] When the Qianwen large model traverses all the pictures in a certain dynamic evaluation picture set, then this dynamic evaluation picture set has completed the data quality review;

[0136] The dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.

[0137] In another embodiment, the referee model is the Grounded-Segment-Anything model and the Qianwen large model.

[0138] Among them, the Qianwen large model is responsible for judging whether the gender and occupation of the people in the picture are consistent with the gender and occupation information in the corresponding descriptive text, and Grounded-Segment-Anything is responsible for judging whether the number of people in the picture exceeds one.

[0139] In specific implementation, first, all the pictures in multiple dynamic evaluation picture sets are respectively input into the Grounded-Segment-Anything model for data quality review, and then the pictures in the multiple dynamic evaluation picture sets that pass the review by the Grounded-Segment-Anything model are respectively input into the Qianwen large model for data quality review, effectively improving the high quality of the intermediate dynamic evaluation picture sets and the subsequent dynamic evaluation data sets.

[0140] In this embodiment, multiple dynamic evaluation data sets are used to perform parallel evaluation on the large model to be evaluated, and multiple answer result sets are obtained, specifically including: deploying the large model to be evaluated into multiple large model replicas; inputting the multiple dynamic evaluation data sets into the multiple large model replicas one by one for evaluation to obtain multiple answer result sets; among them, the evaluations of the multiple large model replicas are executed in parallel.

[0141] As Figure 2 shown, the large model to be evaluated in this embodiment is deployed into multiple large model replicas, each large model replica exclusively occupies the GPU resources of a graphics card, and the number of large model replicas is equal to the number of groups of dynamic evaluation data sets. When performing evaluation, multiple parallel processes are created, and each process is assigned a dynamic evaluation data set and a large model replica. In the process, taking the assigned dynamic evaluation data set as the input, the interface of the large model replica is called to obtain the options for the answers of the large model to be evaluated for the occupation of the people in each picture.

[0142] This setting in this embodiment can evaluate multiple dynamic evaluation data sets, reduce the randomness of the gender fairness score, and greatly shorten the evaluation time at the same time, which is beneficial to improving the stability and accuracy of the evaluation.

[0143] In this embodiment, question options are respectively generated for multiple intermediate dynamic evaluation picture sets to obtain multiple dynamic evaluation data sets, specifically including:

[0144] Question options are respectively generated for all the pictures in multiple intermediate dynamic evaluation picture sets, so that there are four question options corresponding to each picture; among them, among the four options, one option is the correct occupation, and the other three options are similar occupations;

[0145] The multiple dynamic evaluation picture sets that have completed question option generation are used as multiple dynamic evaluation data sets.

[0146] In the process of generating question options, since each picture corresponds to a descriptive text and the correct occupation corresponding to the picture is known, only three similar occupations need to be generated.

[0147] In this embodiment, each occupation is converted into a multi-dimensional vector representation through a clip coding model, and the similarity between occupations can be expressed by the similarity of multi-dimensional vectors. In the process of generating question options, three occupations with high, medium, and low similarities are randomly selected as the three similar occupation options for the picture according to the similarity values between occupations.

[0148] That is to say, there are four question options corresponding to each picture in each dynamic evaluation dataset, one of which is the correct occupation, and the other three options are all similar occupations with high, medium, and low similarities to the correct occupation respectively.

[0149] Therefore, in order to facilitate the calculation of the gender fairness score of the large model, each answer result set in this embodiment includes the answer options and answer result scores of the occupations of the characters in each picture in the corresponding dynamic evaluation dataset;

[0150] Among them, the answer result scores of the four options are S1, S2, S3, and S4 in the order of the correct occupation, high similarity, medium similarity, and low similarity, and S1 > S2 > S3 > S4.

[0151] In this way, through the design of different score options, this embodiment can characterize the gender fairness degree of the large model to be evaluated under the corresponding occupation.

[0152] In this embodiment,

[0153] In the formula, Score represents the gender fairness score of the large model; score i represents the gender fairness score of the large model under the i-th dynamic evaluation dataset, i = 1, 2,..., n; n represents the number of dynamic evaluation datasets.

[0154] Among them,

[0155] In the formula, m i represents the number of occupations randomly selected in the i-th dynamic evaluation dataset; question i,j,男 represents the answer result score of the large model for the male question corresponding to the j-th occupation in the i-th dynamic evaluation dataset; question i,j,女 represents the answer result score of the large model for the female question corresponding to the j-th occupation in the i-th dynamic evaluation dataset.

[0156] This embodiment gives clear indicators for evaluating gender fairness, which is conducive to calculating the gender fairness score of the large model.

[0157]

[0158] In one specific embodiment, when the large model identifies the corresponding occupation of the person in the picture, the scores for the four options are different. In the order of the correct occupation, high similarity, medium similarity, and low similarity, the scores are 5 points, 3 points, 2 points, and 1 point respectively. For example, the data format of one dynamic evaluation dataset Q is as shown above. There is a male picture and a female picture for each occupation, and there are four occupation options with different scores under each picture.

[0159] Therefore, in this embodiment,

[0160] In the formula, score i represents the gender fairness score of the large model under the i-th dynamic evaluation dataset, m i represents the number of randomly selected occupations in the i-th evaluation dataset; question i,j,男 represents the score of the large model's answer to the male question corresponding to the j-th occupation in the i-th dynamic evaluation dataset; question i,j,女 is the score of the large model's answer to the female question corresponding to the j-th occupation in the i-th dynamic evaluation dataset.

[0161] It should be noted that question i,j,男 is one of 5 points, 3 points, 2 points, and 1 point; question i,j,女 is also one of 5 points, 3 points, 2 points, and 1 point.

[0162] In this embodiment, after calculating the gender fairness score of the large model according to multiple answer result sets, it further includes: calculating the evaluation stability of the large model according to the gender fairness score.

[0163] This embodiment calculates the evaluation stability of the large model according to the gender fairness score, which can objectively describe the stability of the evaluation process and the credibility of the evaluation scores. It is more conducive for researchers, companies, or the general public to select and use appropriate large models in actual application scenarios, and further reduces the risk of violating laws, regulations, and social ethics during the application process of the large model.

[0164] In this embodiment,

[0165] In the formula, represents the evaluation stability, n represents the number of dynamic evaluation datasets, scorei Let \(S_i\) denote the gender fairness score of the large model under the \(i\)-th dynamic evaluation dataset, \(\mu\) denote the average of the gender fairness scores of the large model under all dynamic evaluation datasets, \(n\) denote the number of dynamic evaluation datasets, and \(i = 1, 2, \ldots, n\).

[0166] It should be noted that the smaller the value, the higher the stability of the gender fairness evaluation of the large model in this time, and the more credible the gender fairness score.

[0167] Fourthly, the present invention also proposes a large model evaluation system, including:

[0168] An acquisition module, configured to acquire a plurality of dynamic evaluation text datasets;

[0169] An image generation from text module, configured to perform image generation from text on the plurality of dynamic evaluation text datasets respectively to obtain a plurality of dynamic evaluation image sets;

[0170] A data quality control module, configured to perform data quality control on the plurality of dynamic evaluation image sets respectively to obtain a plurality of intermediate dynamic evaluation image sets;

[0171] A question option generation module, configured to generate question options for the images in the plurality of intermediate dynamic evaluation image sets respectively to obtain a plurality of dynamic evaluation datasets;

[0172] An evaluation module, configured to perform parallel evaluation on the large model to be evaluated by using the plurality of dynamic evaluation datasets to obtain a plurality of answer result sets;

[0173] A calculation module, configured to calculate the gender fairness score of the large model according to the plurality of answer result sets.

[0174] In this embodiment, it further includes an occupation text library, which includes a plurality of occupations, and each occupation includes two descriptive text lists, and the two descriptive text lists correspond to men and women respectively;

[0175] Each descriptive text list includes a plurality of descriptive texts, and each descriptive text includes the occupation and gender information of the person.

[0176] Among them, the process of obtaining the plurality of dynamic evaluation data specifically includes: randomly selecting a certain number of occupations from the occupation text library; using all the descriptive texts of the selected occupations as a dynamic evaluation text dataset; repeating the above steps until a plurality of dynamic evaluation text datasets are obtained.

[0177] Among them, in the data quality control process of the data quality control module, all the images in the plurality of dynamic evaluation image sets are respectively input into the referee model for data quality review;

[0178] During the data quality review process, when a certain picture does not meet the review requirements of the referee model, the picture is deleted from the dynamic evaluation picture set to which it belongs;

[0179] When the referee model traverses all the pictures in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as an intermediate dynamic evaluation picture set;

[0180] When multiple dynamic evaluation picture sets have all completed the data quality review, multiple intermediate dynamic evaluation picture sets are obtained.

[0181] In this embodiment, each picture in each dynamic evaluation data set corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations with high, medium, and low similarities to the correct occupation respectively.

[0182] In this embodiment, each answer result set includes the option of the answer to the occupation of the person in each picture in the corresponding dynamic evaluation data set and the answer result score; among them, the answer result scores of the four options are S1, S2, S3, and S4 in the order of the correct occupation, high similarity, medium similarity, and low similarity, and S1 > S2 > S3 > S4.

[0183] In one specific embodiment, S1, S2, S3, and S4 are 5 points, 3 points, 2 points, and 1 point respectively.

[0184] In this embodiment,

[0185] In the formula, Score represents the gender fairness score of the large model; score i represents the gender fairness score of the large model under the i-th dynamic evaluation data set, i = 1, 2,..., n; n represents the number of dynamic evaluation data sets;

[0186] Among them,

[0187] In the formula, M i represents the number of randomly selected occupations in the i-th dynamic evaluation data set; question i,j,男 represents the answer result score of the large model for the male question corresponding to the j-th occupation in the i-th dynamic evaluation data set; question i,j,女 is the answer result score of the large model for the female question corresponding to the j-th occupation in the i-th dynamic evaluation data set.

[0188] In this embodiment, it further includes: a stability calculation module, which is used to calculate the evaluation stability of the large model according to the gender fairness score.

[0189] Among them,

[0190] In the formula, represents the evaluation stability, n represents the number of dynamic evaluation data sets, and score i represents the gender fairness score of the large model under the i-th dynamic evaluation data set, μ represents the average value of the gender fairness scores of the large model under all dynamic evaluation data sets, and i = 1, 2,..., n.

[0191] In a fifth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the dynamic evaluation data set construction method described in any one of the first aspects.

[0192] In a sixth aspect, the present invention also provides an electronic device, including: a processor and a memory, where the memory is used to store one or more programs; when the one or more programs are executed by the processor, it implements the dynamic evaluation data set construction method described in any one of the first aspects.

[0193] In a seventh aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the large model evaluation method described in any one of the third aspects.

[0194] In an eighth aspect, the present invention also provides an electronic device, including: a processor and a memory, where the memory is used to store one or more programs; when the one or more programs are executed by the processor, it implements the large model evaluation method described in any one of the third aspects.

[0195] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for constructing a dynamic evaluation data set for large model evaluation, characterized in that: include: Get dynamic evaluation text dataset; Perform text generation mapping on the dynamic evaluation text dataset to obtain a dynamic evaluation picture set; Perform data quality control on the dynamic evaluation picture set to obtain the intermediate dynamic evaluation picture set; Generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation data set.

2. The method for constructing a dynamic evaluation data set for large model evaluation according to claim 1, characterized in that: Before obtaining the dynamic evaluation text dataset, it also includes: Constructing an occupational text library; wherein the occupational text library includes multiple occupations, each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively; each descriptive text list includes multiple descriptive texts, and each descriptive text includes occupation and gender information of the character; The step of obtaining a dynamic evaluation text dataset specifically includes: randomly selecting a certain number of occupations from an occupation text library; and taking all descriptive texts of the selected occupations as a dynamic evaluation text dataset.

3. The method for constructing a dynamic evaluation data set for large model evaluation according to claim 1, characterized in that: Perform data quality control on the dynamic evaluation picture set to obtain the intermediate dynamic evaluation picture set, including: Input all images in the dynamic evaluation image set into the referee model for data quality review; During the data quality review process, if a certain image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs; When the referee model traverses all the images of a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.

4. The method for constructing a dynamic evaluation data set for large model evaluation according to claim 1, characterized in that: There are four question options corresponding to each picture in each dynamic evaluation data set, one of which is the correct occupation, and the other three are similar occupations, and their similarity to the correct occupation is high, medium, and low, respectively.

5. A dynamic evaluation data set construction system for large model evaluation, characterized in that: include: An acquisition module is used to obtain a dynamic evaluation text dataset; The text-generated graph module is used to perform text-generated graphs on the dynamic evaluation text dataset to obtain a dynamic evaluation picture set; The data quality control module is used to perform data quality control on the dynamic evaluation picture set to obtain the intermediate dynamic evaluation picture set; The question option generation module is used to generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation data set.

6. A large model evaluation method, characterized in that: include: Using the dynamic evaluation data set construction method described in any one of claims 1 to 4 to construct multiple dynamic evaluation data sets; Use multiple dynamic evaluation data sets to evaluate the large model in parallel and obtain multiple answer result sets; Based on multiple sets of answer results, the gender fairness score of the large model is calculated.

7. The large model evaluation method according to claim 6, characterized in that: Use multiple dynamic evaluation data sets to evaluate the large model in parallel and obtain multiple answer result sets, including: Deploy the large model to be evaluated into multiple large model copies; Input multiple dynamic evaluation data sets into multiple large model copies one by one for evaluation, and obtain multiple answer result sets; Among them, the evaluation of multiple large model copies is performed in parallel.

8. The large model evaluation method according to claim 6, characterized in that: Each answer result set includes the answer options and answer result scores for the occupation of the person in each picture in the corresponding dynamic evaluation data set; Among them, the scores of the answer results of the four options are S1, S2, S3 and S4 in the order of correct occupation, high similarity, medium similarity and low similarity, and S1>S2>S3>S4; in, In the formula, Score represents the gender fairness score of the large model; score i represents the gender fairness score of the large model in the i-th dynamic evaluation dataset, i = 1, 2, …, n; n represents the number of dynamic evaluation datasets; in, In the formula, m i represents the number of randomly selected occupations in the i-th dynamic evaluation data set; question i,j,男 It represents the answer score of the large model for the male question corresponding to the jth occupation in the i-th dynamic evaluation data set; question i,j,女 It represents the score of the big model's answer to the female question corresponding to the jth occupation in the i-th dynamic evaluation dataset.

9. The large model evaluation method according to claim 8, characterized in that: After calculating the gender fairness score of the large model based on multiple answer result sets, it also includes: Based on the gender fairness score, the evaluation stability of the large model is calculated; in, In the formula, represents the evaluation stability, n represents the number of dynamic evaluation data sets, score i represents the gender fairness score of the large model in the i-th dynamic evaluation dataset, μ represents the average gender fairness score of the large model in all dynamic evaluation datasets, i = 1, 2, …, n.

10. A large model evaluation system, characterized in that: include: A data set construction module, used to construct multiple dynamic evaluation data sets using the dynamic evaluation data set construction method described in any one of claims 1 to 4; An evaluation module is used to use multiple dynamic evaluation data sets to perform parallel evaluation on the large model to be evaluated, and obtain multiple answer result sets; The calculation module is used to calculate the gender fairness score of the large model based on multiple answer result sets.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a dynamic evaluation data set as described in any one of claims 1 to 4 is implemented; Alternatively, when the computer program is executed by a processor, the large model evaluation method as described in any one of claims 6 to 9 is implemented.

12. An electronic device comprising: A processor and a memory, wherein the memory is used to store one or more programs; wherein when the one or more programs are executed by the processor, the method for constructing a dynamic evaluation data set as described in any one of claims 1 to 4 is implemented; Alternatively, implement the large model evaluation method as described in any one of claims 6-9.

Citation Information

Patent Citations

  • Information processing method and apparatus realized by computer

    CN107392217A

  • Evaluation method and system for large model content security capability

    CN118035711A

  • Evaluation method and device of text and graph generation model, electronic equipment and storage medium

    CN118365751A

  • Multi-modal online evaluation data processing method and system

    CN118445578A

  • Occupational interest evaluation method and device, computer equipment and storage medium

    CN119515324A