Large model evaluation methods, systems, media and equipment
By constructing a parallel evaluation method for dynamic evaluation datasets, the problems of unclear gender fairness indicators and unstable evaluation results in large-scale model evaluation are solved, objective evaluation and stability scoring of gender fairness of large models are achieved, and application risks are reduced.
Patent Information
- Application Number
- CN202510340119.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Existing large-scale model evaluation methods lack a clear definition of gender fairness indicators. The evaluation data sets are small and fixed, resulting in incomplete and unfair evaluation results. The evaluation process is random and difficult to reflect the actual performance of the model in terms of gender fairness. There are legal and regulatory and social ethical risks.
Construct a dynamic evaluation dataset, obtain a dynamic evaluation text dataset, generate a picture set from text, perform data quality control, generate question options, use multiple dynamic evaluation datasets for parallel evaluation, and calculate the gender fairness score.
It achieves objective evaluation of gender fairness of large models, ensures the authenticity, reliability and stability of gender fairness scores, reduces evaluation time, and reduces legal, regulatory and social ethical risks.
Smart Images

Figure CN120163253B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model evaluation, and in particular to a large model evaluation method, system, medium and equipment. Background Art
[0002] In today's society, gender equality is a frequent topic of discussion in various news headlines. When large models provide answers to user questions, they should also maintain basic gender fairness, avoid gender discrimination or bias, and adhere to existing laws, regulations, and social ethics. Therefore, evaluating large models based on gender fairness is a fundamental requirement for ensuring their widespread practical application. However, current methods for evaluating the gender fairness of large models still have the following issues:
[0003] (1) Lack of clear definition of gender equity indicators;
[0004] (2) Evaluation datasets are generally manually screened and labeled, and the process of generating datasets consumes a lot of manpower and material resources. Therefore, the scale of evaluation datasets is generally small. Once the evaluation dataset is manually labeled, the evaluation dataset will remain fixed during the evaluation of the large model. This allows the evaluation model that is trained and adjusted based on a known fixed dataset to obtain a higher false score when being evaluated. The limited dataset has certain limitations in the breadth of the evaluation, and the evaluation coverage of gender equality is not comprehensive enough, so the evaluation results are inevitably unfair.
[0005] (3) The evaluation process is random: Since the underlying logic of the large model is deep learning, its output is related to probability, and the output may be slightly different when the same data is input.
[0006] Therefore, a single evaluation on a fixed and limited data set is not sufficient to reflect the actual performance of the model in terms of gender fairness. The large models selected by researchers, companies or the general public based on this may run the risk of violating laws, regulations and social ethics during actual application. Summary of the Invention
[0007] In order to solve the technical problems existing in the background technology, the present invention proposes a large model evaluation method, system, medium and equipment.
[0008] In a first aspect, the present invention proposes a method for constructing a dynamic evaluation dataset for large model evaluation, comprising:
[0009] Obtain dynamic evaluation text dataset;
[0010] Perform text-based mapping on the dynamic evaluation text dataset to obtain a dynamic evaluation picture set;
[0011] Perform data quality control on the dynamic evaluation image set to obtain the intermediate dynamic evaluation image set;
[0012] Generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation dataset.
[0013] Preferably, before obtaining the dynamic evaluation text dataset, the method further includes:
[0014] Constructing an occupational text library; wherein the occupational text library includes multiple occupations, each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the character;
[0015] The acquisition of the dynamic evaluation text dataset specifically includes: randomly selecting a certain number of occupations from the occupation text library; and taking all descriptive texts of the selected occupations as a dynamic evaluation text dataset.
[0016] Preferably, data quality control is performed on the dynamic evaluation picture set to obtain an intermediate dynamic evaluation picture set, specifically including:
[0017] All images in the dynamic evaluation image set are input into the referee model for data quality review;
[0018] During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0019] When the referee model traverses all the images in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.
[0020] Preferably, each picture in each dynamic evaluation data set corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations and their similarities to the correct occupation are high, medium and low, respectively.
[0021] In a second aspect, the present invention further proposes a dynamic evaluation data set construction system for large model evaluation, comprising:
[0022] Acquisition module, used to obtain dynamic evaluation text dataset;
[0023] The text-based graph module is used to perform text-based graphing on the dynamic evaluation text dataset to obtain a dynamic evaluation image set;
[0024] The data quality control module is used to perform data quality control on the dynamic evaluation image set and obtain the intermediate dynamic evaluation image set;
[0025] The question option generation module is used to generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation data set.
[0026] In a third aspect, the present invention provides a large model evaluation method, comprising: constructing multiple dynamic evaluation data sets using the dynamic evaluation data set construction method for large model evaluation described in any one of the first aspects;
[0027] Use multiple dynamic evaluation data sets to conduct parallel evaluation of the large model to be evaluated, and obtain multiple sets of answer results;
[0028] Based on multiple sets of answer results, the gender fairness score of the large model is calculated.
[0029] Preferably, multiple dynamic evaluation data sets are used to perform parallel evaluation on the large model to be evaluated, and multiple answer result sets are obtained, specifically including:
[0030] Deploy the large model to be evaluated into multiple large model copies;
[0031] Multiple dynamic evaluation data sets are input into multiple large model copies one by one for evaluation to obtain multiple answer result sets; wherein, the evaluation of multiple large model copies is executed in parallel.
[0032] Preferably, each answer result set includes the answer options and answer result scores of the occupation of the person in each picture in the corresponding dynamic evaluation data set; wherein the answer result scores of the four options are in the order of correct occupation, high similarity, medium similarity, and low similarity. S1, S2, S3 and S4 ,and .
[0033] Preferably, ;
[0034] Where, represents the gender fairness score of the large model; Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; Indicates the number of dynamic evaluation data sets.
[0035] Preferably, ;
[0036] Where, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; Indicates the number of dynamic evaluation data sets; Indicates the The number of randomly selected occupations in the evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; For large models The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
[0037] Preferably, after calculating the gender fairness score of the large model based on multiple answer result sets, the method further includes:
[0038] Based on the gender fairness score, the evaluation stability of the large model is calculated.
[0039] Preferably, ;
[0040] Where, Indicates the evaluation stability. Indicates the number of dynamic evaluation data sets, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, represents the average gender fairness score of the large model under all dynamic evaluation datasets, .
[0041] In a fourth aspect, the present invention further proposes a large model evaluation system, comprising:
[0042] A data set construction module, configured to construct a plurality of dynamic evaluation data sets using the dynamic evaluation data set construction method for large model evaluation described in any one of the first aspects;
[0043] The evaluation module is used to use multiple dynamic evaluation data sets to perform parallel evaluation on the large model to be evaluated, and obtain multiple answer result sets;
[0044] The calculation module is used to calculate the gender fairness score of the large model based on multiple answer result sets.
[0045] In a fifth aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for constructing a dynamic evaluation data set for large model evaluation as described in any one of the first aspects.
[0046] In the sixth aspect, the present invention also proposes an electronic device, comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the method for constructing a dynamic evaluation data set for large model evaluation as described in any one of the first aspects is implemented.
[0047] In the seventh aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large model evaluation method as described in any one of the third aspects.
[0048] In an eighth aspect, the present invention further proposes an electronic device comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the large model evaluation method as described in any one of the third aspects is implemented.
[0049] The proposed large-scale model evaluation method, system, medium, and device obtain a dynamic evaluation text dataset, dynamically generate a dynamic evaluation image set based on the dynamic evaluation text dataset, and perform data quality control on the dynamic evaluation image set to obtain an intermediate dynamic evaluation image set; question options are generated for the images in the intermediate dynamic evaluation image set to obtain a dynamic evaluation dataset. The present invention ensures the high quality and randomness of the constructed dynamic evaluation dataset, effectively avoiding the instability of the gender fairness score caused by the targeted training of the large model to be evaluated based on the static evaluation dataset or the randomness of the data and model, thereby ensuring the authenticity and reliability of the gender fairness score. Furthermore, the present invention uses multiple dynamic evaluation datasets to conduct multiple rounds of parallel evaluation on the large model to be evaluated, obtaining multiple answer result sets. Then, based on the multiple answer result sets, the gender fairness score of the large model is calculated, achieving an objective evaluation of the large model on gender fairness, while reducing the randomness of the gender fairness score and significantly shortening the evaluation time.
[0050] The present invention achieves an objective evaluation of the gender fairness of large models, which facilitates scientific researchers, companies or the general public to choose and use appropriate large models in actual application scenarios in the future, effectively reduces the risk of violating laws, regulations and social ethics during the application of large models, and effectively improves the experience of scientific researchers, companies or the general public in the process of choosing to use large models. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of a method for constructing a dynamic evaluation data set for large model evaluation in one embodiment of the present invention.
[0052] Figure 2 A schematic diagram of parallel evaluation in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0054] First, as Figure 1 As shown, the present invention proposes a method for constructing a dynamic evaluation data set for large model evaluation, including:
[0055] Obtain dynamic evaluation text dataset;
[0056] Perform text-based mapping on the dynamic evaluation text dataset to obtain a dynamic evaluation picture set;
[0057] Perform data quality control on the dynamic evaluation image set to obtain the intermediate dynamic evaluation image set;
[0058] Generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation dataset.
[0059] The present invention obtains a dynamic evaluation text dataset, dynamically generates a dynamic evaluation picture set based on the dynamic evaluation text dataset, and performs data quality control on the dynamic evaluation picture set to obtain an intermediate dynamic evaluation picture set; generates question options for pictures in the intermediate dynamic evaluation picture set to obtain a dynamic evaluation dataset, thereby ensuring the high quality and randomness of the constructed multiple dynamic evaluation datasets, effectively avoiding the subsequent large model to be evaluated from being targetedly trained based on the static evaluation dataset or the instability of the gender fairness score due to the randomness of the data and the model, thereby ensuring the authenticity and reliability of the gender fairness score.
[0060] In order to obtain the dynamic evaluation text dataset, in this embodiment, before obtaining the dynamic evaluation text dataset, the following steps are further included: constructing an occupation text library; wherein the occupation text library includes multiple occupations, each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the person;
[0061] The acquisition of the dynamic evaluation text dataset specifically includes: randomly selecting a certain number of occupations from the occupation text library; and taking all descriptive texts of the selected occupations as a dynamic evaluation text dataset.
[0062] The data format of the occupational text library is as follows.
[0063]
[0064] In this embodiment, a dynamic evaluation text dataset is subjected to a text graph to obtain a dynamic evaluation picture set, specifically comprising: inputting all descriptive texts of all occupations in the dynamic evaluation text dataset into a text graph model to obtain a dynamic evaluation picture set.
[0065] This embodiment is configured in such a way that a dynamic evaluation picture set is dynamically generated based on a dynamic evaluation text data set through a text-based graph model, thereby ensuring the randomness of the subsequently obtained dynamic evaluation data set. This can effectively avoid the large model to be evaluated from being subjected to targeted training based on a static evaluation data set or the instability of the gender fairness score due to the randomness of the data and the model, thereby effectively ensuring the authenticity and reliability of the gender fairness score.
[0066] In one specific embodiment, the cultural graph model is a stable-diffusion3 model.
[0067] In another specific embodiment, the cultural graph model is a Janus-Pro multimodal large model.
[0068] In another specific embodiment, the Emu3-Gen model is an Emu3-Gen model.
[0069] In this embodiment, data quality control is performed on the dynamic evaluation picture set to obtain an intermediate dynamic evaluation picture set, specifically including:
[0070] All images in the dynamic evaluation image set are input into the referee model for data quality review;
[0071] During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0072] When the referee model traverses all the images in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.
[0073] This embodiment uses the referee model to control the quality of the generated data, which can ensure the randomness of the subsequent dynamic evaluation data set while maintaining high quality. It further avoids the targeted training of large models based on static evaluation data sets or the instability of gender fairness scores due to the randomness of data and models, thereby further ensuring the authenticity and reliability of gender fairness scores.
[0074] In one specific embodiment, the referee model is a Grounded-Segment-Anything model.
[0075] In the specific implementation, all images in the dynamic evaluation image set are input into the Grounded-Segment-Anything model for data quality review;
[0076] During the data quality review process, if the Grounded-Segment-Anything model detects that there are multiple people in a certain picture, the picture will be deleted from the dynamic evaluation picture set to which it belongs; when the Grounded-Segment-Anything model traverses all pictures in a certain dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review; the dynamic evaluation picture set that has completed the data quality review will be used as the intermediate dynamic evaluation picture set.
[0077] In another specific embodiment, the referee model is a Qianwen model.
[0078] During the specific implementation, all images in the dynamic evaluation image set will be input into the Qianwen model for data quality review; if the Qianwen model identifies that the gender and occupation of the person in a certain image are inconsistent with the gender and occupation information in the corresponding descriptive text, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0079] When the Qianwen model traverses all the pictures in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review; the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.
[0080] In another embodiment, the judging model is a Grounded-Segment-Anything model and a Qianwenda model. The Qianwenda model is responsible for determining whether the gender and occupation of the person in the image are consistent with the gender and occupation information in the corresponding descriptive text, and the Grounded-Segment-Anything model is responsible for determining whether there is more than one person in the image.
[0081] During the specific implementation, all images in the dynamic evaluation image set are first input into the Grounded-Segment-Anything model for data quality review, and then the images in the dynamic evaluation image set that have passed the review of the Grounded-Segment-Anything model are input into the Qianwen model for data quality review, which effectively improves the high quality of the intermediate dynamic evaluation image set and subsequent dynamic evaluation data sets.
[0082] In this embodiment, each picture in each dynamic evaluation data set corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations, and their similarities to the correct occupation are high, medium, and low, respectively.
[0083] In a second aspect, the present invention further proposes a dynamic evaluation data set construction system for large model evaluation, comprising:
[0084] Acquisition module, used to obtain dynamic evaluation text dataset;
[0085] The text-based graph module is used to perform text-based graphing on the dynamic evaluation text dataset to obtain a dynamic evaluation image set;
[0086] The data quality control module is used to perform data quality control on the dynamic evaluation image set and obtain the intermediate dynamic evaluation image set;
[0087] The question option generation module is used to generate question options for the pictures in the intermediate dynamic evaluation picture set to obtain the dynamic evaluation data set.
[0088] In this embodiment, during the acquisition of the dynamic evaluation text data set by the acquisition module, a certain number of occupations are randomly selected from the pre-built occupation text library;
[0089] All descriptive texts of the selected occupations are used as a dynamic evaluation text dataset.
[0090] In this embodiment, in the text graph process of the text graph module, all descriptive texts of all occupations in the dynamic evaluation text dataset are input into the text graph model to obtain a dynamic evaluation picture set.
[0091] In this embodiment, during the data quality control process of the data quality control module, all images in the dynamic evaluation image set are input into the referee model for data quality review;
[0092] During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0093] When the referee model traverses all the images in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.
[0094] In this embodiment, each picture in each dynamic evaluation data set corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations, and their similarities to the correct occupation are high, medium, and low, respectively.
[0095] In a third aspect, the present invention proposes a large model evaluation method, comprising:
[0096] Obtain multiple dynamic evaluation text datasets;
[0097] Perform text-generated graphs on multiple dynamic evaluation text datasets to obtain multiple dynamic evaluation image sets;
[0098] Perform data quality control on multiple dynamic evaluation picture sets respectively to obtain multiple intermediate dynamic evaluation picture sets;
[0099] Generate question options for pictures in multiple intermediate dynamic evaluation picture sets respectively, and obtain multiple dynamic evaluation data sets;
[0100] Use multiple dynamic evaluation data sets to conduct parallel evaluation of the large model to be evaluated, and obtain multiple sets of answer results;
[0101] Based on multiple sets of answer results, the gender fairness score of the large model is calculated.
[0102] The present invention obtains multiple dynamic evaluation text data sets, dynamically generates multiple dynamic evaluation picture sets based on the multiple dynamic evaluation text data sets, and performs data quality control on the multiple dynamic evaluation picture sets respectively, thereby ensuring the high quality and randomness of the multiple dynamic evaluation data sets obtained subsequently, effectively avoiding the large model to be evaluated from being targetedly trained based on the static evaluation data set or the instability of the gender fairness score due to the randomness of the data and the model, thereby ensuring the authenticity and reliability of the gender fairness score. Moreover, the present invention uses multiple dynamic evaluation data sets to conduct multiple rounds of parallel evaluation on the large model to be evaluated, obtains multiple answer result sets, and then calculates the gender fairness score of the large model based on the multiple answer result sets, thereby achieving an objective evaluation of the large model in terms of gender fairness, while reducing the randomness of the gender fairness score and significantly shortening the evaluation time.
[0103] The large-scale model evaluation method proposed in the present invention realizes the objective evaluation of the gender fairness of the large-scale model, which facilitates scientific researchers, companies or the general public to choose and use appropriate large-scale models in actual application scenarios, effectively reduces the risk of violating laws, regulations and social ethics during the application of large models, and effectively improves the experience of scientific researchers, companies or the general public in the process of choosing to use large-scale models.
[0104] In this embodiment, before obtaining multiple dynamic evaluation text data sets, it also includes: constructing an occupational text library; wherein the occupational text library includes multiple occupations, each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the character's occupation and gender information.
[0105] The data format of the occupational text library is as follows.
[0106]
[0107] Therefore, obtaining multiple dynamic evaluation text data sets in this embodiment specifically includes: randomly selecting a certain number of occupations from the occupation text library; using all descriptive texts of the selected occupations as dynamic evaluation text data sets; repeating the above steps until multiple dynamic evaluation text data sets are obtained.
[0108] In this embodiment, multiple dynamic evaluation text data sets are subjected to text graph models to obtain multiple dynamic evaluation picture sets, specifically including: inputting all descriptive texts of all occupations in the multiple dynamic evaluation text data sets into the text graph model to obtain multiple dynamic evaluation picture sets.
[0109] This embodiment is configured in such a way that a plurality of dynamic evaluation picture sets are dynamically generated based on a plurality of dynamic evaluation text data sets through a text-based graph model, thereby ensuring the randomness of the subsequently obtained dynamic evaluation data sets. This can effectively avoid the large model to be evaluated from being subjected to targeted training based on a static evaluation data set or the instability of the gender fairness score due to the randomness of the data and the model, thereby effectively ensuring the authenticity and reliability of the gender fairness score.
[0110] In one specific embodiment, the cultural graph model is a stable-diffusion3 model.
[0111] In another specific embodiment, the cultural graph model is a Janus-Pro multimodal large model.
[0112] In another specific embodiment, the Emu3-Gen model is an Emu3-Gen model.
[0113] In this embodiment, data quality control is performed on multiple dynamic evaluation picture sets respectively to obtain multiple intermediate dynamic evaluation picture sets, specifically including:
[0114] All images in multiple dynamic evaluation image sets are input into the referee model for data quality review;
[0115] During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0116] When the referee model traverses all the images in a dynamic evaluation image set, the data quality review of the dynamic evaluation image set has been completed;
[0117] The completed data quality review dynamic evaluation picture set is used as the intermediate dynamic evaluation picture set;
[0118] When the data quality review of multiple dynamic evaluation picture sets has been completed, multiple intermediate dynamic evaluation picture sets are obtained.
[0119] This embodiment uses the referee model to control the quality of the generated data, which can ensure the randomness of the subsequent dynamic evaluation data set while maintaining high quality. It further avoids the targeted training of large models based on static evaluation data sets or the instability of gender fairness scores due to the randomness of data and models, thereby further ensuring the authenticity and reliability of gender fairness scores.
[0120] In one specific embodiment, the referee model is a Grounded-Segment-Anything model.
[0121] In specific implementation, all images in multiple dynamic evaluation image sets are input into the Grounded-Segment-Anything model for data quality review;
[0122] During the data quality review process, if the Grounded-Segment-Anything model detects that there are multiple people in an image, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0123] When the Grounded-Segment-Anything model traverses all images in a dynamic evaluation image set, the data quality review of the dynamic evaluation image set has been completed;
[0124] The dynamic evaluation picture set that has completed data quality review is used as the intermediate dynamic evaluation picture set.
[0125] In another specific embodiment, the referee model is a Qianwen model.
[0126] During the specific implementation, all images in multiple dynamic evaluation image sets are input into the Qianwen model for data quality review;
[0127] If the Qianwen model identifies that the gender and occupation of the person in a certain image are inconsistent with the gender and occupation information in the corresponding descriptive text, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0128] When the Qianwen model traverses all the images in a dynamic evaluation image set, the data quality review of the dynamic evaluation image set has been completed;
[0129] The dynamic evaluation picture set that has completed data quality review is used as the intermediate dynamic evaluation picture set.
[0130] In another embodiment, the referee model is the Grounded-Segment-Anything model and the Thousand Questions model.
[0131] Among them, the Qianwen model is responsible for determining whether the gender and occupation of the person in the image are consistent with the gender and occupation information in the corresponding descriptive text, and Grounded-Segment-Anything is responsible for determining whether there is more than one person in the image.
[0132] During the specific implementation, all images in multiple dynamic evaluation picture sets are first input into the Grounded-Segment-Anything model for data quality review, and then the images in multiple dynamic evaluation picture sets that have passed the review of the Grounded-Segment-Anything model are input into the Qianwen model for data quality review, which effectively improves the high quality of the intermediate dynamic evaluation picture sets and subsequent dynamic evaluation data sets.
[0133] In this embodiment, a large model to be evaluated is evaluated in parallel using multiple dynamic evaluation data sets to obtain multiple answer result sets, specifically including: deploying the large model to be evaluated into multiple large model copies; inputting multiple dynamic evaluation data sets into multiple large model copies one by one for evaluation to obtain multiple answer result sets; wherein, the evaluation of multiple large model copies is executed in parallel.
[0134] like Figure 2 As shown, in this embodiment, the large model to be evaluated is deployed as multiple large model replicas, each of which exclusively uses the GPU resources of a graphics card. The number of large model replicas equals the number of dynamic evaluation datasets. During evaluation, multiple parallel processes are created, each of which is assigned a dynamic evaluation dataset and a large model replica. Within each process, the assigned dynamic evaluation dataset is used as input, and the interface of the large model replica is called to obtain the options for the large model to answer the occupation of each person in the image.
[0135] This configuration of the present embodiment can evaluate multiple dynamic evaluation data sets, thereby reducing the randomness of gender fairness scores while significantly shortening the evaluation time, which is conducive to improving the stability and accuracy of the evaluation.
[0136] In this embodiment, question options are generated for multiple intermediate dynamic evaluation image sets to obtain multiple dynamic evaluation data sets, specifically including:
[0137] Generate question options for all images in multiple intermediate dynamic evaluation image sets, so that each image has four corresponding question options; among the four options, one option is the correct occupation, and the other three options are similar occupations;
[0138] The dynamic evaluation image sets generated by multiple completed question selections are used as multiple dynamic evaluation data sets.
[0139] In the process of generating question options, since each picture corresponds to a descriptive text and the correct occupation corresponding to the picture is known, only three similar occupations need to be generated.
[0140] In this example, each occupation is converted into a multidimensional vector representation using a clip encoding model. The similarity between occupations can be expressed through the similarity of these multidimensional vectors. When generating question options, three occupations with high, medium, and low similarity are randomly selected as the three similar occupation options for the image, based on the similarity values between the occupations.
[0141] That is to say, each picture in each dynamic evaluation dataset corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations and their similarity to the correct occupation is high, medium, and low, respectively.
[0142] Therefore, in order to facilitate the calculation of the gender fairness score of the large model, each answer result set in this embodiment includes the answer options and answer result scores of the occupation of the person in each picture in the corresponding dynamic evaluation data set; among them, the answer result scores of the four options are in the order of correct occupation, high similarity, medium similarity, and low similarity. S1, S2, S3 and S4 ,and S1 > S2 > S3 > S4 .
[0143] In this way, this embodiment can characterize the degree of gender fairness of the large model to be evaluated in the corresponding occupation through the design of different score options.
[0144] In this embodiment, ;
[0145] Where, represents the gender fairness score of the large model; Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; Indicates the number of dynamic evaluation data sets.
[0146] in, ;
[0147] Where, Indicates the The number of randomly selected occupations in the dynamic evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; Indicates that the large model is for the The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
[0148] This embodiment provides clear indicators for evaluating gender fairness, which is conducive to calculating the gender fairness score of large models.
[0149]
[0150] In one specific embodiment, when the large model identifies the occupation of the person in the image, it assigns different scores to the four options, with scores of 5, 3, 2, and 1, respectively, for the correct occupation, high similarity, medium similarity, and low similarity. For example, the data format of one dynamic evaluation dataset Q is shown above. Each occupation has one image of a man and one image of a woman, and each image has four occupation options with different scores.
[0151] Therefore, in this embodiment, ;
[0152] Where, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, Indicates the The number of randomly selected occupations in the evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; For large models The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
[0153] It should be known that One of 5 points, 3 points, 2 points and 1 point; It can also be one of 5 points, 3 points, 2 points and 1 point.
[0154] In this embodiment, after calculating the gender fairness score of the large model based on multiple answer result sets, the method further includes: calculating the evaluation stability of the large model based on the gender fairness score.
[0155] This embodiment calculates the evaluation stability of the large model based on the gender fairness score, which can objectively describe the stability of the evaluation process and the credibility of the evaluation score. It is more conducive to scientific researchers, companies or the general public to choose and use appropriate large models in actual application scenarios, and further reduces the risk of violating laws, regulations and social ethics during the application of large models.
[0156] In this embodiment, ;
[0157] Where, Indicates the evaluation stability. Indicates the number of dynamic evaluation data sets, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, represents the average gender fairness score of the large model under all dynamic evaluation datasets, Indicates the number of dynamic evaluation data sets, .
[0158] What you need to know is, The smaller the value, the higher the stability of the gender fairness evaluation of the large model and the more credible the gender fairness score.
[0159] In a fourth aspect, the present invention further proposes a large model evaluation system, comprising:
[0160] An acquisition module is used to obtain multiple dynamic evaluation text data sets;
[0161] The text-generated graph module is used to perform text-generated graphs on multiple dynamic evaluation text datasets to obtain multiple dynamic evaluation image sets;
[0162] A data quality control module is used to perform data quality control on multiple dynamic evaluation image sets respectively to obtain multiple intermediate dynamic evaluation image sets;
[0163] A question option generation module is used to generate question options for pictures in multiple intermediate dynamic evaluation picture sets, thereby obtaining multiple dynamic evaluation data sets;
[0164] The evaluation module is used to use multiple dynamic evaluation data sets to perform parallel evaluation on the large model to be evaluated, and obtain multiple answer result sets;
[0165] The calculation module is used to calculate the gender fairness score of the large model based on multiple answer result sets.
[0166] In this embodiment, an occupational text library is also included. The occupational text library includes multiple occupations. Each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively.
[0167] Each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the character.
[0168] The process of obtaining multiple dynamic evaluation data specifically includes: randomly selecting a certain number of occupations from the occupation text library; taking all descriptive texts of the selected occupations as a dynamic evaluation text dataset; repeating the above steps until multiple dynamic evaluation text datasets are obtained.
[0169] Among them, in the data quality control process of the data quality control module, all images in multiple dynamic evaluation image sets are input into the referee model for data quality review;
[0170] During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs;
[0171] When the referee model traverses all the images in a dynamic evaluation image set, the dynamic evaluation image set has completed the data quality review, and the dynamic evaluation image set that has completed the data quality review is used as the intermediate dynamic evaluation image set;
[0172] When the data quality review of multiple dynamic evaluation picture sets has been completed, multiple intermediate dynamic evaluation picture sets are obtained.
[0173] In this embodiment, each picture in each dynamic evaluation data set corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations, and their similarities to the correct occupation are high, medium, and low, respectively.
[0174] In this embodiment, each answer result set includes the answer options and answer result scores of the occupation of the person in each picture in the corresponding dynamic evaluation data set; among them, the answer result scores of the four options are in the order of correct occupation, high similarity, medium similarity, and low similarity. S1, S2, S3 and S4 ,and S1 > S2 > S3 > S4 .
[0175] In one specific embodiment, S1, S2, S3 and S4 They are 5 points, 3 points, 2 points and 1 point respectively.
[0176] In this embodiment, ;
[0177] Where, represents the gender fairness score of the large model; Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; Indicates the number of dynamic evaluation data sets;
[0178] in,
[0179] Where, Indicates the The number of randomly selected occupations in the dynamic evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; For large models The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
[0180] In this embodiment, it also includes: a stability calculation module, which is used to calculate the evaluation stability of the large model according to the gender fairness score.
[0181] in, ;
[0182] Where, Indicates the evaluation stability. Indicates the number of dynamic evaluation data sets, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, represents the average gender fairness score of the large model under all dynamic evaluation datasets, .
[0183] In a fifth aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for constructing a dynamic evaluation data set for large model evaluation as described in any one of the first aspects.
[0184] In the sixth aspect, the present invention also proposes an electronic device, comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the method for constructing a dynamic evaluation data set for large model evaluation as described in any one of the first aspects is implemented.
[0185] In the seventh aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large model evaluation method as described in any one of the third aspects.
[0186] In an eighth aspect, the present invention further proposes an electronic device comprising: a processor and a memory, the memory being used to store one or more programs; when the one or more programs are executed by the processor, the large model evaluation method as described in any one of the third aspects is implemented.
[0187] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A large model evaluation method, characterized in that: include: Obtain multiple dynamic evaluation text datasets; Perform text-generated graphs on multiple dynamic evaluation text datasets to obtain multiple dynamic evaluation image sets; Perform data quality control on multiple dynamic evaluation picture sets respectively to obtain multiple intermediate dynamic evaluation picture sets; Generate question options for pictures in multiple intermediate dynamic evaluation picture sets respectively, and obtain multiple dynamic evaluation data sets; Use multiple dynamic evaluation data sets to conduct parallel evaluation of the large model to be evaluated, and obtain multiple sets of answer results; Based on multiple sets of answer results, the gender fairness score of the large model is calculated; Each answer result set includes the answer options and answer result scores of the occupation of the person in each picture in the corresponding dynamic evaluation data set; the answer result scores of the four options are in the order of correct occupation, high similarity, medium similarity, and low similarity. S1, S2, S3 and S4 ,and S1>S2>S3>S4 ; in, ; Where, represents the gender fairness score of the large model; Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; represents the number of dynamic evaluation data sets; ; Where, Indicates the The number of randomly selected occupations in the dynamic evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; Indicates that the large model is for the The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
2. The large model evaluation method according to claim 1, characterized in that: Before obtaining multiple dynamic evaluation text datasets, it also includes: Constructing an occupational text library; wherein the occupational text library includes multiple occupations, each occupation includes two descriptive text lists, and the two descriptive text lists correspond to males and females respectively; each descriptive text list includes multiple descriptive texts, and each descriptive text includes the occupation and gender information of the character; The acquisition process of the dynamic evaluation text dataset specifically includes: randomly selecting a certain number of occupations from the occupation text library; and taking all the descriptive texts of the selected occupations as a dynamic evaluation text dataset.
3. The large model evaluation method according to claim 1, characterized in that: Perform data quality control on multiple dynamic evaluation image sets respectively to obtain multiple intermediate dynamic evaluation image sets, including: For each dynamic evaluation picture set, all pictures in the dynamic evaluation picture set are input into the referee model for data quality review; During the data quality review process, if an image does not meet the review requirements of the referee model, the image will be deleted from the dynamic evaluation image set to which it belongs; When the referee model traverses all the images in a dynamic evaluation picture set, the dynamic evaluation picture set has completed the data quality review, and the dynamic evaluation picture set that has completed the data quality review is used as the intermediate dynamic evaluation picture set.
4. The large model evaluation method according to claim 1, characterized in that: Each picture in each dynamic evaluation dataset corresponds to four question options, one of which is the correct occupation, and the other three options are similar occupations, and their similarity to the correct occupation is high, medium, and low, respectively.
5. The large model evaluation method according to claim 1, characterized in that: Use multiple dynamic evaluation data sets to evaluate the large model in parallel, and obtain multiple sets of answer results, including: Deploy the large model to be evaluated into multiple large model copies; Multiple dynamic evaluation data sets are input into multiple large model copies one by one for evaluation, and multiple answer result sets are obtained; Among them, the evaluation of multiple large model copies is performed in parallel.
6. The large model evaluation method according to claim 1, characterized in that: After calculating the gender fairness score of the large model based on multiple answer result sets, it also includes: Based on the gender fairness score, the evaluation stability of the large model is calculated; in, ; Where, Indicates the evaluation stability. Indicates the number of dynamic evaluation data sets, Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, represents the average gender fairness score of the large model under all dynamic evaluation datasets, .
7. A large model evaluation system, characterized in that: include: The data set construction module is used to obtain multiple dynamic evaluation text data sets; perform text generation and mapping on the multiple dynamic evaluation text data sets to obtain multiple dynamic evaluation picture sets; Performing data quality control on multiple dynamic evaluation picture sets respectively to obtain multiple intermediate dynamic evaluation picture sets; generating question options for pictures in multiple intermediate dynamic evaluation picture sets respectively to obtain multiple dynamic evaluation data sets; The evaluation module is used to use multiple dynamic evaluation data sets to perform parallel evaluation on the large model to be evaluated, and obtain multiple answer result sets; A calculation module is used to calculate the gender fairness score of the large model based on multiple sets of answer results; Each answer result set includes the answer options and answer result scores of the occupation of the person in each picture in the corresponding dynamic evaluation data set; the answer result scores of the four options are in the order of correct occupation, high similarity, medium similarity, and low similarity. S1, S2, S3 and S4 ,and S1>S2>S3>S4 ; in, ; Where, represents the gender fairness score of the large model; Indicates that the large model is Gender fairness scores under dynamic evaluation datasets, ; Indicates the number of dynamic evaluation data sets; in, ; Where, Indicates the The number of randomly selected occupations in the dynamic evaluation dataset; Indicates that the large model is for the The dynamic evaluation dataset The score of the male answer to the question corresponding to each occupation; Indicates that the large model is for the The dynamic evaluation dataset The scores of the answers to the questions about women's occupations.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the large model evaluation method according to any one of claims 1 to 6 is implemented.
9. An electronic device comprising: A processor and a memory, the memory being used to store one or more programs; characterized in that when one or more programs are executed by the processor, the large model evaluation method as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Evaluation method and system for large model content security capability
CN118035711A
Evaluation method and device of text and graph generation model, electronic equipment and storage medium
CN118365751A