Model fairness evaluation method and device
By optimizing the data probability distribution of image and text training samples through large-scale pre-trained models and generative adversarial networks, and combining self-supervised learning for credibility detection and sample processing, the problem of inaccurate fairness assessment in untrusted environments in existing technologies is solved, and the reliability and completeness of fairness assessment results in untrusted environments are achieved.
Patent Information
- Application Number
- CN202210379396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-04-12
AI Technical Summary
Existing technologies cannot accurately assess the fairness of artificial intelligence algorithms in untrusted environments, resulting in inaccurate fairness assessment results that fail to reflect the true level of fairness of the algorithms.
By constructing a generative adversarial network through a large-scale pre-trained model and a generative adversarial network model, the probability distribution of the image and text training samples is optimized. The credibility detection and sample processing are combined with self-supervised learning techniques to generate updated evaluation samples and conduct fairness assessment.
It improves the reliability and integrity of fairness assessment results in untrusted environments, and ensures the robustness of the fairness assessment system and the availability of assessment results in uncontrolled environments.
Smart Images

Figure CN114970670B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to two model fairness evaluation methods. Background Art
[0002] As the basic theories and technologies of artificial intelligence continue to make new breakthroughs, graphic algorithms based on artificial intelligence technology have been widely used in many public fields such as finance, education, medical care, and security, giving rise to a series of intelligent applications such as smart security, intelligent customer service, medical consultation, and personalized recommendations. These applications have not only greatly enriched and facilitated people's daily lives, but also promoted and facilitated the development and progress of social economy and science and technology.
[0003] However, the unfairness and even discrimination inherent in AI's automated decision-making processes have been widely debated. This has not only sparked concerns and questions about automated decision-making by algorithms, but has also gradually attracted widespread attention from society and the public. Some regions have also introduced laws and regulations on algorithmic fairness, explicitly stating that the development and application of AI algorithms must meet fairness constraints. Therefore, conducting algorithmic fairness assessments and eliminating algorithmic bias are essential steps throughout the algorithm's lifecycle, both for regulatory compliance and for improving user experience. Summary of the Invention
[0004] In light of this, the embodiments of this specification provide two model fairness assessment methods. One or more embodiments of this specification also involve two model fairness assessment devices, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a model fairness evaluation method is provided, including:
[0006] Determine the true data probability distribution of the image and / or text training samples based on the image and / or text training model;
[0007] Determine the credibility test result of the sample to be evaluated based on the probability distribution of the real data and the generative adversarial network model;
[0008] In the case where the credibility test result satisfies the non-credible condition, performing sample processing on the sample to be evaluated according to the credibility test result to obtain an updated evaluation sample;
[0009] A fairness evaluation is performed on the model to be evaluated based on the samples to be evaluated and the updated evaluation samples.
[0010] According to a second aspect of the embodiments of this specification, a model fairness evaluation device is provided, including:
[0011] The probability distribution determination module is configured to determine a real data probability distribution of the picture and / or text training sample according to a picture and / or text training model;
[0012] The detection result determination module is configured to determine a credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model;
[0013] The sample processing module is configured to perform sample processing on the to-be-evaluated sample according to the credibility detection result, to obtain an updated evaluation sample, in a case where the credibility detection result meets an untrusted condition.
[0014] The evaluation module is configured to perform fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample.
[0015] According to a third aspect of an embodiment of the present specification, a model fairness evaluation method is provided, applied to a model fairness evaluation platform, including:
[0016] Determine a real data probability distribution of a picture and / or text training sample according to a picture and / or text training model;
[0017] Receive a to-be-evaluated sample and a to-be-evaluated model sent by a user;
[0018] Determine a credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model;
[0019] In a case where the credibility detection result meets an untrusted condition, perform sample processing on the to-be-evaluated sample according to the credibility detection result, to obtain an updated evaluation sample;
[0020] Perform fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample.
[0021] Obtain a fairness evaluation result of the to-be-evaluated model, and return the fairness evaluation result to the user.
[0022] According to a fourth aspect of an embodiment of the present specification, a model fairness evaluation method is provided, applied to a model fairness evaluation platform, including:
[0023] The first determination module is configured to determine a real data probability distribution of a picture and / or text training sample according to a picture and / or text training model;
[0024] The data receiving module is configured to receive a to-be-evaluated sample and a to-be-evaluated model sent by a user;
[0025] a second determination module configured to determine a credibility detection result of the to-be-evaluated sample according to the real data probability distribution and a generative adversarial network model;
[0026] a sample updating module configured to, in a case where the credibility detection result satisfies an untrustworthy condition, perform sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample;
[0027] a fairness evaluation module configured to perform fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample;
[0028] a result display module configured to obtain a fairness evaluation result of the to-be-evaluated model and return the fairness evaluation result to the user.
[0029] According to a fifth aspect of an embodiment of the present specification, a computing device is provided, comprising:
[0030] a memory and a processor;
[0031] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the model fairness evaluation method.
[0032] According to a sixth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions, when executed by a processor, implement the steps of the model fairness evaluation method.
[0033] According to a seventh aspect of an embodiment of the present specification, a computer program is provided, and when the computer program is executed in a computer, the computer program causes the computer to execute the steps of the model fairness evaluation method.
[0034] One embodiment of the present specification implements two model fairness evaluation methods and devices. One of the model fairness evaluation methods comprises training a model according to pictures and / or texts, determining a real data probability distribution of picture and / or text training samples, determining a credibility detection result of a to-be-evaluated sample according to the real data probability distribution and a generative adversarial network model, in a case where the credibility detection result satisfies an untrustworthy condition, performing sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample, and performing fairness evaluation on a to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample.
[0035] Specifically, the model fairness evaluation method models the real data probability distribution of the image-text training sample through the image-text training model, performs credibility detection on the to-be-evaluated sample according to the real data probability distribution, and processes the non-credible sample to obtain an updated evaluation sample, so as to improve the reliability and completeness of the evaluation sample in a non-credible environment, thereby ensuring the robustness of the model fairness evaluation method in a credible environment and a non-credible environment and the availability of the model evaluation result, thereby ensuring the accuracy of the model fairness evaluation method in evaluating the to-be-evaluated model, and making the subsequent to-be-evaluated model have good effects in actual application from the aspects of meeting regulatory compliance and improving user experience, and being applicable to algorithm governance. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a specific scene schematic diagram of a model fairness evaluation method provided by an embodiment of the present specification;
[0037] Figure 2 is a flowchart of a model fairness evaluation method provided by an embodiment of the present specification;
[0038] Figure 3 is a process flowchart of a model fairness evaluation method provided by an embodiment of the present specification;
[0039] Figure 4 is a structural schematic diagram of a model fairness evaluation device provided by an embodiment of the present specification;
[0040] Figure 5 is a flowchart of another model fairness evaluation method provided by an embodiment of the present specification;
[0041] Figure 6 is a structural schematic diagram of another model fairness evaluation method provided by an embodiment of the present specification;
[0042] Figure 7 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0043] In the following description, a lot of specific details are set forth in order to facilitate a thorough understanding of the present specification. However, the present specification can be implemented in many different ways than those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present specification, so the present specification is not limited by the specific implementation disclosed below.
[0044] The terminology used in this disclosure, one or more embodiments of the present specification, is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in this disclosure and the appended claims herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used in this disclosure, refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0045] It will be understood that, although the terms first, second, etc. can be employed in this disclosure, one or more embodiments of the present specification, to describe various information, these information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information of the same type. For example, a first can also be referred to as a second, and similarly, a second can also be referred to as a first, without departing from the scope of one or more embodiments of the present specification. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining".
[0046] First, the noun terms related to one or more embodiments of the present specification are explained.
[0047] Image-text algorithm: Image-text algorithm refers to the general term of picture algorithm and text algorithm, which specifically includes picture classification, face recognition, target detection, picture retrieval and other algorithms taking picture as input, as well as text classification, sentiment analysis, machine translation, dialogue generation and other algorithms taking text as input.
[0048] Algorithm fairness: Artificial intelligence algorithm automated decision-making is independent of sensitive attributes (natural attributes and social attributes), that is, for protected sensitive attributes, artificial intelligence algorithm decision-making does not exist prejudice or favor for individuals or groups caused by their inherent or acquired attributes.
[0049] OOD data: OOD (Out-of-Distribution) data, also known as out-of-distribution data, refers to sample data from data different from the distribution of algorithm model training data. If the sample data distribution is the same as the distribution of algorithm model training data, the data is called ID data, i.e. in-distribution data.
[0050] Robustness: Robustness, also known as robustness or robustness, refers to the ability of a computer system to continue normal operation and ensure stable performance when encountering input, operation, etc. Abnormalities when processing errors under certain parameter (structure, size) changes or during execution.
[0051] With the continuous breakthroughs in the basic theories and technologies of artificial intelligence, the text and image algorithm based on artificial intelligence technology is widely used in finance, education, medical care, security and other public fields, giving birth to a series of intelligent applications such as intelligent security, intelligent customer service, medical diagnosis, personalized recommendation, etc. Not only greatly enrich and facilitate people's daily life, but also promote and promote the development and progress of social economy and technology.
[0052] However, with the increasing application scenarios, the legal, ethical issues and risks faced by artificial intelligence algorithms have become increasingly prominent. The unfairness and even discrimination in the automatic decision-making process of artificial intelligence have been controversial in society, not only arousing people's concerns and doubts about algorithmic automated decision-making, but also gradually attracting widespread attention from society and the public. Some places have introduced algorithm fairness-related laws and regulations, clearly stating that artificial intelligence algorithm research and application must meet the fairness constraints. Therefore, from the perspective of meeting regulatory compliance or improving user experience, algorithm fairness evaluation and eliminating algorithm bias are essential steps in the entire life cycle of the algorithm.
[0053] To solve the above problems, the embodiment of the present specification provides a fairness evaluation system, which has the ability to evaluate the fairness of part of artificial intelligence algorithm tasks (such as text classification, image classification, etc.). However, the fairness evaluation system only considers the fairness evaluation of natural input in a trusted environment, and the quantification of its fairness completely depends on the statistical indicators such as accuracy, recall rate, F1-score. In an uncontrolled environment (untrusted environment), these statistical indicators will be greatly affected by factors such as adversarial perturbation and data selection, so the fairness evaluation results produced by the above system will not accurately reflect the true fairness level of the algorithm (model) itself, and therefore cannot guarantee the effectiveness and usability of the evaluation.
[0054] Based on this, in the present specification, two model fairness evaluation methods are provided. One or more embodiments of the present specification simultaneously relate to two model fairness evaluation devices, a computing device, a computer readable storage medium and a computer program, which are described in detail one by one in the following embodiments.
[0055] Referring to Figure 1 , Figure 1 A specific scene schematic diagram of a model fairness evaluation method according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0056] Specifically, the model fairness evaluation method provided by the embodiment of the present specification is applied to a model fairness evaluation platform.
[0057] Step 102: Based on the large-scale image and text training samples collected in the database, a large-scale pre-training model (image and / or text training model) is trained using self-supervised learning technology, and the data probability distribution of the image and text training samples is initially modeled through the pre-training model; and based on the pre-training model, a generative adversarial network model is constructed using deep generation technology, and the data probability distribution of the image and text training samples is further optimized through the generative adversarial network model.
[0058] Among them, image and text training samples can be understood as image and text training samples.
[0059] Step 104: Receive a fairness assessment request sent by the user, which carries the sample to be evaluated and the model to be evaluated; first, perform a credibility check on the sample to be evaluated based on the data probability distribution of the image and text training sample and the generative adversarial network model; second, based on the credibility check result of the sample to be evaluated, if it is determined that the sample to be evaluated includes an adversarial sample, use the adversarial defense technology to denoise and adversarially reconstruct the adversarial sample; if it is determined that the distribution diversity test of the sample to be evaluated is weak, use the generative adversarial network model to generate diversity for the sample to be evaluated; finally, mix the original sample to be evaluated, the evaluation sample after adversarial reconstruction, and the sample generated by diversity, and input them into the fairness assessment module in combination with the model to be evaluated to obtain the fairness assessment result of the model to be evaluated.
[0060] Specifically, the fairness assessment of the model under evaluation in the embodiments of this specification can be understood as statistically analyzing model performance differences across different groups based on sensitive / protected attributes, such as false positive rate, statistical parity, equal opportunity, and inconsistent impact. These metrics can then be used to evaluate the fairness of the model under evaluation.
[0061] Step 106: Return the fairness evaluation result of the model to be evaluated to the user.
[0062] The model fairness assessment method provided in the embodiments of this specification proposes a robust fairness assessment system for graphic and text algorithms. It combines large-scale pre-training technology and deep generation technology to model data probability distribution, and performs reliability testing on the evaluation samples based on the probability distribution. At the same time, it performs denoising, adversarial reconstruction and diversity generation on untrusted samples such as adversarial samples or distribution deviation samples, thereby improving the reliability and integrity of evaluation samples in untrusted environments, thereby ensuring the robustness of the fairness assessment system in uncontrolled environments and the availability of evaluation results.
[0063] See also Figure 2 , Figure 2 A flowchart of a model fairness evaluation method provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0064] Step 202: training a model according to the picture and / or text to determine the real data probability distribution of the picture and / or text training sample.
[0065] Wherein, the picture includes but is not limited to any type, any size, and any content, such as pictures containing animals or people; the text includes but is not limited to any type, any length, and any content, such as academic discussions, literary articles, etc.
[0066] And the picture and / or text training model can be understood as a picture training model, a text training model, or a combination of picture and text training model, etc.; in actual application, the specific type of picture and / or text training model can be determined according to actual needs, and the embodiments of the present application do not make any limitation on this.
[0067] In specific implementation, in order to ensure the accuracy of the real data probability distribution of the picture and / or text training sample, a large-scale picture and / or text training sample is used to train the picture and / or text training model, and then the real data probability distribution of the picture and / or text training sample is modeled according to the trained picture and / or text training model; and the real data probability distribution is optimized by a generative adversarial network model to obtain the optimized real data probability distribution. The specific implementation is as follows:
[0068] The real data probability distribution of the picture and / or text training sample is determined according to the picture and / or text training model, including:
[0069] Obtaining a picture and / or text training sample;
[0070] Training a picture and / or text training model using a self-supervised learning technique according to the training sample;
[0071] Obtaining the adjusted real data probability distribution of the training sample according to the picture and / or text training model;
[0072] Adjusting the real data probability distribution of the training sample according to the generative adversarial network model to obtain the adjusted real data probability distribution of the training sample.
[0073] Wherein, in order to ensure the accuracy of the picture and / or text training model, in the embodiment of the present specification, a large number of picture and / or text training samples are obtained for model training; and when the training sample is a picture, the picture and / or text training model can be understood as a visual Transformer model, etc.; when the training sample is a text, the picture and / or text training model can be understood as a language model BERT model, etc.; when the training sample is a picture-text training sample, the picture and / or text training model can be understood as a combined multi-modal fusion model of visual Transformer model and language model BERT model.
[0074] Specifically, after obtaining a large number of picture and / or text training samples, a picture and / or text training model is trained using a self-supervised learning technique according to the large number of picture and / or text training samples; and then the real data probability distribution of the picture and / or text training sample is initially modeled by the picture and / or text training model. In order to further optimize the real data probability distribution of the picture and / or text training sample, the real data probability distribution of the picture and / or text training sample can be adjusted according to the generative adversarial network model to obtain the adjusted real data probability distribution of the picture and / or text training sample.
[0075] The model fairness evaluation method provided by the embodiment of the present specification first trains a picture and / or text training model through a large number of picture and / or text training samples, preliminarily models the real data probability distribution of the picture and / or text training sample according to the picture and / or text training model, and then optimizes the data probability distribution of the preliminarily modeled picture and / or text training sample according to the generative adversarial network model constructed by the deep generation technology, so as to determine the accuracy and availability of the picture and / or text training sample.
[0076] Before the real data probability distribution of the picture and / or text training sample is adjusted according to the generative adversarial network model, the generative adversarial network model needs to be constructed according to the picture and / or text training model using the deep generation technology to ensure the subsequent availability of the generative adversarial network model. The specific implementation mode is as follows:
[0077] The real data probability distribution of the training sample is adjusted according to the generative adversarial network model to obtain the adjusted real data probability distribution of the training sample, comprising:
[0078] The generative adversarial network model is constructed according to the picture and / or text training model;
[0079] The discriminant module and the generation module of the trained generative adversarial network model are obtained by training the generative adversarial network model according to the training sample;
[0080] According to the real data probability distribution of the training sample is adjusted by the discrimination module, the adjusted real data probability distribution of the training sample is obtained.
[0081] Specifically, first, the generative adversarial network model is constructed according to the picture and / or text training model; second, the generative adversarial network model is trained according to the picture and / or text training sample, and the discrimination module and the generation module of the trained generative adversarial network model are obtained; and finally, the initial probability distribution of the picture and / or text training sample is fine-tuned according to the discrimination module of the generative adversarial network model, and the fine-tuned real data probability distribution of the picture and / or text training sample is obtained.
[0082] In specific implementation, the construction stage of the generative adversarial network model includes two parts, the first part is the construction of the generative adversarial network model, and the second part is the training of the generative adversarial network model; wherein the generative adversarial network model is constructed by the discrimination module and the generation module, then in the construction of the generative adversarial network model, the picture and / or text training sample model trained in the above embodiment can be used as the discrimination module of the generative adversarial network model; and for the generation module, if the generated image data, a plurality of up-sampling deconvolution networks can be used to construct the generation module of the generative adversarial network model, and if the generated text data, the Transformer can be used as the generation module of the generative adversarial network model. The specific implementation mode is as follows:
[0083] The generative adversarial network model is constructed according to the picture and / or text training model, comprising:
[0084] According to the model parameters of the picture and / or text training model, the module parameters of the discrimination module of the generative adversarial network model are initialized, and the discrimination module of the generative adversarial network model is constructed;
[0085] According to the deconvolution network and / or text generation network, the generation module of the generative adversarial network model is constructed;
[0086] The generative adversarial network model is constructed according to the discrimination module and the generation module.
[0087] Wherein, the picture and / or text training sample model trained in the above embodiment is used as the discrimination module of the generative adversarial network model, which can be understood as that the module parameters of the discrimination module of the generative adversarial network model are initialized according to the model parameters of the picture and / or text training model, so as to construct the discrimination module of the generative adversarial network model; and the generation module can be constructed based on the type of the data to be generated, and the deconvolution network or the text generation network is selected; finally, the discrimination module and the generation module construct the generative adversarial network model.
[0088] After the generative adversarial network model is constructed, the generative adversarial network model can be trained. Specifically, the generation module and the discrimination module of the generative adversarial network model are alternately trained by constructing a zero-sum game adversarial loss function, so that the data generated by the generation module is closer to the real data distribution, and meanwhile, the discrimination module can better distinguish the real data and the generated data.
[0089] Specifically, in the embodiments of the present specification, the parameters of the pre-training model (i.e., the model parameters of the picture and / or text training model) are used to initialize the parameters of the discriminator (i.e., the discrimination module). The advantage is that the pre-training model is trained based on large-scale image-text training samples. By initializing the discriminator with the pre-training model, the knowledge learned by the pre-training model from large-scale training samples can be transferred to the discriminator, i.e., the pre-training and fine-tuning technology in deep learning is realized.
[0090] The model fairness evaluation method provided by the embodiments of the present specification constructs a generative adversarial network model according to the picture and / or text training model, and trains the generative adversarial network model according to the real data probability distribution of the picture and / or text training sample. Subsequently, the real data probability distribution of the picture and / or text training sample can be optimized according to the generative adversarial network model obtained by training, to obtain the adjusted real data probability distribution of the picture and / or text training sample, so as to improve the authenticity of the picture and / or text training sample.
[0091] Step 204: determining the credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model.
[0092] After obtaining the adjusted real data probability distribution of the picture and / or text training sample, the credibility of the to-be-evaluated sample can be detected in combination with the generative adversarial network model.
[0093] Specifically, the determining of the credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model comprises:
[0094] obtaining a sample data probability distribution of the to-be-evaluated sample according to the picture and / or text training model;
[0095] determining a similarity of the to-be-evaluated sample to the real data probability distribution of the training sample according to the sample data probability distribution;
[0096] obtaining a sample prediction result of the to-be-evaluated sample according to the discrimination module of the generative adversarial network model;
[0097] determining the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result.
[0098] wherein, for the convenience of understanding, the real data probability distribution in the following embodiments can be understood as the real data probability distribution adjusted by the picture and / or text training sample; and the sample prediction result of the to-be-evaluated sample can be understood as that the to-be-evaluated sample is a generated sample or a real sample.
[0099] In a specific implementation, first, the sample data probability distribution of the to-be-evaluated sample is obtained according to the picture and / or text training model, the similarity (i.e., the log-likelihood) of the to-be-evaluated sample belonging to the real data (picture and / or text training sample) distribution is calculated according to the sample data probability distribution of the to-be-evaluated sample; at the same time, the sample prediction result of the to-be-evaluated sample is obtained according to the discriminator module of the generative adversarial network model; and then the credibility detection result of the to-be-evaluated sample is determined according to the similarity and the sample prediction result.
[0100] The model fairness evaluation method provided by the embodiments of the present specification performs credibility detection on the to-be-evaluated sample according to the target probability distribution of the picture and / or text training sample and the generative adversarial network model, so as to determine whether the to-be-evaluated sample includes an adversarial sample or the diversity of the to-be-evaluated sample is weak, etc. In the case where the credibility of the to-be-evaluated sample is determined to have a problem, the to-be-evaluated sample can be processed subsequently, thereby improving the reliability and integrity of the to-be-evaluated sample in a non-credible environment.
[0101] In actual application, the credibility detection of the to-be-evaluated sample can be understood as the detection of whether the to-be-evaluated sample is an adversarial sample and the distribution diversity of the to-be-evaluated sample. The specific implementation manner is as follows:
[0102] The determining of the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result comprises:
[0103] The determining of whether the to-be-evaluated sample is an adversarial sample and the distribution diversity of the to-be-evaluated sample according to the similarity and the sample prediction result.
[0104] In actual application, after the log-likelihood and the sample prediction result are obtained, it can be determined according to the log-likelihood and the sample prediction result whether the to-be-evaluated sample is an OOD sample, such as an adversarial sample. In theory, the smaller the log-likelihood (in actual application, a log-likelihood threshold can be set for the sample set and the model task, and if the log-likelihood is smaller than the threshold, it can be considered that the log-likelihood is smaller), the greater the probability that the to-be-evaluated sample is an adversarial sample.
[0105] For the detection of the distribution diversity of the to-be-evaluated sample, the distribution of the logarithmic likelihood of the to-be-evaluated sample belonging to the real data distribution is calculated. If the distribution is more dispersed, it indicates that the distribution diversity of the to-be-evaluated sample is stronger, and if the distribution is more concentrated, it indicates that the distribution diversity of the to-be-evaluated sample is weaker. In actual applications, the detection of the distribution diversity of the to-be-evaluated sample is the detection of the distribution diversity of the entire to-be-evaluated sample, rather than the measurement of a single to-be-evaluated sample. The distribution diversity of the to-be-evaluated sample can be detected by using indicators such as variance, standard deviation, median, and central tendency to measure the dispersion of the distribution, and the embodiments of the present specification do not make any limitation in this regard.
[0106] The model fairness evaluation method provided by the embodiments of the present specification can detect the credibility of the to-be-evaluated sample according to the logarithmic likelihood of the to-be-evaluated sample belonging to the real data distribution and the sample prediction result after obtaining the logarithmic likelihood and the sample prediction result. That is, whether the to-be-evaluated sample is an adversarial sample and the distribution diversity of the to-be-evaluated sample are detected. When it is determined that the to-be-evaluated sample is an adversarial sample or the distribution diversity of the to-be-evaluated sample is weak, it can be determined that the to-be-evaluated sample is not credible. Subsequently, the to-be-evaluated sample that is not credible can be processed to improve the reliability and integrity of the to-be-evaluated sample in a non-credible environment.
[0107] Step 206: In the case where the credibility detection result meets the non-credible condition, performing sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample.
[0108] Specifically, in the case where it is determined that the to-be-evaluated sample is an adversarial sample, the to-be-evaluated sample can be reconstructed by denoising and adversarial reconstruction. In the case where it is determined that the distribution diversity of the to-be-evaluated sample is weak, the to-be-evaluated sample can be generated in diversity. The specific implementation modes are as follows:
[0109] In the case where the credibility detection result meets the non-credible condition, performing sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample, including:
[0110] In the case where the to-be-evaluated sample is an adversarial sample, performing sample processing on the to-be-evaluated sample according to a first preset processing mode to obtain a first updated evaluation sample; and / or
[0111] In the case where the distribution diversity of the to-be-evaluated sample meets a preset distribution condition, performing sample processing on the to-be-evaluated sample according to a second preset processing mode to obtain a second updated evaluation sample.
[0112] The first preset processing mode and the second preset processing mode can be set according to actual applications, and the embodiments of the present specification do not make any limitation in this regard.
[0113] In actual application, the trustworthiness detection of the to-be-evaluated sample can detect whether the to-be-evaluated sample is an adversarial sample and whether the distribution diversity of the to-be-evaluated sample is weak. In the case that the to-be-evaluated sample is an adversarial sample or the distribution diversity of the to-be-evaluated sample is weak, the to-be-evaluated sample can be considered as untrustworthy. At this time, the to-be-evaluated sample needs to be processed.
[0114] In specific implementation, in the case that the trustworthiness detection result of the to-be-evaluated sample is that the to-be-evaluated sample is an adversarial sample, it can be determined that the trustworthiness detection result of the to-be-evaluated sample satisfies the untrustworthiness condition. At this time, the to-be-evaluated sample can be processed according to the first preset processing mode to obtain a first updated evaluation sample. In the case that the trustworthiness detection result of the to-be-evaluated sample is that the distribution diversity of the to-be-evaluated sample is weak, it can also be determined that the trustworthiness detection result of the to-be-evaluated sample satisfies the untrustworthiness condition. At this time, the to-be-evaluated sample can be processed according to the second preset processing mode to obtain a second updated evaluation sample.
[0115] Then, in the case that the first preset processing mode includes data compression, data randomization or an adversarial error correction denoising method, the specific manner of processing the to-be-evaluated sample according to the first preset processing mode to obtain the first updated evaluation sample is as follows:
[0116] The processing manner of processing the to-be-evaluated sample according to the first preset processing mode to obtain the first updated evaluation sample in the case that the to-be-evaluated sample is an adversarial sample includes:
[0117] In the case that the to-be-evaluated sample is an adversarial sample, the to-be-evaluated sample is reconstructed by a data compression, data randomization or adversarial error correction denoising method to obtain the first updated evaluation sample.
[0118] In the case that the second preset processing method is to generate a new evaluation sample, the specific processing manner of processing the to-be-evaluated sample according to the second preset processing mode to obtain the second updated evaluation sample is as follows:
[0119] The processing manner of processing the to-be-evaluated sample according to the second preset processing mode to obtain the second updated evaluation sample in the case that the distribution diversity of the to-be-evaluated sample satisfies the preset distribution condition includes:
[0120] In the case that the distribution diversity of the to-be-evaluated sample satisfies the preset distribution condition, a new evaluation sample is generated by a generation module of the generative adversarial network model according to the sample data probability distribution of the to-be-evaluated sample.
[0121] According to the new evaluation sample, a second updated evaluation sample is obtained.
[0122] Specifically, in the case that the credibility of the to-be-evaluated sample is weak due to the adversarial sample and the weak distribution diversity, the to-be-evaluated sample can be subjected to adversarial reconstruction and diversity generation. The adversarial reconstruction can be understood as reconstructing the adversarial sample detected in the to-be-evaluated sample by using a denoising method including but not limited to data compression, data randomization, adversarial error correction, etc., to eliminate the interference of the adversarial noise. The diversity generation can be understood as generating a new evaluation sample based on the sample data probability distribution of the to-be-evaluated sample by using the generation module of the generative adversarial network model, to expand the diversity of the to-be-evaluated sample.
[0123] The model fairness evaluation method provided by the embodiments of the present specification can detect the credibility of the to-be-evaluated sample according to the log-likelihood and the sample prediction result after obtaining the log-likelihood and the sample prediction result of the to-be-evaluated sample belonging to the real data distribution, that is, whether the to-be-evaluated sample is an adversarial sample and the distribution diversity of the to-be-evaluated sample. When it is determined that the to-be-evaluated sample is an adversarial sample or the distribution diversity of the to-be-evaluated sample is weak, it can be determined that the to-be-evaluated sample is not credible. Subsequently, the data processing can be performed on the to-be-evaluated sample that is not credible, to improve the reliability and integrity of the to-be-evaluated sample in the non-credible environment, and thus the robustness of the fairness evaluation method in the non-credible environment and the availability of the evaluation result are ensured.
[0124] In addition, in order to ensure the accuracy of the new evaluation sample, after generating the new evaluation sample by using the generation module of the generative adversarial network model, the quality of the new evaluation sample is detected by using the discrimination module of the generative adversarial network model, to ensure the accuracy of the new evaluation sample. The specific implementation manner is as follows:
[0125] The second updated evaluation sample is obtained according to the new evaluation sample.
[0126] The prediction result of the new evaluation sample is obtained by inputting the new evaluation sample into the discrimination module of the generative adversarial network model.
[0127] The new evaluation sample is pruned according to the prediction result of the new evaluation sample, to obtain the second updated evaluation sample.
[0128] The model fairness evaluation method provided by the embodiments of the present specification can ensure the accuracy of the new evaluation samples. After the new evaluation samples are generated by the generation module of the generative adversarial network model, the new evaluation samples are filtered based on the generation quality of the new evaluation samples to ensure the accuracy of the new evaluation samples. In addition, the generation quality of the new evaluation samples can also be filtered according to the quality evaluation indicators of the original to-be-evaluated samples. For example, when the original to-be-evaluated samples are images, the quality of the images can be filtered according to the Inception Score (initial score). When the original to-be-evaluated samples are texts, the quality of the texts can be filtered according to the fluency of the texts, and so on.
[0129] Step 208: According to the to-be-evaluated sample and the updated evaluation sample, the fairness of the to-be-evaluated model is evaluated.
[0130] The updated evaluation sample includes the first updated evaluation sample and / or the second updated evaluation sample. The to-be-evaluated model can also be understood as a type of image-text recognition model.
[0131] When the updated evaluation sample includes the first updated evaluation sample, the to-be-evaluated sample and the first updated evaluation sample are mixed to generate a mixed sample, and then the fairness of the to-be-evaluated model is evaluated. When the updated evaluation sample includes the second updated evaluation sample, the to-be-evaluated sample and the second updated evaluation sample are mixed to generate a mixed sample, and then the fairness of the to-be-evaluated model is evaluated. When the updated evaluation sample includes the first updated evaluation sample and the second updated evaluation sample, the to-be-evaluated sample, the first updated evaluation sample, and the second updated evaluation sample are mixed to generate a mixed sample, and then the fairness of the to-be-evaluated model is evaluated. The specific implementation mode is as follows:
[0132] The fairness of the to-be-evaluated model is evaluated according to the to-be-evaluated sample and the updated evaluation sample, including:
[0133] The to-be-evaluated sample and the updated evaluation sample are mixed to obtain a mixed evaluation sample.
[0134] The mixed evaluation sample and the to-be-evaluated model are input into a fairness evaluation module to obtain a fairness evaluation indicator of the to-be-evaluated model.
[0135] The fairness of the to-be-evaluated model is evaluated according to the fairness evaluation indicator of the to-be-evaluated model.
[0136] The fairness evaluation indicator includes but is not limited to false positive rate, statistical equality, opportunity equality, inconsistent impact, and the like.
[0137] In actual implementation, the mixed evaluation sample and the to-be-evaluated model are input into the fairness evaluation module, the mixed evaluation sample is input into the to-be-evaluated model in the fairness evaluation module, the prediction value of the mixed evaluation sample output by the to-be-evaluated model is obtained, the prediction accuracy of the to-be-evaluated model is calculated according to the prediction value and the true value of the mixed evaluation sample, and the false positive rate, statistical equality, opportunity equality, and inconsistent influence of the to-be-evaluated model are calculated according to the prediction accuracy after the fairness evaluation module determines the prediction accuracy of the to-be-evaluated model. Subsequent users or systems can evaluate the fairness of the to-be-evaluated model according to the indexes.
[0138] The mixed evaluation sample and the to-be-evaluated model are input into the fairness evaluation module.
[0139] The fairness evaluation module outputs the fairness evaluation indexes of the to-be-evaluated model determined according to the comparison result of the true value and the prediction value of the mixed evaluation sample,
[0140] The prediction value is output by the to-be-evaluated model according to the mixed evaluation sample.
[0141] Specifically, the comparison result of the true value and the prediction value of the mixed evaluation sample can be understood as the prediction accuracy of the to-be-evaluated model.
[0142] In actual application, the mixed evaluation sample and the to-be-evaluated model are input into the fairness evaluation module, the mixed evaluation sample is input into the to-be-evaluated model in the fairness evaluation module, the prediction value of the mixed evaluation sample output by the to-be-evaluated model is obtained, the prediction accuracy of the to-be-evaluated model is calculated according to the prediction value and the true value of the mixed evaluation sample, and the false positive rate, statistical equality, opportunity equality, and inconsistent influence of the to-be-evaluated model are calculated according to the prediction accuracy after the fairness evaluation module determines the prediction accuracy of the to-be-evaluated model. Subsequent users or systems can evaluate the fairness of the to-be-evaluated model according to the indexes.
[0143] The model fairness evaluation method provided in the specification provides a graph-text training model, models the true data probability distribution of the graph-text training sample, performs credibility detection on the to-be-evaluated sample according to the true data probability distribution, processes the non-credible sample, obtains the updated evaluation sample, and improves the reliability and completeness of the evaluation sample in the non-credible environment, thereby ensuring the robustness of the model fairness evaluation method in the credible environment and the non-credible environment and the availability of the model evaluation result, thereby ensuring the accuracy of the model fairness evaluation method in evaluating the fairness of the to-be-evaluated model, and making the subsequent to-be-evaluated model have good effects in actual application from the aspects of meeting regulatory compliance and improving user experience.
[0144] The following describes the model fairness evaluation method provided in the specification in combination with the accompanying Figure 3 The model fairness evaluation method provided in the specification is used for evaluating the fairness of a recommendation model, and the model fairness evaluation method is further described. Wherein, Figure 3 FIG. 1 shows a process flowchart of a model fairness evaluation method according to an embodiment of the specification, specifically including the following steps.
[0145] Step 302: According to the collected large-scale unsupervised graph-text training samples, a pre-training model is trained by using a self-supervised learning technique, and the real data probability distribution of the graph-text training samples is initially modeled by the pre-training model.
[0146] Step 304: According to the pre-training model, a generative adversarial network model is constructed by using a deep generation technique, the generative adversarial network model is trained according to the graph-text training samples, and the real data probability distribution of the graph-text training samples is optimized according to the generative adversarial network model, to obtain an optimized real data probability distribution of the graph-text training samples.
[0147] Step 306: According to the optimized real data probability distribution of the graph-text training samples and the generative adversarial network model, the credibility of the to-be-evaluated sample is detected.
[0148] The specific implementation of detecting the credibility of the to-be-evaluated sample according to the optimized real data probability distribution of the graph-text training samples and the generative adversarial network model can be referred to the detailed description of the above embodiments, and will not be repeated here.
[0149] Step 308: According to the credibility detection result of the to-be-evaluated sample, the to-be-evaluated sample is reconstructed by using an adversarial defense technique, or the to-be-evaluated sample is generated by using a generator of the generative adversarial network model.
[0150] The specific implementation of the adversarial reconstruction and the diversity generation of the to-be-evaluated sample can be referred to the detailed description of the above embodiments, and will not be repeated here.
[0151] Step 310: The to-be-evaluated sample is mixed with the adversarial reconstructed to-be-evaluated sample and / or the diversity generated to-be-evaluated sample to obtain a mixed sample, and the mixed sample and the recommendation model are input into a fairness evaluation module for evaluation to obtain a fairness evaluation result of the recommendation model.
[0152] The model fairness evaluation method provided by the embodiments of the present specification proposes an algorithm fairness evaluation technology for an uncontrolled (credible) environment, models a real data probability distribution by using a combination of large-scale pre-training technology and deep generation technology based on large-scale easily accessible unsupervised graph-text training data, detects untrusted samples in a to-be-evaluated sample based on the modeled real data (unsupervised graph-text training data) probability distribution, and simultaneously performs adversarial denoising reconstruction by using an adversarial defense technique, which can effectively eliminate the influence of adversarial noise on fairness evaluation.
[0153] For the problem of distribution bias of the data to be evaluated (the sample distribution is single and cannot cover the entire data distribution), the model fairness evaluation method of the embodiment of this specification proposes a data distribution based on the sample to be evaluated itself, and combines it with deep generation technology for diversity generation, which can achieve the purpose of improving the reliability and integrity of the sample to be evaluated in an untrusted environment. In addition, in the entire fairness evaluation process, the model fairness evaluation method provided by the embodiment of this specification does not require additional evaluation data or manual intervention, which greatly improves the intelligence level of the evaluation and reduces the evaluation cost. In summary, the model fairness evaluation method provided by the embodiment of this specification not only has the ability to evaluate fairness in a trusted environment, but also can ensure the robustness of fairness evaluation in an uncontrolled environment and the availability of evaluation results. Therefore, it will be applicable to fairness evaluation of algorithms including but not limited to intelligent customer service, personalized recommendation, intelligent risk control, etc. on e-commerce platforms, online social platforms, online social media and other platforms, so as to eliminate algorithm bias and encourage the algorithm to meet regulatory compliance and improve user experience.
[0154] Corresponding to the above method embodiment, this specification also provides an embodiment of a model fairness evaluation device, Figure 4 FIG1 shows a schematic diagram of the structure of a model fairness evaluation device provided by an embodiment of this specification. Figure 4 As shown, the device includes:
[0155] The probability distribution determination module 402 is configured to determine the real data probability distribution of the image and / or text training samples based on the image and / or text training model;
[0156] The detection result determination module 404 is configured to determine the credibility detection result of the sample to be evaluated based on the probability distribution of the real data and the generative adversarial network model;
[0157] The sample processing module 406 is configured to perform sample processing on the sample to be evaluated according to the credibility detection result to obtain an updated evaluation sample when the credibility detection result satisfies the non-credible condition;
[0158] The evaluation module 408 is configured to perform fairness evaluation on the model to be evaluated based on the sample to be evaluated and the updated evaluation sample.
[0159] Optionally, the probability distribution determination module 402 is further configured to:
[0160] Obtain image and / or text training samples;
[0161] Using self-supervised learning technology according to the training samples, training to obtain an image and / or text training model;
[0162] According to the picture and / or text training model, the real data probability distribution of the training sample is adjusted to obtain an adjusted real data probability distribution of the training sample.
[0163] According to the generated adversarial network model, the real data probability distribution of the training sample is adjusted to obtain an adjusted real data probability distribution of the training sample.
[0164] Optionally, the probability distribution determination module 402 is further configured to:
[0165] According to the picture and / or text training model, the generated adversarial network model is constructed;
[0166] According to the training sample, the generated adversarial network model is trained to obtain a discriminant module and a generation module of the trained generated adversarial network model;
[0167] According to the discriminant module, the real data probability distribution of the training sample is adjusted to obtain an adjusted real data probability distribution of the training sample.
[0168] Optionally, the probability distribution determination module 402 is further configured to:
[0169] According to the model parameters of the picture and / or text training model, the module parameters of the discriminant module of the generated adversarial network model are initialized to construct the discriminant module of the generated adversarial network model;
[0170] According to the deconvolution network and / or text generation network, the generation module of the generated adversarial network model is constructed;
[0171] The generated adversarial network model is constructed according to the discriminant module and the generation module.
[0172] Optionally, the detection result determination module 404 is further configured to:
[0173] According to the picture and / or text training model, the sample data probability distribution of the to-be-evaluated sample is obtained;
[0174] According to the sample data probability distribution, the similarity of the to-be-evaluated sample belonging to the real data probability distribution of the training sample is determined;
[0175] According to the discriminant module of the generated adversarial network model, the sample prediction result of the to-be-evaluated sample is obtained;
[0176] According to the similarity and the sample prediction result, the credibility detection result of the to-be-evaluated sample is determined.
[0177] Optionally, the detection result determination module 404 is further configured to:
[0178] determine whether the to-be-evaluated sample is an adversarial sample and distribution diversity of the to-be-evaluated sample according to the similarity and the sample prediction result.
[0179] Optionally, the sample processing module 406 is further configured to:
[0180] in a case where the to-be-evaluated sample is an adversarial sample, performing sample processing on the to-be-evaluated sample according to a first preset processing mode to obtain a first updated evaluation sample; and / or
[0181] in a case where the distribution diversity of the to-be-evaluated sample meets a preset distribution condition, performing sample processing on the to-be-evaluated sample according to a second preset processing mode to obtain a second updated evaluation sample.
[0182] Optionally, the sample processing module 406 is further configured to:
[0183] in a case where the to-be-evaluated sample is an adversarial sample, reconstructing the to-be-evaluated sample by a data compression, data randomization or adversarial error correction denoising method to obtain a first updated evaluation sample.
[0184] Optionally, the sample processing module 406 is further configured to:
[0185] in a case where the distribution diversity of the to-be-evaluated sample meets a preset distribution condition, generating a new evaluation sample by a generation module of the generative adversarial network model according to a sample data probability distribution of the to-be-evaluated sample;
[0186] obtaining a second updated evaluation sample according to the new evaluation sample.
[0187] Optionally, the sample processing module 406 is further configured to:
[0188] inputting the new evaluation sample into a discriminant module of the generative adversarial network model to obtain a prediction result of the new evaluation sample;
[0189] performing pruning on the new evaluation sample according to the prediction result of the new evaluation sample to obtain a second updated evaluation sample.
[0190] Optionally, the evaluation module 408 is further configured to:
[0191] mixing the to-be-evaluated sample and the updated evaluation sample to obtain a mixed evaluation sample;
[0192] inputting the mixed evaluation sample and a to-be-evaluated model into a fairness evaluation module to obtain a fairness evaluation index of the to-be-evaluated model;
[0193] According to the fairness evaluation index of the to-be-evaluated model, the fairness of the to-be-evaluated model is evaluated.
[0194] Optionally, the evaluation module 408 is further configured to:
[0195] inputting the mixed evaluation sample and the to-be-evaluated model into a fairness evaluation module;
[0196] receiving the fairness evaluation index of the to-be-evaluated model determined by the fairness evaluation module according to the comparison result of the true value and the predicted value of the mixed evaluation sample,
[0197] wherein the predicted value is output by the to-be-evaluated model according to the mixed evaluation sample.
[0198] The model fairness evaluation device provided by the embodiment of the present specification improves the reliability and completeness of the evaluation sample in the untrusted environment by modeling the true data probability distribution of the image-text training sample through the image-text training model, performing credibility detection on the to-be-evaluated sample according to the true data probability distribution, and processing the untrusted sample to obtain an updated evaluation sample, thereby ensuring the robustness of the model fairness evaluation method in the trusted environment and the untrusted environment and the availability of the model evaluation result, thereby ensuring the accuracy of the fairness evaluation of the to-be-evaluated model by the model fairness evaluation method, and making the subsequent to-be-evaluated model have good effects in actual application from the aspects of meeting regulatory compliance and improving user experience.
[0199] The above is a schematic scheme of the model fairness evaluation device of the embodiment. It should be noted that the technical scheme of the model fairness evaluation device belongs to the same concept as the technical scheme of the model fairness evaluation method described above, and the details of the technical scheme of the model fairness evaluation device that are not described in detail can be referred to the description of the technical scheme of the model fairness evaluation method.
[0200] Referring to Figure 5 , Figure 5 A flowchart of another model fairness evaluation method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0201] Step 502: According to the image and / or text training model, the true data probability distribution of the image and / or text training sample is determined.
[0202] Step 504: Receive the to-be-evaluated sample and the to-be-evaluated model sent by the user.
[0203] Step 506: According to the true data probability distribution and the generative adversarial network model, the credibility detection result of the to-be-evaluated sample is determined.
[0204] Step 508: in the case that the trustworthiness detection result meets the untrustworthy condition, performing sample processing on the to-be-evaluated sample according to the trustworthiness detection result to obtain an updated evaluation sample.
[0205] Step 510: performing fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample.
[0206] Step 512: obtaining a fairness evaluation result of the to-be-evaluated model and returning the fairness evaluation result to the user.
[0207] The model fairness evaluation method provided by the embodiment of the present specification is applied to a model fairness evaluation platform.
[0208] In an actual application scenario, a user wants to perform fairness evaluation on a project model in a model fairness evaluation platform. The user can then send the to-be-evaluated sample and the to-be-evaluated model to the model fairness evaluation platform. After receiving the to-be-evaluated sample and the to-be-evaluated model sent by the user, the model fairness evaluation platform can perform adversarial reconstruction or diversity generation on the to-be-evaluated sample according to the above-mentioned embodiment, so as to ensure that the evaluation result given by the model fairness evaluation platform is closer to the real situation of the to-be-evaluated model.
[0209] The model fairness evaluation method provided by the embodiment of the present specification first models and trains data probability distribution based on large-scale easily accessible unsupervised image-text training data, using a combination of large-scale pre-training technology and deep generation technology, to serve as real data probability distribution. For the to-be-evaluated sample, before entering the fairness evaluation module, the to-be-evaluated sample will be subjected to reliability detection according to the real data probability distribution, effectively alleviating the interference of untrustworthy samples on the evaluation result. For the detected untrustworthy samples such as adversarial samples or distribution deviation samples, the present scheme respectively performs denoising, reconstruction on the evaluation sample based on adversarial defense technology and diversity generation based on a deep generation model, achieving the purpose of improving the reliability and integrity of the evaluation sample in an untrustworthy environment, ensuring the robustness of the fairness evaluation system in an uncontrolled environment and the availability of the evaluation result, effectively making up for the shortcomings of existing systems that are sensitive to adversarial disturbance and evaluation sample distribution deviation in an uncontrolled environment.
[0210] Corresponding to the above method embodiment, the present specification also provides another model fairness evaluation device embodiment, Figure 6 shows the structure schematic diagram of another model fairness evaluation device provided by an embodiment of the present specification. As shown in the figure, Figure 6 the device is applied to a model fairness evaluation platform and includes:
[0211] The first determination module 602 is configured to determine a real data probability distribution of the picture and / or text training sample according to a picture and / or text training model.
[0212] The data receiving module 604 is configured to receive a to-be-evaluated sample and a to-be-evaluated model sent by a user.
[0213] The second determination module 606 is configured to determine a credibility detection result of the to-be-evaluated sample according to the real data probability distribution and a generative adversarial network model.
[0214] The sample updating module 608 is configured to perform sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample in a case where the credibility detection result satisfies an untrusted condition.
[0215] The fairness evaluation module 610 is configured to perform fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample.
[0216] The result display module 612 is configured to obtain a fairness evaluation result of the to-be-evaluated model and return the fairness evaluation result to the user.
[0217] The model fairness evaluation device provided by the embodiment of the present specification first models and trains a data probability distribution as a real data probability distribution by using a combination of large-scale pre-training technology and deep generation technology based on large-scale easily accessible unsupervised picture-text training data. For a to-be-evaluated sample, a reliability detection is performed on the to-be-evaluated sample according to the real data probability distribution before the to-be-evaluated sample enters a fairness evaluation module, which effectively alleviates the interference of untrusted samples on the evaluation result. For the detected untrusted samples such as adversarial samples or distribution deviation samples, the present solution respectively performs denoising, reconstruction on the evaluation sample based on an adversarial defense technology and diversity generation based on a deep generation model, which achieves the purpose of improving the reliability and integrity of the evaluation sample in an untrusted environment, guarantees the robustness of the fairness evaluation system in an uncontrolled environment and the availability of the evaluation result, and effectively makes up for the shortcomings of the existing system that is sensitive to adversarial disturbance and evaluation sample distribution deviation in an uncontrolled environment.
[0218] The above is a schematic solution of the model fairness evaluation device of the embodiment. It should be noted that the technical solution of the model fairness evaluation device belongs to the same concept as the technical solution of the model fairness evaluation method described above, and the details of the technical solution of the model fairness evaluation device that are not described in detail can be referred to the description of the technical solution of the model fairness evaluation method.
[0219] Figure 7A structural block diagram of a computing device 700 according to one embodiment of the present specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected with the memory 710 through a bus 730, and a database 750 is used to save data.
[0220] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include the public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 can include one or more of any type of network interface (e.g., network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like, either wired or wireless.
[0221] In one embodiment of the present specification, the above-mentioned components of the computing device 700 and other components not shown in the above-mentioned components can be connected with each other, for example, through a bus. It should be understood that, Figure 7 the computing device structure block diagram shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art. Figure 7
[0222] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC. The computing device 700 can also be a mobile or stationary server.
[0223] The processor 720 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned model fairness evaluation method.
[0224] The above is a schematic scheme of a computing device according to the present embodiment. It should be noted that the technical scheme of the computing device belongs to the same concept as the technical scheme of the above-mentioned model fairness evaluation method, and the details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the above-mentioned model fairness evaluation method.
[0225] The embodiment of the present specification also provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions realize the steps of the model fairness evaluation method when executed by a processor.
[0226] The above is a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the model fairness evaluation method belong to the same concept, and the details of the technical scheme of the storage medium which are not described in detail can be referred to the description of the technical scheme of the model fairness evaluation method.
[0227] The embodiment of the present specification also provides a computer program, which causes a computer to execute the steps of the model fairness evaluation method when the computer program is executed in the computer.
[0228] The above is a schematic scheme of the computer program of the embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the model fairness evaluation method belong to the same concept, and the details of the technical scheme of the computer program which are not described in detail can be referred to the description of the technical scheme of the model fairness evaluation method.
[0229] The specific embodiments of the present specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order in order to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0230] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0231] It should be noted that, for the aforementioned method embodiments, the sequences of the described actions are not necessarily required to implement the present application, and certain actions can be performed in other sequences, or even at the same time, in accordance with the present application. Furthermore, certain actions can not be required to implement the present application. Additionally, the described embodiments are not necessarily the only possible implementation of the present application.
[0232] In the above embodiments, the description of each embodiment is focused on the aspects of the embodiment, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0233] The preferred embodiments of the present application disclosed above are only used to clarify the present application. The alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, according to the content of the embodiments of the present application, many modifications and changes can be made. The embodiments are selected and described in detail in order to better explain the principles and practical applications of the embodiments of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their full scope and equivalents.
Claims
1. A model fairness evaluation method, comprising: training a picture and / or text model to determine a real data probability distribution of picture and / or text training samples; determining a credibility detection result of a to-be-evaluated sample according to the real data probability distribution and a generative adversarial network model; in a case where the credibility detection result meets an untrusted condition, performing sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample; performing fairness evaluation on a to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample; wherein the determining of the credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model comprises: obtaining a sample data probability distribution of the to-be-evaluated sample according to the picture and / or text model; determining a similarity of the to-be-evaluated sample to the real data probability distribution of the training sample according to the sample data probability distribution; obtaining a sample prediction result of the to-be-evaluated sample according to a discriminant module of the generative adversarial network model; and determining the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result.
2. The model fairness evaluation method of claim 1, wherein the determining of the real data probability distribution of the picture and / or text training sample according to the picture and / or text model comprises: obtaining picture and / or text training samples; training a picture and / or text model using a self-supervised learning technique according to the training samples; obtaining an adjusted real data probability distribution of the training samples according to the picture and / or text model; and adjusting the real data probability distribution of the training samples according to a generative adversarial network model to obtain an adjusted real data probability distribution of the training samples.
3. The model fairness evaluation method of claim 2, wherein the adjusting of the real data probability distribution of the training samples according to the generative adversarial network model to obtain an adjusted real data probability distribution of the training samples comprises: constructing a generative adversarial network model according to the picture and / or text model; training the generative adversarial network model according to the training samples to obtain a discriminant module and a generative module of the trained generative adversarial network model; adjusting the real data probability distribution of the training samples according to the discriminant module to obtain an adjusted real data probability distribution of the training samples.
4. The model fairness evaluation method of claim 3, wherein the constructing of the generative adversarial network model according to the picture and / or text model comprises: initializing module parameters of a discriminant module of the generative adversarial network model according to model parameters of the picture and / or text model to construct the discriminant module of the generative adversarial network model; constructing a generative module of the generative adversarial network model according to a deconvolution network and / or a text generative network; and constructing the generative adversarial network model according to the discriminant module and the generative module.
5. The model fairness evaluation method according to claim 1, wherein the determining the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result comprises: determining whether the to-be-evaluated sample is an adversarial sample and the distribution diversity of the to-be-evaluated sample according to the similarity and the sample prediction result.
6. The model fairness evaluation method according to claim 5, wherein the sample processing of the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample when the credibility detection result meets an untrusted condition comprises: performing sample processing on the to-be-evaluated sample according to a first preset processing mode to obtain a first updated evaluation sample when the to-be-evaluated sample is an adversarial sample; and / or performing sample processing on the to-be-evaluated sample according to a second preset processing mode to obtain a second updated evaluation sample when the distribution diversity of the to-be-evaluated sample meets a preset distribution condition.
7. The model fairness evaluation method according to claim 6, wherein the performing sample processing on the to-be-evaluated sample according to a first preset processing mode to obtain a first updated evaluation sample when the to-be-evaluated sample is an adversarial sample comprises: reconstructing the to-be-evaluated sample by data compression, data randomization or an adversarial error correction denoising method to obtain the first updated evaluation sample.
8. The model fairness evaluation method according to claim 6, wherein the performing sample processing on the to-be-evaluated sample according to a second preset processing mode to obtain a second updated evaluation sample when the distribution diversity of the to-be-evaluated sample meets a preset distribution condition comprises: generating a new evaluation sample by a generation module of the generative adversarial network model according to the sample data probability distribution of the to-be-evaluated sample when the distribution diversity of the to-be-evaluated sample meets the preset distribution condition; and obtaining the second updated evaluation sample according to the new evaluation sample.
9. The model fairness evaluation method according to claim 8, wherein the obtaining the second updated evaluation sample according to the new evaluation sample comprises: inputting the new evaluation sample into a discriminant module of the generative adversarial network model to obtain a prediction result of the new evaluation sample; and performing pruning on the new evaluation sample according to the prediction result of the new evaluation sample to obtain the second updated evaluation sample.
10. The model fairness evaluation method according to claim 1, wherein the performing fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample comprises: mixing the to-be-evaluated sample and the updated evaluation sample to obtain mixed evaluation samples; inputting the mixed evaluation samples and the to-be-evaluated model into a fairness evaluation module to obtain a fairness evaluation index of the to-be-evaluated model; and performing fairness evaluation on the to-be-evaluated model according to the fairness evaluation index of the to-be-evaluated model. 11. The model fairness evaluation method according to claim 10, wherein the mixed evaluation sample and the to-be-evaluated model are input into a fairness evaluation module to obtain a fairness evaluation index of the to-be-evaluated model, comprising: inputting the mixed evaluation sample and the to-be-evaluated model into the fairness evaluation module; receiving the fairness evaluation index of the to-be-evaluated model determined by the fairness evaluation module according to the comparison result of the true value and the predicted value of the mixed evaluation sample, wherein the predicted value is output by the to-be-evaluated model according to the mixed evaluation sample.
12. A model fairness evaluation device, comprising: a probability distribution determination module configured to determine a true data probability distribution of a picture and / or text training sample according to a picture and / or text training model; a detection result determination module configured to determine a credibility detection result of a to-be-evaluated sample according to the true data probability distribution and a generative adversarial network model; a sample processing module configured to perform sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample if the credibility detection result meets an untrusted condition; an evaluation module configured to perform fairness evaluation on a to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample; wherein the detection result determination module is further configured to: obtain a sample data probability distribution of the to-be-evaluated sample according to the picture and / or text training model; determine a similarity of the to-be-evaluated sample to the true data probability distribution of the training sample according to the sample data probability distribution; obtain a sample prediction result of the to-be-evaluated sample according to a discriminator of the generative adversarial network model; and determine the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result.
13. A model fairness evaluation method applied to a model fairness evaluation platform, comprising: determining a true data probability distribution of a picture and / or text training sample according to a picture and / or text training model; receiving a to-be-evaluated sample and a to-be-evaluated model sent by a user; determining a credibility detection result of the to-be-evaluated sample according to the true data probability distribution and a generative adversarial network model; performing sample processing on the to-be-evaluated sample according to the credibility detection result to obtain an updated evaluation sample if the credibility detection result meets an untrusted condition; performing fairness evaluation on the to-be-evaluated model according to the to-be-evaluated sample and the updated evaluation sample; obtaining a fairness evaluation result of the to-be-evaluated model and returning the fairness evaluation result to the user; The determining the credibility detection result of the to-be-evaluated sample according to the real data probability distribution and the generative adversarial network model comprises: obtaining a sample data probability distribution of the to-be-evaluated sample according to the picture and / or text training model; determining a similarity of the to-be-evaluated sample belonging to the real data probability distribution of the training sample according to the sample data probability distribution; obtaining a sample prediction result of the to-be-evaluated sample according to a discriminant module of the generative adversarial network model; and determining the credibility detection result of the to-be-evaluated sample according to the similarity and the sample prediction result.
Citation Information
Patent Citations
Model security detection method and device and electronic equipment
CN107808098A
Security risk assessment method of depth learning model based on antagonistic samples
CN109034632A